Information processing apparatus and method

The apparatus generates synthetic data to assess AI prediction reliability and feature contributions, addressing the challenge of detecting environmental changes that degrade AI accuracy, thereby preventing damage in AI prediction systems.

JP7705737B2Active Publication Date: 2025-07-10HITACHI LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021091281
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-31
Publication Date
2025-07-10
Estimated Expiration
2041-05-31

AI Technical Summary

Technical Problem

Existing prediction systems using AI struggle to quickly detect environmental changes that cause accuracy deterioration, leading to potential damage due to delayed recognition of incorrect predictions, especially in critical scenarios like emergency vehicle dispatch.

Method used

An information processing apparatus and method that generates synthetic data by combining target and reference data, calculates reliability and contribution degrees of feature amounts, and outputs these contributions to detect environmental changes, enabling timely maintenance and prevention of damage.

Benefits of technology

Enables rapid detection of environmental changes affecting prediction accuracy, allowing for proactive maintenance to prevent potential harm, such as delayed emergency responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705737000001
    Figure 0007705737000001
  • Figure 0007705737000002
    Figure 0007705737000002
  • Figure 0007705737000003
    Figure 0007705737000003
Patent Text Reader

Abstract

To provide an information processing device and a method for preventing occurrence of damage caused by environmental changes that cause deterioration in accuracy of a prediction system.SOLUTION: In an information processing system in which a plurality of terminal devices and an information processing device are connected via a network, the information processing device 4 includes: a reference data database 26 that stores a plurality of items of reference data prepared in advance; a composite data generating unit 30 that generates each item of first composite data obtained by compositing target data and reference data based on the target data to be predicted and the plurality of items of reference data; a predictor 31 that performs prediction for each item of first composite data; a reliability degree calculation unit 32 that calculates a reliability degree of a prediction result for each item of first composite data; a reliability degree contribution degree calculation unit 33 that calculates a contribution degree of each feature amount of the target data with respect to the reliability degree of the prediction result for the first composite data; and an output unit 34 that outputs the contribution degree of each feature amount with respect to the calculated reliability degree of the prediction result for the first composite data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus and method, and is suitable for application to, for example, a prediction system utilizing AI (Artificial Intelligence).

Background Art

[0002] In recent years, the social penetration of AI has advanced, and many prediction systems utilizing AI have come to be operated. When operating such a system, it is necessary to prevent the occurrence of damage associated with the deterioration of the accuracy of AI due to environmental changes.

[0003] For example, when the jurisdiction of a certain fire station develops, the number of dispatches of emergency vehicles such as ambulances and fire trucks increases, and there is a possibility that a situation may occur where, when receiving a power supply for an emergency vehicle dispatch request, the emergency vehicle cannot be immediately directed to the site because it is already on a dispatch.

[0004] Therefore, for example, when constructing a prediction system that predicts, by AI, the time from receiving a power supply for a dispatch request of such an emergency vehicle until the emergency vehicle arrives at the site, it is necessary to appropriately perform maintenance of the prediction system as the target area develops.

[0005] If such maintenance is neglected, the accuracy of the AI may deteriorate and it may predict a shorter time than the actual arrival time of the emergency vehicle, and there is a risk that a situation where lives are lost may occur. It can be said that damage has already occurred when the deterioration of the accuracy of the AI is found.

[0006] Regarding this point, for example, Non-Patent Document 1 discloses a method for detecting the occurrence of environmental changes using a method called LossSHAP (Shapley Additive exPlanations). Specifically, it is disclosed that the occurrence of environmental changes is detected by observing the change over time in the contribution of each feature quantity of the data to be predicted to the prediction error of the AI. For example, if the contribution of the feature quantity "number of nearby hospitals" to the prediction error, which was previously low, has increased, this is regarded as the occurrence of an environmental change.

Prior Art Documents

Non-Patent Documents

[0007]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0008] By the way, in the technology disclosed in Non-Patent Document 1, since the prediction error of the AI is utilized, there is a problem that for cases where the correct value is obtained, the contribution of each feature quantity of the data to be predicted to the prediction error of the AI can only be calculated retrospectively. However, in actual cases, for example, in the review of housing loans, it may take a considerable amount of time to obtain the correct value, or in the prediction of the arrival time of an ambulance, serious damage may occur after the correct value is known, and it is impossible to wait for the correct value to be obtained.

[0009] The present invention has been made in consideration of the above points, and aims to propose an information processing apparatus and method that can quickly present information for detecting environmental changes that cause deterioration in the accuracy of a prediction system, and can prevent the occurrence of damage caused by such environmental changes.

Means for Solving the Problems

[0010] In order to solve such problems, in the present invention, in an information processing apparatus that presents information for detecting environmental changes in a prediction system using a machine learning model, based on target data to be predicted and a plurality of prepared reference data, a synthetic data generation unit that generates first synthetic data obtained by synthesizing the target data and the reference data respectively, a predictor that makes predictions for each of the first synthetic data, a reliability calculation unit that calculates the reliability of the prediction results of the predictor for each of the first synthetic data, a reliability contribution calculation unit that calculates the contribution degree of each feature amount of the target data to the reliability of the prediction result for the target data based on the reliability of the prediction results for each of the first synthetic data, and an output unit that outputs the contribution degree of each of the feature amounts to the reliability of the prediction result for the target data calculated by the reliability contribution calculation unit are provided.

[0011] Further, in the present invention, there is provided an information processing method executed by an information processing apparatus that presents information for detecting environmental changes in a prediction system using a machine learning model. The method includes a first step of generating first synthetic data obtained by synthesizing the target data and the reference data respectively based on the target data to be predicted and a plurality of prepared reference data, a second step of making predictions for each of the first synthetic data, and for each of the first synthetic data ru yuA third step of calculating the reliability of each measurement result, and a fourth step of calculating the contribution degree of each feature amount of the target data to the reliability of the prediction result of the target data based on the reliability of the prediction result for each of the first synthetic data, and a fifth step of outputting the contribution degree of each of the feature amounts to the reliability of the prediction result for the calculated target data are provided.

[0012] According to the information processing apparatus and method of the present invention, the user can recognize the presence or absence of the occurrence of an environmental change that causes deterioration of the prediction accuracy of the prediction system based on the contribution degree of each feature amount to the reliability of the prediction result for the presented target data, and when recognizing the occurrence of the environmental change, by performing maintenance of the prediction system, it is possible to prevent the occurrence of damage caused by the environmental change.

Effects of the Invention

[0013] According to the present invention, it is possible to quickly present information for detecting an environmental change that causes deterioration of the prediction accuracy of a prediction system, and to realize an information processing apparatus and method that can prevent the occurrence of damage caused by such an environmental change.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Embodiments for Carrying Out the Invention

[0015] The following describes an embodiment of the present invention in detail with reference to the drawings.

[0016] (1) First Embodiment (1-1) Configuration of the Information Processing System According to the Present Embodiment In FIG. 1, reference numeral 1 denotes an information processing system according to the present embodiment as a whole. This information processing system 1 is a system having a function (hereinafter referred to as an environmental change information presentation function) of providing a user with information for detecting an environmental change that causes deterioration in the prediction accuracy of AI in a prediction system utilizing AI. The information processing system 1 includes a plurality of terminal devices 3 connected via a network 2 and an information processing device 4.

[0017] The terminal device 3 is a computer device used by a user and is composed of a personal computer, a notebook personal computer, a tablet, or the like. The terminal device 3 executes processes such as transmitting necessary commands and data to the information processing device 4 according to a user operation, and displaying a screen based on screen data transmitted from the information processing device 4.

[0018] The information processing device 4 is composed of a general-purpose computer device including information processing resources such as a CPU 10, a main storage device 11, an auxiliary storage device 12, a communication device 13, an input device 14, and an output device 15.

[0019] The CPU 10 is an arithmetic device that comprehensively controls the operation of the entire information processing device 4, and is composed of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an AI chip, or the like.

[0020] The main memory device 11 is a semiconductor memory used as the working memory of the CPU 10, and is configured to include a ROM (Read Only Memory) and a RAM (Random Access Memory). The ROM is composed of a mask ROM, a PROM (Programmable ROM), etc., and the RAM is composed of an SRAM (Static RAM), an NVRAM (Non-Volatile RAM), a DRAM (Dynamic RAM), etc. The synthetic data generation program 20, the AI program 21, the reliability calculation program 22, the reliability contribution degree calculation program 23, and the output program 24, which will be described later, are read from the auxiliary storage device 12 when the information processing device 4 is started up or when necessary, and are stored and held in the main memory device 11.

[0021] The auxiliary storage device 12 is a non-volatile large-capacity storage device used for storing and holding programs and data to be stored long-term, and is composed of a hard disk device, a flash memory, an SSD (Solid State Drive), and / or an optical storage device, etc. As the optical storage device, a CD (Compact Disc) drive, a DVD (Digital Versatile Disc) drive, or a Blu-ray drive, etc. is used. The teacher data database 25 and the reference data database 26, which will be described later, are also stored and held in the auxiliary storage device 12.

[0022] The communication device 13 is a communication interface for communicating with the terminal device 3 via the network 2, and is composed of a NIC (Network Interface Card), a serial communication module, etc. As the communication device 13, in addition to a NIC, a serial communication module, etc., a USB (Universal Serial Interface) may be provided.

[0023] The input device 14 is a user interface for the user to input various instructions and information, and is composed of a keyboard, a mouse, a card reader, and / or a touch panel, etc. The output device 15 is a user interface that provides various information to the user visually and / or aurally, and is composed of a display device such as a liquid crystal display or an organic EL (Electro-Luminescence) display, a speaker, and / or a printer, etc.

[0024] (1-2) Environmental change information presentation function according to this embodiment Next, the environmental change information presentation function installed in the information processing device 4 will be described. In this regard, first, the Trust Score and SHAP (Shapley Additive exPlanations) will be described.

[0025] When an environmental change occurs, a large amount of data that the AI has not learned before (data unknown to the AI) begins to appear frequently. However, the AI makes a prediction without hesitation even if it lacks confidence. For this reason, the accuracy rate of the AI prediction decreases, and the prediction accuracy of the AI deteriorates.

[0026] In this case, when an environmental change that leads to a deterioration in the prediction accuracy of the AI occurs, the "degree of confidence" when the AI derives a predicted value also changes. As a method for evaluating such "degree of confidence" in the prediction accuracy of the AI, in recent years, many methods for calculating the reliability of the prediction results of predictions using machine learning models have been proposed. One of them is the "Trust Score".

[0027] The Trust Score is a method that, although limited to classification problems, calculates the comparison result between the distance from the target data (hereinafter referred to as the target data) to the closest data within the predicted class and the distance from the target data to the closest data outside the predicted class as the reliability of the prediction.

[0028] By applying this trust score to AI prediction, when recognizing a handwritten image of, for example, "4", a recognition result such as "the probability that the image is '4' is 90%, and the confidence level is 5.5 (= reliable)" can be obtained. When showing an image of a dog, a recognition result such as "the probability that the image is '4' is 90%, and the confidence level is 0.98 (= not reliable)" can be obtained.

[0029] Therefore, it is considered that environmental changes can be detected by monitoring the confidence level of the prediction results of AI prediction using such a trust score. However, as a practical problem, even if such a confidence level is constant, there may still be environmental changes.

[0030] On the other hand, even if such a confidence level seems unchanged, there are cases where signs are emerging at the level of the contribution degree of each feature amount of the target data that is the basis for the prediction result. Therefore, it is considered that by observing the contribution degree of each feature amount of the target data to this confidence level rather than this confidence level, environmental changes can be detected with higher accuracy.

[0031] Here, there is SHAP (SHapley Additive exPlanations) as a technology for calculating how much each feature amount (the value of each feature included in the target data) of the target data contributes to the prediction result of AI. By using this SHAP, for example, when the predicted time for ambulance deployment is 8 minutes for target data of "age = ○○, address = ××", an output such as "Regarding the deployment time, compared to the average of 10 minutes, 'age = ○○' has an impact of -3 minutes and 'address = ××' has an impact of +1 minute, and the prediction is 8 minutes" can be obtained.

[0032] In SHAP, a large amount of reference data is prepared separately from the target data, and a large amount of synthetic data is generated by replacing some of the feature amounts of each reference data with the corresponding feature amounts of the target data. Based on the generated synthetic data, AI is made to perform predictions, and based on the prediction results, the contribution degree of each feature amount of the target data to the prediction result is calculated respectively.

[0033] At this time, from the viewpoints of simplifying and accelerating the arithmetic processing, usually, synthetic data in which the feature amounts of the reference data and the target data are not much swapped (for example, synthetic data in which the number of feature amounts derived from the reference data is one or less) is preferentially generated. In the following, such a method for generating synthetic data will be referred to as the "conventional method of SHAP".

[0034] By using such a SHAP technique in combination with a technique for calculating the reliability of the prediction result of an AI prediction such as a trust score, it is possible to calculate the contribution degree of each feature amount to the reliability of the prediction result of the AI prediction, and it is presumed that environmental changes can be detected with higher accuracy by observing the contribution degree for each of these feature amounts. Here, the "contribution degree" is a value indicating how much each feature amount of the target data has affected the reliability.

[0035] Therefore, the information processing apparatus 4 according to the present embodiment generates synthetic data based on the target data and the reference data, calculates the reliability of the prediction result for each generated synthetic data, calculates the contribution degree of each feature amount of the target data to these calculated reliabilities, respectively, and has an environmental change information presentation function that presents these contribution degrees of the respective feature amounts to the user as information for detecting environmental changes. Note that a series of processes related to such an environmental change information presentation function are performed in parallel with the prediction process for the target data at the timing when the target data to be predicted is given from any of the terminal devices 3.

[0036] As means for realizing such an environmental change information presentation function, as shown in FIG. 1, a synthetic data generation program 20, an AI program 21, a reliability calculation program 22, a reliability contribution degree calculation program 23, and an output program 24 are stored in the main storage device 11 of the information processing apparatus 4, and a teacher data database 25 and a reference data database 26 are stored in the auxiliary storage device 12.

[0037] Details of the synthetic data generation program 20, AI program 21, reliability calculation program 22, reliability contribution degree calculation program 23, and output program 24 will be described later.

[0038] The teacher data database 25 is a database that stores a plurality of teacher data used when the predictor 31 described later learns target events such as the arrival time of emergency vehicles and insurance risks. As shown in FIG. 2, this teacher data database 25 has a table structure including an ID column 25A and a feature amount column 25B. In the teacher data database 25 of FIG. 2, one row corresponds to one teacher data.

[0039] The ID column 25A stores an identifier (teacher data ID) unique to the corresponding teacher data assigned to the teacher data. The feature amount column 25B is divided into a plurality of feature columns 25BA corresponding to each feature amount constituting the teacher data, and the value of the corresponding feature is stored as a feature amount in each feature column 25BA.

[0040] Therefore, in the case of the example in FIG. 2, in the teacher data to which the teacher data ID of "1" is assigned, the value (feature amount) of the feature ("feat_1") of "age" is "30", and the value (feature amount) of the feature ("feat_2") of "gender", which is "feature 2 (feat_2)", is "male", the value (feature amount) of the feature ("feat_3") of "height" is "170", the value (feature amount) of the feature ("feat_4") of "weight" is "64",..., and the value (feature amount) of the feature ("feat_N") of "blood pressure" is "120".

[0041] The reference data database 26 is a database that stores a plurality of reference data for generating the above synthetic data by swapping the target data and the feature amounts. In the case of the present embodiment, a part of the teacher data registered in the teacher data database 25 is stored in the reference data database 26 as reference data.

[0042] The reference data database 26 has the same configuration as the teacher data database 25. Specifically, as shown in FIG. 3, the reference data database 26 has a table structure including an ID column 26A and a feature amount column 26B. In the reference data database 26 of FIG. 3, one row corresponds to one piece of reference data.

[0043] And in the ID column 26A, an identifier unique to the corresponding reference data (reference data ID) assigned to the reference data is stored. Further, the feature amount column 26B is divided into a plurality of feature columns 26BA corresponding to the feature amounts of the respective features constituting the reference data, and the values of the corresponding features are stored as feature amounts in each feature column 26B.

[0044] FIG. 4 shows the logical configuration of the information processing apparatus 4 related to the environmental change information presentation function of the above-described embodiment. As shown in this FIG. 4, the information processing apparatus 4 includes a synthetic data generation unit 30, a predictor 31, a reliability calculation unit 32, and a reliability contribution calculation unit 33.

[0045] The synthetic data generation unit 30 is a functional unit realized by the CPU 10 (FIG. 1) of the information processing apparatus 4 executing the synthetic data generation program 20 (FIG. 1) stored in the main storage device 11 (FIG. 1). The synthetic data generation unit 30 has a function of generating a plurality of pieces of synthetic data obtained by synthesizing each piece of reference data stored in the reference data database 26 and data to be predicted (target data) for a predetermined item given from the terminal device 3 (FIG. 1) via the network 2 by the above-described conventional SHAP method. Then, the synthetic data generation unit 30 outputs the synthetic data generated in this way to the predictor 31, the reliability calculation unit 32, and the reliability contribution calculation unit 33.

[0046] The predictor 31 is a functional unit realized by the CPU 10 executing the AI program 21 (Fig. 1) stored in the main memory device 11. The predictor 31 holds a machine learning model generated by pre-performing machine learning on the reference data pre-registered in the reference data database 26, and has a function of inputting each synthetic data given from the synthetic data generation unit 30 into the machine learning model to perform predictions on these synthetic data. Then, the predictor 31 outputs the prediction results for each obtained synthetic data to the confidence calculation unit 32.

[0047] The confidence calculation unit 32 is a functional unit realized by the CPU 10 executing the confidence calculation program 22 (Fig. 1) stored in the main memory device 11. The confidence calculation unit 32 has a function of calculating the confidence of the prediction results for each synthetic data as an existing technique, for example, the above-mentioned trust score, based on each reference data stored in the reference data database 26, the target data given from the terminal device 3, and the prediction results for each synthetic data given from the predictor 31. The confidence calculation unit 32 outputs the calculated confidence of the prediction results for each synthetic data to the confidence contribution calculation unit 33.

[0048] The confidence contribution calculation unit 33 is a functional unit realized by the CPU 10 executing the confidence contribution calculation program 23 (Fig. 1) stored in the main memory device 11. The confidence contribution calculation unit 33 has a function of calculating the contribution degree of each feature amount of the target data to the confidence by an existing method, for example, a method similar to SHAP, which calculates the contribution degree of the feature amount of the perturbation-based based on each synthetic data given from the synthetic data generation unit 30, the prediction results for each synthetic data by the predictor 31, and the confidence of the prediction results of the predictor 31 for each synthetic data given from the confidence calculation unit 32. Then, the confidence contribution calculation unit 33 outputs the calculated contribution degree for each feature amount to the output unit 34.

[0049] The output unit 34 is a functional unit realized by the CPU 10 executing the output program 24 stored in the main memory device 11. The output unit 34 generates screen data of a reliability contribution degree calculation result screen 40, which will be described later with reference to FIG. 5, based on the contribution degrees of each feature amount with respect to the reliability given from the reliability contribution degree calculation unit 33, and has a function of transmitting the generated screen data to the corresponding terminal device 3. As a result, based on this screen data, the reliability contribution degree calculation result screen 40 is displayed on the terminal device 3.

[0050] FIG. 5 shows a configuration example of the reliability contribution degree calculation result screen 40. In this configuration example, the reliability contribution degree calculation result screen 40 includes a contribution degree display area 41 for each feature and an explanation display area 42.

[0051] In the contribution degree display area 41 for each feature, the magnitudes of the contribution degrees of each feature amount of the target data with respect to the reliability of the prediction result of the predictor 31 calculated by the reliability contribution degree calculation unit 33 are displayed as the magnitudes of bar graphs. In the example of FIG. 5, the feature amounts of the target data are "age", "gender", "height", "weight", and "blood pressure". Among them, each feature amount of "age", "gender", and "weight" contributes in the direction of increasing the reliability of the prediction result of the predictor, and it is shown that "height" and "blood pressure" contribute in the direction of decreasing such reliability.

[0052] In the explanation display area 42, text representing an explanation about the contribution degree of each feature amount of the target data with respect to the reliability of the prediction result of the predictor 31 is displayed. In the case of the example of FIG. 5, as is clear from the graphs of each feature amount displayed in the contribution degree display area 41 for each feature, among the contribution degrees of each feature amount with respect to such reliability, since the magnitude of the contribution of "age" to such reliability is the largest, an example is shown where the explanation "Age has a great influence on the reliability." is displayed.

[0053] Therefore, based on the contribution degrees of each feature amount displayed on the reliability contribution degree calculation result screen 40 displayed on the terminal device 3, the user can recognize that some environmental change has occurred when the contribution degree of any feature amount to the reliability of the prediction result of the predictor 31 has fluctuated significantly compared to before.

[0054] However, it is also possible to provide a functional unit that observes the change over time of the contribution degree of each feature amount to the reliability of the prediction result of the predictor 31, and when the change amount of the change over time of any feature amount exceeds a certain threshold value, notifies the user by displaying a warning to that effect on the terminal device 3 corresponding thereto.

[0055] (1-3) Processing of each functional unit related to the environmental change information presentation function Next, the specific processing contents of each process executed by the composite data generation unit 30 and the reliability contribution degree calculation unit 33 of the information processing apparatus 4 in relation to the environmental change information presentation function according to the present embodiment will be described. In the following, the processing subject of each process will be described as the composite data generation unit 30 or the reliability contribution degree calculation unit 33, but in practice, it goes without saying that the CPU 10 of the information processing apparatus 4 executes the process based on the corresponding program (composite data generation program 20 or reliability contribution degree calculation program 23).

[0056] (1-3-1) Composite data generation process FIG. 6 shows the flow of the composite data generation process executed by the composite data generation unit 30 in relation to such an environmental change information presentation function. The composite data generation unit 30 generates composite data according to the processing procedure shown in FIG. 6.

[0057] In practice, when the composite data generation unit 30 is given target data and an instruction to execute a prediction on the target data from any one of the terminal devices 3 in response to a user operation, the composite data generation process shown in FIG. 6 is started.

[0058] Then, the synthetic data generation unit 30 first selects one piece of reference data whose steps after step S2 are unprocessed from the reference data stored in the reference data database 26 (S1). Also, the synthetic data generation unit 30 uses the reference data selected in step S1 to generate one or more synthetic data by, for example, a conventional SHAP method (S2), and outputs the generated synthetic data to the predictor 31, the reliability calculation unit 32, and the reliability contribution calculation unit 33, respectively (S3).

[0059] After that, the synthetic data generation unit 30 determines whether or not the process of step S2 (synthetic data generation process) has been completed for all or a predetermined number of reference data registered in the reference data database 26 (S4). If the synthetic data generation unit 30 obtains a negative result in this determination, it returns to step S1, and then repeats the processes of steps S1 to S4 while sequentially switching the reference data selected in step S1 to other reference data for which step S2 is unprocessed.

[0060] When the synthetic data generation unit 30 finally obtains an affirmative result in step S4 by finishing generating synthetic data based on all or a predetermined number of reference data registered in the reference data database 26, it ends this synthetic data generation process.

[0061] (1-3-2) Reliability calculation process On the other hand, FIG. 7 shows the reliability calculation process executed by the reliability calculation unit 32 in relation to such an environmental change information presentation function. The reliability calculation unit 32 calculates the reliability of the prediction results for each synthetic data according to the processing procedure shown in this FIG. 7.

[0062] In practice, when each synthetic data is given from the synthetic data generation unit 30 and the prediction results for these synthetic data are given from the predictor 31, the reliability calculation unit 32 starts the reliability calculation process shown in this FIG. 7. First, it selects one piece of synthetic data whose steps after step S11 are unprocessed from the synthetic data sequentially given from the synthetic data generation unit 30 (S10).

[0063] Subsequently, the reliability calculation unit 32 calculates the reliability of the prediction result for the synthetic data selected in step S10 (hereinafter referred to as selected synthetic data) (S11). In the present embodiment, the reliability calculation unit 32 calculates the trans score of the prediction result as such reliability.

[0064] Next, the reliability calculation unit 32 determines whether or not the process of step S11 has been completed for all the synthetic data (S12). If the reliability calculation unit 32 obtains a negative result in this determination, it returns to step S10, and then repeats the processes of steps S10 to S12 while sequentially switching the synthetic data selected in step S10 to other synthetic data for which step S11 has not been processed.

[0065] Then, when the reliability calculation unit 32 finally obtains an affirmative result in step S12 by calculating the reliability of the prediction result for all the synthetic data given from the synthetic data generation unit 30, it ends this reliability calculation process.

[0066] (1-3-3) Reliability contribution calculation process On the other hand, FIG. 8 shows a reliability contribution calculation process executed by the reliability contribution calculation unit 33 in relation to such an environmental change information presentation function. The reliability contribution calculation unit 33 calculates the contribution of each feature amount to the reliability of the prediction result for the target data according to the processing procedure shown in FIG. 8.

[0067] Actually, when all the synthetic data are given from the synthetic data generation unit 30 and the reliability of each prediction result of the predictor 31 for these synthetic data is given from the reliability calculation unit 32, the reliability contribution calculation unit 33 starts the reliability contribution calculation process shown in FIG. 8.

[0068] Then, the reliability contribution calculation unit 33 calculates the contribution of each feature amount of the target data to the reliability of the prediction result of the target data by using an existing method (e.g., SHAP) for calculating the contribution of the perturbation-based feature amount (S15). Then, the reliability contribution calculation unit 33 outputs the calculated contribution of each feature amount to the output unit 34 (S16), and then ends this reliability contribution calculation process.

[0069] (1-4) Effects of this embodiment As described above, in the information processing apparatus 4 of this embodiment, synthetic data is generated based on the target data and the reference data, the reliability of the prediction result for each generated synthetic data is calculated respectively, and based on these calculated reliabilities, the contribution of each feature amount of the target data to the reliability of the prediction result for the target data is calculated respectively, and a reliability contribution calculation result screen 40 in which the contributions of these feature amounts are displayed is displayed.

[0070] Therefore, the user can recognize the presence or absence of the occurrence of an environmental change that causes deterioration in the accuracy of AI prediction based on such contributions for each feature amount of the target data displayed on the reliability contribution calculation result screen 40. When the occurrence of the environmental change is recognized, by performing maintenance of the AI, it is possible to prevent the occurrence of damage caused by the environmental change.

[0071] As described above, according to this embodiment, it is possible to quickly present information for detecting an environmental change that causes deterioration in the accuracy of AI prediction, and to realize an information processing apparatus that can prevent the occurrence of damage caused by such an environmental change.

[0072] (2) Second embodiment FIG. 9, which shows corresponding parts with the same reference numerals as in FIG. 1, shows an information processing system 50 according to the second embodiment. This information processing system 50 is the same as the information processing system 1 of the first embodiment, except that a similarity determination program 52 and a similarity calculation program 53 are additionally stored in the main storage device 11 of the information processing device 51, a similarity information database 54 is additionally stored in the auxiliary storage device 12 of the information processing device 51, and the function of the composite data generation program 55 is different.

[0073] The functions of the similarity determination program 52, the similarity calculation program 53, and the composite data generation program 55 will be described later.

[0074] The similarity information database 54 is a database that stores the determination results as to whether or not there is similarity between all or a predetermined number of reference data registered in the reference data database 26, which is determined by a similarity determination unit 61 (FIG. 11) described later, and the data to be predicted (target data) given from the terminal device 3.

[0075] As shown in FIG. 10, this similarity information database 54 has a table structure including an ID column 54A and a similarity column 54B. In the similarity information database 54 of FIG. 10, one row corresponds to one reference data registered in the reference data database 26.

[0076] The reference data ID of the corresponding reference data is stored in the ID column 54A. In the similarity column 54B, "1" is stored when the corresponding reference data is similar to the target data, and "0" is stored when it is not similar.

[0077] Therefore, in the example of FIG. 10, it is shown that the reference data given the reference data ID of "1" is not similar to the target data given from the terminal device 3 at that time, and the reference data given the reference data ID of "2" is determined by the similarity determination unit 61 (FIG. 11) to be similar to such target data.

[0078] FIG. 11, which shows corresponding parts with the same reference numerals as FIG. 4, shows the logical configuration of the information processing apparatus 51 related to the environmental change information presentation function according to the present embodiment. As shown in FIG. 11, in addition to the predictor 31, the reliability calculation unit 32, the reliability contribution calculation unit 33, and the output unit 34, the information processing apparatus 51 includes a similarity calculation unit 60, a similarity determination unit 61, and a composite data generation unit 62.

[0079] The similarity calculation unit 60 is a functional unit realized by the CPU 10 (FIG. 9) of the information processing apparatus 51 executing a similarity calculation program 53 (FIG. 9) stored in the main storage device 11 (FIG. 9). The similarity calculation unit 60 has a function of calculating the similarity between the target data and the reference data given from the similarity determination unit 61 described later by an existing method. The similarity calculation unit 60 outputs the calculated similarity between the target data and the reference data to the similarity determination unit 61.

[0080] The similarity determination unit 61 is a functional unit realized by the CPU 10 of the information processing apparatus 51 executing a similarity determination program 52 (FIG. 9) stored in the main storage device 11. The similarity determination unit 61 has a function of outputting all or a predetermined number of reference data registered in the reference data database 26 and the target data to be predicted given from the terminal device 3 to the similarity calculation unit 60. Based on the similarity between each reference data and the target data calculated by the similarity calculation unit 60 as a result, the similarity determination unit 61 determines the presence or absence of similarity between the reference data and the target data, and registers the determination result in the similarity information database 54.

[0081] The composite data generation unit 62 has a function of generating composite data between the reference data and the target data while switching the composite method according to the presence or absence of similarity between the reference data and the target data registered in the similarity information database 54 for all or a predetermined number of reference data stored in the reference data database 26.

[0082] In practice, for reference data that is not similar to the target data, in order to generate synthetic data that does not depend on the mixing ratio, as shown in FIG. 12 for example, in the entire finally generated synthetic data, the number of feature quantities derived from the reference data is uniformly distributed without bias in the number of feature quantities derived from the reference data. The synthetic data is generated by swapping the feature quantities of the reference data and the corresponding feature quantities of the target data. Further, for reference data that is similar to the target data, the synthetic data generation unit 62 generates synthetic data using the reference data by the conventional SHAP method.

[0083] Note that switching the synthetic data generation method depending on whether the target data and the reference data are similar or not is to accurately calculate the contribution degree of each feature quantity of the target data to the reliability of the prediction result calculated for the synthetic data while improving the efficiency.

[0084] In practice, synthetic data with a low mixing ratio between the feature quantities of the target data and the feature quantities of the reference data (almost the target data or almost the reference data) has high reliability, and synthetic data with a high mixing ratio tends to have low reliability. Therefore, the conventional SHAP method will generate synthetic data with high reliability in a biased manner, and it is impossible to accurately calculate the contribution degree of each feature quantity of the target data to the reliability of the prediction result calculated for the synthetic data.

[0085] Therefore, when the target data and the reference data are not similar, in the entire finally generated synthetic data, the synthetic data is generated by swapping the feature quantities of the reference data and the corresponding feature quantities of the target data so that the number of feature quantities derived from the reference data is uniformly distributed without bias in the number of feature quantities derived from the reference data, thereby generating synthetic data in which synthetic data with high reliability and synthetic data with low reliability exist in the same degree, and thereby improving the accuracy of the contribution degree of each feature quantity of the target data to such reliability calculated by the reliability contribution degree calculation unit 33.

[0086] On the other hand, when the target data and the reference data are similar, the synthetic data generated by swapping any number of the feature amounts of the target data and the feature amounts of the reference data does not change much. Therefore, from the viewpoints of simplifying and accelerating the arithmetic processing, synthetic data is generated by the conventional SHAP method.

[0087] Then, the synthetic data generation unit 62 outputs the generated synthetic data to the predictor 31, the reliability calculation unit 32, and the reliability contribution calculation unit 33, respectively.

[0088] FIG. 13 shows the processing contents of the similarity determination process executed by the similarity determination unit 61 (FIG. 11) of the information processing apparatus 51 in relation to the environment change information presentation function of the present embodiment. The similarity determination unit 61 determines the presence or absence of similarity between each reference data and the target data according to the processing procedure of FIG. 13.

[0089] In practice, when the target data is given from any of the terminal devices 3, the similarity determination unit 61 starts the similarity determination process shown in FIG. 13. First, one piece of reference data whose subsequent steps S21 and later are unprocessed is selected from the reference data registered in the reference data database 26 (S20).

[0090] Subsequently, the similarity determination unit 61 requests the similarity calculation unit 60 (FIG. 11) to calculate the similarity of the reference data selected in step S20 with respect to the target data (hereinafter, in the description of FIG. 13, this is referred to as the selected reference data) (S21). As a result, the similarity between the target data and the selected reference data is calculated by the similarity calculation unit 60 and notified to the similarity determination unit 61.

[0091] When the similarity determination unit 61 is notified of such similarity from the similarity calculation unit 60, it determines whether the target data and the selected reference data are similar based on the notified similarity (S22), and registers the determination result in the similarity information database (S23).

[0092] Specifically, the similarity determination unit 61 compares the similarity notified from the similarity calculation unit 60 with a preset threshold value (hereinafter referred to as the similarity determination threshold value). When the similarity is equal to or greater than the similarity determination threshold value, the similarity determination unit 61 determines that the selection reference data and the target data are similar, and stores "1" in the similarity column 54B (FIG. 10) of the row corresponding to the selection reference data in the similarity information database 54. When the similarity is less than the similarity determination threshold value, the similarity determination unit 61 determines that the selection reference data and the target data are not similar, and stores "0" in the similarity column 54B of the row corresponding to the selection reference data in the similarity information database 54.

[0093] Next, the similarity determination unit 61 determines whether or not the processing after step S21 has been completed for all the reference data stored in the reference data database 26 (S24). If the determination results in a negative result, the similarity determination unit 61 returns to step S20, and thereafter, while sequentially switching the reference data selected in step S20 to other unprocessed reference data, repeats the processing of steps S20 to S24.

[0094] When the similarity determination unit 61 finally obtains an affirmative result in step S24 by determining the presence or absence of similarity with the target data for all the reference data stored in the reference data database 26, the similarity determination process ends.

[0095] On the other hand, FIG. 14 shows the processing content of the synthetic data generation process executed by the synthetic data generation unit 62 in relation to the environmental change information presentation function of the present embodiment. The synthetic data generation unit 62 generates synthetic data based on each reference data stored in the reference data database 26 according to the processing procedure shown in FIG. 13.

[0096] In practice, when the similarity determination unit 61 finishes determining the presence or absence of similarity with the target data for all or a predetermined number of reference data registered in the reference data database 26, the synthetic data generation unit 62 starts the synthetic data generation process shown in FIG. 14. First, it selects one piece of reference data whose steps after step S31 are unprocessed from among the reference data stored in the reference data database 26 (S30).

[0097] Subsequently, the synthetic data generation unit 62 refers to the similarity information database 54 (FIG. 10) to determine whether the reference data selected in step S30 (hereinafter referred to as the selected reference data in the description of FIG. 14) is similar to the target data (S31).

[0098] When the synthetic data generation unit 62 obtains an affirmative result in this determination, it generates synthetic data using the selected reference data by the above-described conventional method (S32), and outputs the generated synthetic data to the predictor 31, the reliability calculation unit 32, and the reliability contribution calculation unit 33, respectively (S34).

[0099] On the other hand, when the synthetic data generation unit 62 obtains a negative result in the determination of step S31, it generates synthetic data so that the number of features derived from the reference data is uniform without bias in the number of features derived from the reference data (S33), and outputs the generated synthetic data to the predictor 31, the reliability calculation unit 32, and the reliability contribution calculation unit 33, respectively (S34).

[0100] Next, the synthetic data generation unit 62 determines whether it has finished executing the processing after step S31 (synthetic data generation processing) for all or a preset predetermined number of reference data registered in the reference data database 26 (S35). When the synthetic data generation unit 62 obtains a negative result in this determination, it returns to step S30, and thereafter, repeats the processing of steps S30 to S35 while sequentially switching the reference data selected in step S30 to other unprocessed reference data whose steps after step S31 are unprocessed.

[0101] When the synthetic data generation unit 62 finally finishes generating synthetic data based on all or a preset number of reference data registered in the reference data database 26 and obtains an affirmative result in step S35, this synthetic data generation process ends.

[0102] As described above, in the information processing apparatus 51 of the present embodiment, by determining the presence or absence of similarity between the target data and each reference data and switching the generation method of the synthetic data obtained by synthesizing the target data and the reference data based on whether the reference data is similar to the target data, in addition to the effects obtained by the first embodiment, while improving efficiency, it is also possible to obtain the effect of accurately calculating the contribution degree of each feature amount of the target data to the reliability of the prediction result calculated for the synthetic data.

[0103] (3) Third Embodiment FIG. 15, in which the corresponding parts to FIG. 9 are denoted by the same reference numerals, shows an information processing system 70 according to the third embodiment. This information processing system 70 is different from the information processing system 50 according to the second embodiment in that, in addition to the environmental change information presentation function similar to that of the first embodiment, it is equipped with a weak point tendency presentation function that analyzes the weak point tendency (tendency of feature amounts with low reliability of prediction results) of the predictor 31 (FIG. 17) before the start of operation and presents it to the user.

[0104] Actually, in this information processing system 70, in addition to the configuration of the information processing apparatus 4 of the first embodiment described above with respect to FIG. 9, a data selection program 72 and a weak point tendency analysis program 73 are stored in the main storage device 11 of the information processing apparatus 71, and a reliability contribution degree database 74 is stored in the auxiliary storage device 12 of the information processing apparatus 71. Details of the data selection program 72 and the weak point tendency analysis program 73 will be described later.

[0105] The reliability contribution database 74 is a database used to store and hold the contribution degrees of each feature quantity with respect to the reliability of the prediction result of the AI prediction for the teacher data selected as temporary target data (hereinafter referred to as this temporary target data) by the reliability contribution calculation unit 33 (FIG. 17) for each teacher data as described later. As shown in FIG. 16, the reliability contribution database 74 has a table structure including an ID column 74A and a feature quantity column 74B. In the reliability contribution database 74 of FIG. 16, one row corresponds to one temporary target data.

[0106] And in the ID column 74A, an identifier unique to the corresponding temporary target data (temporary target data ID) given to the corresponding temporary target data is stored. Further, the feature quantity column 74B is divided into a plurality of feature columns 74BA corresponding to each feature constituting the temporary target data, and in these feature columns 74BA, the contribution degrees of the corresponding feature quantities of the temporary target data with respect to the reliability of the prediction result of the predictor 31 for the temporary target data, calculated by the reliability contribution calculation unit 33 as described later, are respectively stored.

[0107] Therefore, in the case of the example of FIG. 16, it is shown that the contribution degree of the value (feature quantity) of the feature "age" with respect to the reliability of the prediction result of the AI prediction for the temporary target data given the temporary target data ID of "1" is "+5", the contribution degree of the value (feature quantity) of the feature "gender" is "+5", the contribution degree of the value (feature quantity) of the feature "height" is "+3", the value (feature quantity) of the feature "weight" is "+7", ……, and the contribution degree of the value (feature quantity) of the feature "blood pressure" is "+2".

[0108] FIG. 17, in which the same reference numerals are given to the corresponding parts as in FIG. 4, shows the logical configuration of the information processing apparatus 71 regarding the disliked tendency analysis function of the present embodiment. Since the logical configuration of the present information processing apparatus 71 regarding the environmental change information presentation function is the same as the logical configuration of the information processing apparatus 4 of the first embodiment described above with respect to FIG. 4, the illustration and description here are omitted.

[0109] As shown in FIG. 17, the information processing apparatus 71 includes a data selection unit 80, a synthetic data generation unit 30, a predictor 31, a reliability calculation unit 32, a reliability contribution calculation unit 33, a dislike tendency analysis unit 81, and an output unit 82 in relation to the dislike tendency analysis function.

[0110] The data selection unit 80 is a functional unit realized by the CPU 10 of the information processing apparatus 71 executing a target data selection program 72 (FIG. 15) stored in the main storage device 11. The data selection unit 80 selects one piece of teacher data from the teacher data registered in the teacher data database 25 as temporary target data (hereinafter referred to as this temporary target data), and selects a predetermined number of teacher data other than this temporary target data as temporary reference data in advance, and transmits these temporary target data and each temporary reference data to the synthetic data generation unit 30. Further, the data selection unit 80 also outputs each temporary reference data to the reliability calculation unit 32, the reliability contribution calculation unit 33, and the dislike tendency analysis unit 81.

[0111] And then, based on these respective temporary reference data and the temporary target data, the synthetic data generation unit 30, the predictor 31, the reliability calculation unit 32, and the reliability contribution calculation unit 33 each execute the respective processes described above with respect to FIG. 4. As a result, the reliability contribution calculation unit 33 calculates the contribution of each feature amount of the temporary target data to the reliability of the prediction result of the temporary target data, and these contributions are respectively registered in the reliability contribution database 74.

[0112] Similarly, for a plurality of mutually different temporary target data, the contribution of each feature amount of the temporary target data to the reliability of the prediction result is calculated respectively, and the calculation results are respectively stored in the reliability contribution database 74.

[0113] The difficulty tendency analysis unit 81 is a functional unit realized by the CPU 10 of the information processing apparatus 71 executing a difficulty tendency analysis program 73 (FIG. 15) stored in the main memory device 11. The difficulty tendency analysis unit 81 calculates, based on the contribution degrees of the feature amounts of the respective features of each temporary target data to the reliability of the prediction result of the temporary target data registered in the reliability contribution degree database 74, the average value of the contribution degrees to such reliability for each of these categories when the feature amounts of the feature are divided into a plurality of categories.

[0114] Specifically, for features whose feature amounts can take continuous values as features, such as "age" and "height", the difficulty tendency analysis unit 81 divides the feature amounts of the feature into a plurality of continuous categories, such as "0 to 10 years old", "10 to 20 years old", "20 to 30 years old",..., "90 to 100 years old" and "100 years old and above", or "0 to 100 cm", "100 to 110 cm", "110 to 120 cm",..., "190 to 200 cm" and "200 cm and above", refers to the reliability contribution degree database 74, and calculates the average value of the contribution degrees of the feature amounts of these categories to such reliability respectively. Also, for features whose feature amounts can take non - continuous values as features, such as "gender", the difficulty tendency analysis unit 81 calculates the average value of the contribution degrees of the feature amounts to such reliability for each value ("male" and "female") respectively.

[0115] Then, the difficulty tendency analysis unit 81 outputs to the output unit 82 the contribution degrees to such reliability for each category of the feature amounts of each feature calculated in this way.

[0116] The output unit 82 is a functional unit realized by the CPU 10 of the information processing apparatus 71 executing an output program 75 (FIG. 15) stored in the main memory device 11. The output unit 82 generates, based on the contribution degrees to such reliability for each category of the feature amounts of each feature notified from the difficulty tendency analysis unit 81, screen data of a difficulty tendency analysis result screen 90 as shown in FIG. 18 for example, and transmits the generated screen data to the corresponding terminal device 3. Thus, based on this screen data, such a difficulty tendency analysis result screen 90 is displayed on the terminal device 3.

[0117] This difficulty tendency analysis result screen 90 is configured to include a feature selection pull-down button 91, a selected feature display column 92, and a difficulty tendency analysis result display area 93. On the difficulty tendency analysis result screen 90, by clicking the feature selection pull-down button 91, a pull-down menu 94 in which all features included in the teacher data and the target data are listed can be displayed.

[0118] Thus, the user selects by clicking or tapping, etc., the desired feature at that time from among the features listed in the pull-down menu 94. A character string representing the name of the feature selected at this time is displayed in the selected feature display column 92.

[0119] In addition, in the difficulty tendency analysis result display area 93, for the feature selected at this time (the feature whose name is displayed in the selected feature display column 92), the average value of the contribution degree to such reliability for each category of the feature amount of the feature is displayed as a bar graph with a length and direction corresponding to the average value.

[0120] Thus, the user can confirm, based on the magnitude of the contribution degree to such reliability for each category of the feature amount of the feature displayed in the difficulty tendency analysis result display area 93, the difficulty tendency of the predictor 31, for example, the category for each feature amount that is a factor reducing the reliability of the prediction result of the predictor 31.

[0121] FIG. 19 shows the processing content of the data selection process executed by the data selection unit 80 in relation to the difficulty tendency presentation function. The data selection unit 80 selects pseudo-target data and pseudo-reference data from the teacher data stored in the teacher data database 25 according to the processing procedure shown in FIG. 19 and outputs them to the composite data generation unit 30 and the like.

[0122] In practice, the data selection unit 80 starts the data selection process shown in FIG. 19 in response to a request from, for example, any one of the terminal devices 3. First, it selects any one piece of teacher data from the teacher data stored in the teacher data database 25 as pseudo-target data (S40).

[0123] Subsequently, the data selection unit 80 selects, as provisional reference data, a predetermined number of teacher data other than the teacher data selected in step S40 from among the teacher data stored in the teacher data database 25 (S41).

[0124] Then, the data selection unit 80 transmits the teacher data (provisional target data) selected in step S40 and each teacher data (provisional reference data) selected in step S41 to the composite data generation unit 30 and the dislike tendency analysis unit 81, and outputs each teacher data (provisional reference data) selected in step S41 to the reliability calculation unit 32 and the reliability contribution calculation unit 33, respectively (S42). After that, this target data selection process ends.

[0125] Thus, thereafter, using these provisional target data and provisional reference data, the same processes as those in the first embodiment are respectively executed in the composite data generation unit 30, the predictor 31, the reliability calculation unit 32, and the reliability contribution calculation unit 33. As a result, the contribution degree of each feature amount to the reliability of the prediction result for the provisional target data obtained is calculated by the reliability contribution calculation unit 33 and stored in the reliability contribution database 74.

[0126] Note that the data selection unit 80 repeats the process of FIG. 19 a predetermined number of times while sequentially switching the teacher data selected as the provisional target data to other teacher data. As a result, the contribution degree of each feature amount to the reliability of the prediction result for a plurality of provisional target data is calculated by the reliability contribution calculation unit 33 each time and stored in the reliability contribution database 74.

[0127] On the other hand, FIG. 20 shows the processing content of the dislike tendency analysis process executed by the dislike tendency analysis unit 81 in relation to the dislike tendency presentation function. The dislike tendency analysis unit 81 analyzes the dislike tendency (tendency of data with low prediction reliability) of the predictor 31 according to the processing procedure shown in FIG. 20.

[0128] Actually, when the contribution degree of each feature amount to the reliability of the prediction results of a predetermined number of synthetic data is registered in the reliability contribution degree database 74 (FIG. 16), the difficulty tendency analysis unit 81 starts the difficulty tendency analysis process shown in FIG. 20. First, one feature whose process feature is from step S51 and later is selected from among the features for which the contribution degree of the feature amount is registered in the reliability contribution degree database 74 (S50).

[0129] Subsequently, the difficulty tendency analysis unit 81 determines whether the value (feature amount) of the feature selected in step S50 (hereinafter referred to as the selected feature) can take continuous values (S51). If the difficulty tendency analysis unit 81 obtains a negative result in this determination, it proceeds to step S53.

[0130] On the other hand, when the difficulty tendency analysis unit 81 obtains an affirmative result in the determination of step S51, it classifies the feature amount range of the selected feature into a plurality of categories by dividing it into a plurality of sections (S52). Then, the difficulty tendency analysis unit 81 selects one unprocessed category from among the categories classified in step S52 for which the process from step S54 and later is yet to be performed (S53).

[0131] Subsequently, for each value (feature amount) of the feature included in the category selected in step S53 (hereinafter referred to as the selected category), the difficulty tendency analysis unit 81 calculates the contribution degree to the reliability of the prediction result for the temporary target data, and calculates the average value of these contribution degrees in the selected category based on the calculation result (S54).

[0132] Next, the difficulty tendency analysis unit 81 determines whether it has finished executing the process of step S54 for all categories of the selected feature (S55). If the difficulty tendency analysis unit 81 obtains a negative result in this determination, it returns to step S53. After that, while sequentially switching the category selected in step S53 to other categories for which step S54 is unprocessed, the processes of steps S53 to S55 are repeated.

[0133] Then, when the disliked tendency analysis unit 81 finally calculates the average value of the contribution degrees to the reliability of the prediction results for the temporary target data in each category of all the selected features and obtains an affirmative result in step S55, it determines whether the processing after step S51 has been completed for all features (S56).

[0134] If the disliked tendency analysis unit 81 obtains a negative result in this determination, it returns to step S50. After that, while sequentially switching the features selected in step S50 to other unprocessed features after step S51, it repeats the processing of steps S50 to S56 in the same manner as described above.

[0135] Then, when the disliked tendency analysis unit 81 finally obtains an affirmative result in step S56 by completing the processing after step S51 for all features, it outputs to the output unit 82 the average value of the contribution degrees to the reliability of the prediction results for the temporary target data in each category of each feature obtained by the processing of steps S50 to S56 (S57). After that, it ends this disliked tendency analysis process.

[0136] As described above, the information processing apparatus 71 of the present embodiment analyzes the disliked tendency of the predictor 31 and causes the terminal device 3 to display the disliked tendency analysis result screen 90 based on the analysis result. Therefore, the user can recognize the disliked tendency of the predictor 31 based on the disliked tendency analysis result screen 90 displayed on the terminal device 3. Thus, according to this information processing apparatus 71, the user can determine to what extent the prediction result for the subsequent target data can be trusted based on such a recognition result.

[0137] (4) Other embodiments In the above-described first to third embodiments, the case where the environmental change information presentation function according to each embodiment is mounted on one information processing apparatus has been described. However, the present invention is not limited to this, and such an environmental change information presentation function may be decomposed into a plurality of functions, and each function may be mounted on different computer devices constituting a distributed computing system.

[0138] Also, in the above-described first to third embodiments, the case where the reliability of each composite data calculated by the reliability calculation unit 32 is calculated using the trust score technique has been described. However, the present invention is not limited to this, and such reliability may be calculated using a technique other than the trust score, for example, a technique such as Dropout.

[0139] Similarly, in the above-described first to third embodiments, the case where the contribution degree of each feature amount of the target data to the reliability is calculated using the SHAP technique has been described. However, the present invention is not limited to this. In short, as long as it is a technique capable of calculating the contribution degree of perturbation-based feature amounts, a technique other than SHAP, such as LIME (Locally Interpretable Model-agnostic Explanations), may be applied.

[0140] Furthermore, in the above-described first to third embodiments, the case where the output units 34 and 82 display the contribution degree of each feature amount of the target data to the reliability calculated by the reliability contribution degree calculation unit 33 and the analysis result of the dislike tendency analysis unit 81 on the terminal device 3 to present it to the user has been described. However, the present invention is not limited to this. For example, it may be printed out or output as voice, and various other presentation methods can be applied as the method of presenting this information to the user.

[0141] Furthermore, in the above-described third embodiment, the case where the dislike tendency presentation function of the third embodiment is applied to the information processing device 71 equipped with the environmental change information presentation function similar to that of the first embodiment has been described. However, the present invention is not limited to this, and the dislike tendency presentation function of the third embodiment may be applied to an information processing device equipped with the environmental change information presentation function similar to that of the second embodiment.

Industrial Applicability

[0142] The present invention can be widely applied to a prediction system utilizing a machine learning model.

Explanation of Signs

[0143] 1, 50, 70... information processing systems, 3... terminal devices, 4, 51, 71... information processing devices, 10... CPUs, 20, 55... synthetic data generation programs, 21... AI programs, 22... reliability calculation programs, 23... reliability contribution calculation programs, 24, 75... output programs, 25... teacher data databases, 26... reference data databases, 30, 62... synthetic data generation units, 31... predictors, 32... reliability calculation units, 33... reliability contribution calculation units, 34, 82... output units, 40... reliability contribution calculation result screens, 52... similarity determination programs, 53... similarity calculation programs, 54... similarity information databases, 60... similarity calculation units, 61... similarity determination units, 72... data selection programs, 73... disliked tendency analysis programs, 74... reliability contribution databases, 80... data selection units, 81... disliked tendency analysis units, 90... disliked tendency analysis result screens.

Claims

1. In an information processing apparatus that presents information for detecting an environmental change in a prediction system using a machine learning model, a synthetic data generation unit that generates first synthetic data obtained by synthesizing the target data and the reference data, respectively, based on the target data to be predicted and a plurality of reference data prepared in advance; a predictor that performs a prediction on each of the first synthetic data; a reliability calculation unit that calculates the reliability of the prediction result of the predictor for each of the first synthetic data; a reliability contribution calculation unit that calculates the contribution degree of each feature amount of the target data to the reliability of the prediction result for the target data based on the reliability of the prediction result for each of the first synthetic data; and an output unit that outputs the contribution degree of each feature amount to the reliability of the prediction result for the target data calculated by the reliability contribution calculation unit. An information processing apparatus characterized by comprising the above.

2. The information processing apparatus further includes a similarity determination unit that determines the presence or absence of similarity between the target data and each of the reference data, wherein the synthetic data generation unit switches the generation method of the first synthetic data between the reference data determined to be similar to the target data by the similarity determination unit and the reference data determined not to be similar to the target data by the similarity determination unit. The information processing apparatus according to claim 1, characterized by the above.

3. For the reference data determined not to be similar to the target data by the similarity determination unit, the synthetic data generation unit generates the first synthetic data by swapping the feature amount derived from the reference data and the corresponding feature amount of the target data so that the number of feature amounts derived from the reference data is evenly distributed without bias in the entire first synthetic data finally generated, and for the reference data determined to be similar to the target data by the similarity determination unit, generates the first synthetic data in which the number of feature amounts derived from the reference data is one or less. The information processing apparatus according to claim 2, characterized by the above.

4. An information processing method executed by an information processing apparatus that presents information for detecting an environmental change in a prediction system using a machine learning model, A first step of generating first synthetic data obtained by synthesizing the target data and the reference data based on the target data to be predicted and a plurality of reference data prepared in advance, respectively; A second step of making a prediction for each of the first synthetic data; A third step of calculating the reliability of the prediction result for each of the first synthetic data, respectively; A fourth step of calculating the contribution degree of each feature amount of the target data to the reliability of the prediction result for the target data based on the reliability of the prediction result for each of the first synthetic data; A fifth step of outputting the contribution degree of each of the feature amounts to the reliability of the prediction result for the calculated target data An information processing method characterized by comprising the above steps.

5. In the first step, Determine whether there is similarity between the target data and each of the reference data, respectively, Switch the generation method of the first synthetic data between the reference data determined to be similar to the target data and the reference data determined not to be similar to the target data The information processing method according to claim 4, characterized in that.

6. In the first step, For the reference data determined not to be similar to the target data, the number of feature amounts derived from the reference data is evenly distributed without bias in the entire first synthetic data finally generated, and the feature amounts derived from the reference data and the corresponding feature amounts of the target data are exchanged to generate the first synthetic data, For the reference data determined to be similar to the target data, generate the first synthetic data in which the number of feature amounts derived from the reference data is one or less The information processing method according to claim 5, characterized in that.

Citation Information

Patent Citations

  • Estimation method and device

    JP1996249007A

  • Predicting device with reliability scale

    JP2003323601A

  • Medical treatment information analysis apparatus, method, and program

    JP2006163465A

  • Data analysis device and data analysis method

    JP2018147280A

  • Prediction basis presentation system for model and prediction basis presentation method for model

    JP2020095398A