Risk analysis support system, risk analysis support method, and risk analysis support program

The risk analysis support system uses MFCVAE to enhance risk analysis by assigning common features and keywords from a keyword-assigned dataset to an unassigned dataset, addressing the oversight of potential risks in machine learning models.

JP7798818B2Active Publication Date: 2026-01-14HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023017250
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-02-08
Publication Date
2026-01-14
Estimated Expiration
2043-02-08

AI Technical Summary

Technical Problem

Existing methods fail to extract all variations in training data that could be risk factors for machine learning models, particularly when there is no identical domain for the training data, leading to overlooked risks in risk analysis.

Method used

A risk analysis support system that utilizes Multi-Facet Clustering Variational Autoencoders (MFCVAE) to generate distributions of keyword-assigned and keyword-unassigned datasets in a latent space, selecting and assigning common features and keywords from the assigned dataset to the unassigned dataset to identify potential risks.

Benefits of technology

Reduces the omission of variations in the target dataset that can be risk factors for machine learning models, enriching the dataset with relevant keywords and features, thereby enhancing risk analysis accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798818000001
    Figure 0007798818000001
  • Figure 0007798818000002
    Figure 0007798818000002
  • Figure 0007798818000003
    Figure 0007798818000003
Patent Text Reader

Abstract

To reduce extraction omission of the variation of target datasets which become a factor of a risk of a machine learning model.SOLUTION: A risk analysis support device reads a keyword-assigned dataset which is assigned with a feature of each data and a keyword associated with the feature and a keyword-unassigned dataset which is not assigned with the feature of each data and the keyword associated with the feature. The risk analysis support device generates the distribution of each data of the keyword-assigned dataset and the keyword-unassigned dataset in a latent space, and selects the first data of the keyword-assigned dataset that is similar to the second data of the keyword-assigned dataset on the basis of the distance of the feature amount of each data. The risk analysis support device assigns the feature and the keyword associated with the feature to the second data when there is the feature of the first data that is determined to be common with the second data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention is a risk analysis support Systems and risk analysis support Methodology and Risk Analysis support Regarding the program. [Background technology]

[0002] The quality of a machine learning model depends on the training data and test data (hereinafter referred to as the "target dataset") used to learn the machine learning model. Therefore, in order to prevent overlooking risks in a risk analysis of a machine learning model, it is important to thoroughly extract variations in the target dataset that are factors contributing to the risk of the machine learning model.

[0003] As a technique for preventing risk analysis omissions, Patent Document 1 discloses a technique for reusing the results of risk analysis of a past system having a control structure similar to or identical to the system to be analyzed.

[0004] Hazard and Operability Studies (HAZOP) is a well-known method for thoroughly identifying potential hazards in new chemical processes, assessing their impact, and developing safety measures. HAZOP can be applied to the evaluation of machine learning models, and variations in the target dataset that are risk factors for the machine learning model can be extracted using keywords for each domain of the target dataset. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Publication No. 2020-017238 Summary of the Invention [Problem to be solved by the invention]

[0006] However, the technology disclosed in Patent Document 1 extracts all control patterns of the system, but is unable to extract all variations in the learning data that could be a risk factor for the machine learning model.

[0007] Furthermore, when applying HAZOP to the evaluation of machine learning models, if there is no existing domain that is identical to the training data or test data, keywords cannot be used and variations in the training data that may be a risk factor for the machine learning model cannot be extracted.

[0008] The present invention has been made in consideration of the above-mentioned problems, and aims to reduce the overlooking of extraction of variations in a target dataset, which is a risk factor for machine learning models. [Means for solving the problem]

[0009] In order to solve this problem, one aspect of the present invention is a risk analysis support system that supports risk analysis of machine learning models, the risk analysis support device having a memory and a control unit, the control unit reads a keyword-assigned dataset in which each piece of data is assigned a feature of the data and a keyword linked to the feature, and a keyword-unassigned dataset in which each piece of data is not assigned a feature of the data and a keyword linked to the feature, generates a distribution of each piece of data in the keyword-assigned dataset and the keyword-unassigned dataset in a latent space based on features of a predetermined dataset, selects first data from the keyword-assigned dataset that is similar to second data in the keyword-unassigned dataset based on the distance between the feature of each piece of data in the keyword-assigned dataset and the feature of each piece of data in the keyword-unassigned dataset, and if there is a feature of the first data that is determined to be common to the second data, assigns the feature and the keyword linked to the feature to the second data. [Effects of the Invention]

[0010] According to the present invention, it is possible to reduce the omission of variations in a target dataset that can be a risk factor for a machine learning model. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a risk analysis system according to a first embodiment. [Figure 2A] FIG. 2 is a diagram showing an example of a keyword-assigned dataset, features of the keyword-assigned dataset, and keywords of the keyword-assigned dataset according to the first embodiment. [Figure 2B] FIG. 2 is a diagram showing an example of a keyword-unassigned data set according to the first embodiment. [Figure 3] 4 is a flowchart showing an example of main processing of the risk analysis system according to the first embodiment. [Figure 4] 6 is a flowchart showing an example of a distribution generation process according to the first embodiment. [Figure 5] 10 is a flowchart showing an example of a keyword-assigned data set selection process according to the first embodiment. [Figure 6] 10 is a flowchart showing an example of feature selection processing according to the first embodiment. [Figure 7] 10 is a flowchart showing an example of a keyword selection process according to the first embodiment. [Figure 8] FIG. 2 is a diagram showing an example of a GUI according to the first embodiment. [Figure 9] 10 is a flowchart showing an example of a keyword-assigned data set selection process according to the second embodiment. [Figure 10] 10 is a flowchart showing an example of main processing of the risk analysis system according to the third embodiment. [Figure 11] 10 is a flowchart showing an example of a distribution generation process according to the fourth embodiment. [Figure 12] FIG. 13 is a diagram showing an example of a GUI for displaying the distribution of a keyword-assigned dataset according to the fourth embodiment. [Figure 13] FIG. 1 is a diagram showing an example of the hardware configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments according to the disclosure of the present application will be described with reference to the drawings. The embodiments, including the drawings, are examples for explaining the present application. In the embodiments, appropriate omissions and simplifications have been made for clarity of explanation. Unless otherwise specified, components of the embodiments may be singular or plural. Furthermore, a combination of one embodiment with another embodiment is also included in the embodiments according to the present application.

[0013] Identical or similar components are given the same reference numerals, and descriptions of the previously described components in later embodiments may be omitted, or descriptions may be focused on differences. Furthermore, when there are multiple identical or similar components, they may be described with the same reference numerals but with different subscripts. Furthermore, when it is not necessary to distinguish between these multiple components, the subscripts may be omitted.

[0014] In the following embodiment, keywords that serve as hints for recalling training data previously used to build a machine learning model (hereinafter referred to as "keyword-assigned dataset") are assigned to the training data. The keywords depend on the domain (e.g., field, area, category, etc.) to which the training data belongs, and represent, for example, the attributes of the target represented by each datum in the training data.

[0015] The training data for building a new machine learning model is called a “keyword-unassigned dataset.” Consider the case where keywords are assigned to a keyword-unassigned dataset.

[0016] If a keyword-assigned dataset exists in the same domain as the keyword-unassigned dataset, the keywords assigned to the keyword-assigned dataset can be reused as keywords to be assigned to the keyword-unassigned dataset.

[0017] However, if there is no keyword-assigned dataset for the same domain as the unkeyworded dataset, the keywords assigned to the keyword-assigned dataset cannot be reused. This means that variations in the unkeyworded dataset may be missed, and the potential risks of new machine learning models using unkeyworded datasets may not be identified in advance.

[0018] In this embodiment, even if there is no keyword-assigned data set in the same domain as the keyword-unassigned data set, the keywords assigned to the keyword-assigned data set can be reused.

[0019] [Embodiment 1] (Configuration of risk analysis system 1S according to embodiment 1) Fig. 1 is a diagram showing an example of the configuration of a risk analysis system 1S according to embodiment 1. Fig. 2A is a diagram showing an example of a keyword-assigned dataset 1D, features 1F of the keyword-assigned dataset 1D, and keywords 1K of the keyword-assigned dataset 1D according to embodiment 1. Fig. 2B is a diagram showing an example of a keyword-unassigned dataset 2D according to embodiment 1.

[0020] The risk analysis system 1S has a keyword-assigned data set input unit 11a, a keyword-unassigned data set input unit 11b, a feature input unit 11c, and a keyword input unit 11d.

[0021] The risk analysis system 1S also includes a keyword-assigned data set holding unit 12a, a keyword-unassigned data set holding unit 12b, a feature holding unit 12c, and a keyword holding unit 12d.

[0022] The risk analysis system 1S also includes a distribution generation unit 13a, a data set selection unit 13b, a feature selection unit 13c, and a keyword selection unit 13d. The processing functions of the distribution generation unit 13a, the data set selection unit 13b, the feature selection unit 13c, and the keyword selection unit 13d will be described later.

[0023] The keyword-assigned dataset input unit 11a accepts input of the keyword-assigned dataset 1D to the risk analysis system 1S. In this embodiment, as shown in Fig. 2A , the keyword-assigned dataset 1D consists of traffic light images 1D1, ..., 1D10, which are road images #1, ..., #10, pedestrian crossing images 1D11, ..., 1D20, which are road images #11, ..., #20, and sign images 1D21, ..., 1D30, which are road images #21, ..., #30.

[0024] Furthermore, the keyword-assigned data set 1D is assigned a feature 1F and a keyword 1K linked to the feature 1F. The feature 1F includes a feature 1Fx of each data 1Dx (x = 1 to 30). Each feature 1Fx (x = 1 to 30) is associated with a keyword 1Kx that evokes each 1Fx. That is, each data 1Dx (x = 1 to 30) is assigned a feature 1Fx of each data 1Dx and a keyword 1Kx linked to each 1Fx.

[0025] The keyword assignment data set 1D input by the keyword assignment data set input unit 11a is stored in the keyword assignment data set holding unit 12a.

[0026] The keyword-unassigned dataset input unit 11b accepts input of a keyword-unassigned dataset 2D to the risk analysis system 1S. In this embodiment, the keyword-unassigned dataset 2D includes, as shown in FIG. 2B , horse images 2D1, ..., 2D10, which are animal images #1, ..., #10, zebra images 2D11, ..., 2D20, which are animal images #11, ..., #20, and cat images 2D21, ..., 2D30, which are animal images #21, ..., #30.

[0027] The keyword-unassigned dataset 2D input by the keyword-unassigned dataset input unit 11b is stored in the keyword-unassigned dataset holding unit 12b.

[0028] The feature input unit 11c accepts input of the features 1F of the keyword-assigned dataset 1D to the risk analysis system 1S. The features 1F of the keyword-assigned dataset 1D input by the feature input unit 11c are stored in the keyword-unassigned dataset holding unit 12b.

[0029] The keyword input unit 11d accepts input of keywords 1K linked to features 1F of the keyword-assigned data set 1D to the risk analysis system 1S. The keywords 1K linked to features 1F of the keyword-assigned data set 1D input by the keyword input unit 11d are stored in the keyword holding unit 12d.

[0030] (Main processing according to the first embodiment) FIG. 3 is a flowchart showing an example of the main processing of the risk analysis system 1S according to the first embodiment.

[0031] First, in step S11, the keyword-assigned dataset input unit 11a reads the input of the keyword-assigned dataset 1D and stores it in the keyword-assigned dataset holding unit 12a. Next, in step S12, the keyword-unassigned dataset input unit 11b reads the input of the keyword-unassigned dataset 2D and stores it in the keyword-unassigned dataset holding unit 12b.

[0032] Next, in step S13, the feature input unit 11c reads the feature 1F of the keyword-assigned data set 1D and stores it in the feature storage unit 12c. Next, in step S14, the keyword input unit 11d reads the keyword 1K linked to the feature 1F of the keyword-assigned data set 1D and stores it in the keyword storage unit 12d.

[0033] Next, in step S15, the distribution generating unit 13a executes a distribution generating process, the details of which will be described later with reference to FIG.

[0034] Next, in step S16, the data set selection unit 13b executes a keyword-assigned data set selection process, the details of which will be described later with reference to FIG.

[0035] Next, in step S17, the feature selection unit 13c executes a feature selection process, the details of which will be described later with reference to FIG.

[0036] Next, in step S18, the keyword selection unit 13d executes a keyword selection process, the details of which will be described later with reference to FIG.

[0037] (Distribution Generation Process According to the First Embodiment) FIG. 4 is a flowchart illustrating an example of the distribution generation process according to the first embodiment.

[0038] First, in step S15a, the distribution generation unit 13a generates a machine learning model of MFCVAE (Multi-Facet Clustering Variational Autoencoders) (hereinafter referred to as the "MFCVAE model") using the keyword-assigned dataset 1D stored in the keyword-assigned dataset holding unit 12a.

[0039] MFCVAE is an extended variational auto-encoder (VAE) that can output latent variables (features expressed as vectors) from multiple perspectives. A variational auto-encoder is a generative model that uses a neural network and assumes a probability distribution as the space of latent variables. The "perspective" in MFCVAE refers to the type of latent variable output by MFCVAE, and is the basis for identification by the machine learning model. In the case of character data, "perspective" would be "character type" or "character shape (thickness, angle, etc.)". "Features" are assigned to the dataset by humans based on the identification results of the machine learning model.

[0040] Variational autoencoders are described in reference 1: Diederik P Kingma, Max Welling, “Auto-Encoding Variational Bayes,” May 2014. [Retrieved January 27, 2023], Internet<URL:https: / / arxiv.org / abs / 1312.6114> MFCVAE is disclosed in Reference 2, "Fabian Falck et al., "Multi-Facet Clustering Variational Autoencoders, Oct. 2021. [Retrieved January 27, 2023], Internet<URL:https: / / arxiv.org / abs / 2106.05241> " is disclosed in

[0041] Next, in step S15b, the distribution generation unit 13a generates distributions of the keyword-assigned dataset 1D and the keyword-unassigned dataset 2D in a latent space based on the MFCVAE model trained in step S15a. The latent space based on the MFCVAE model is a vector space based on multiple latent variables of the MFCVAE.

[0042] Note that the machine learning model generated in step S15a is not limited to the MFCVAE model based on the keyword-assigned dataset 1D. Furthermore, the latent space in which the distribution is generated in step S15b is not limited to the latent space of the MFCVAE. That is, in this distribution generation process, the distributions of the keyword-assigned dataset 1D and the keyword-unassigned dataset 2D may be generated in any latent space based on a machine learning model having any encoder (including VAE) trained using any data. The arbitrary data may be the keyword-unassigned dataset 2D or a dataset that is a union of the keyword-assigned dataset 1D and the keyword-unassigned dataset 2D.

[0043] (Keyword-assigned data set selection process according to the first embodiment) FIG. 5 is a flowchart illustrating an example of a keyword-assigned data set selection process according to the first embodiment.

[0044] First, in step S16a, the dataset selection unit 13b sets 1 to the indexes i and j.

[0045] Next, in step S16b, the dataset selection unit 13b determines whether the index i is less than or equal to the number M of the keyword - added datasets 1D. When i ≤ M (YES in step S16b), the dataset selection unit 13b transfers the process to step S16c; when i > M (NO in step S16b), the dataset selection unit 13b transfers the process to step S16i.

[0046] In step S16c, the dataset selection unit 13b determines whether the index j is less than or equal to the number N of the keyword - non - added datasets 2D. When j ≤ N (YES in step S16c), the dataset selection unit 13b transfers the process to step S16d; when j > N (NO in step S16c), the dataset selection unit 13b transfers the process to step S16h.

[0047] In step S16d, the dataset selection unit 13b calculates the distance D between the i - th data of the keyword - added dataset 1D and the j - th data of the keyword - non - added dataset 2D. The distance D is, for example, the Euclidean distance of the feature amounts of the two data, but is not limited to this.

[0048] Next, in step S16e, the dataset selection unit 13b determines whether the distance D calculated in step S16d is less than the minimum value Min_D. When D < Min_D (YES in step S16e), the dataset selection unit 13b transfers the process to step S16f; when D ≥ Min_D (NO in step S16e), the dataset selection unit 13b transfers the process to step S16g.

[0049] In step S16f, the dataset selection unit 13b substitutes the distance D calculated in step S16d into the minimum value Min_D and holds the indexes i and j at that time. Next, in step S16g, the dataset selection unit 13b adds 1 to j and returns the process to step S16c.

[0050] In step S16h, the data set selection unit 13b adds 1 to i, sets j=1, and returns the process to step S16b.

[0051] In step S16i, the data set selection unit 13b outputs the minimum value Min_D and the values ​​of the indexes i and j held in step S16f.

[0052] The combination of data in the keyword-assigned data set 1D and data in the keyword-unassigned data set 2D output by this keyword-assigned data set selection process is not limited to one with the smallest distance, but may be a plurality of combinations with distances equal to or less than a predetermined value.

[0053] (Feature Selection Process According to the First Embodiment) FIG. 6 is a flowchart illustrating an example of the input feature selection process according to the first embodiment.

[0054] First, in step S17a, the feature selection unit 13c determines whether or not a selection of feature 1F of the keyword-assigned data set 1D has been input. If a selection of feature 1F of the keyword-assigned data set 1D has been input (step S17a), the feature selection unit 13c proceeds to step S17b, and if no input has been input (step S17a NO), the feature selection unit 13c proceeds to step S17c.

[0055] In step S17b, the feature selection unit 13c accepts the selection of the feature 1F of the keyword-assigned dataset 1D inputted by the user of the risk analysis system 1S via the feature input unit 11c. The feature 1F selected here is the feature 1F of the keyword-assigned dataset 1D that the user judges to be common between the data of the keyword-assigned dataset 1D and the data of the keyword-unassigned dataset 2D that are determined to be closest in the keyword-assigned dataset selection process (FIG. 5).

[0056] The feature 1F selected in step S17b is not limited to a feature common to the data of the keyword-assigned dataset 1D and the data of the keyword-unassigned dataset 2D that are determined to be closest in the keyword-assigned dataset selection process, but may also be a similar feature whose Euclidean distance, etc., is less than a predetermined value.

[0057] Meanwhile, in step S17c, the feature selection unit 13c receives, via the feature input unit 11c, an input of a new feature F_new of the keyword-unassigned data set 2D that is determined by the user of the risk analysis system 1S to be common to the keyword-assigned data set 1D.

[0058] In step S17d, the feature selection unit 13c saves the feature 1F of the keyword-assigned dataset 1D selected in step S17b or the new feature F_new input in step S17c in a list L1 stored in the feature holding unit 12c. Next, in step S17e, the feature selection unit 13c outputs the list L1.

[0059] (Keyword selection process according to the first embodiment) FIG. 7 is a flowchart illustrating an example of a keyword selection process according to the first embodiment.

[0060] First, in step S18a, the keyword selection unit 13d determines whether the feature read from the list L1 is feature 1F of the keyword-assigned data set 1D to which the keyword is linked. If the feature read from the list L1 is feature 1F of the keyword-assigned data set 1D (step S18a YES), the keyword selection unit 13d proceeds to step S18b, and if the feature is a new feature F_new (step S18a NO), the keyword selection unit 13d proceeds to step S18c.

[0061] In step S18b, the keyword selection unit 13d accepts, via the keyword input unit 11d, a selection by the user of the risk analysis system 1S of a keyword 1K linked to the feature 1F of the keyword-assigned data set 1D.

[0062] Meanwhile, in step S18c, the keyword selection unit 13d accepts, via the keyword input unit 11d, an input of a keyword to be linked to a new feature F_new, selected by the user of the risk analysis system 1S from the existing keywords 1K linked to the feature 1F of the keyword-assigned dataset 1D. In step S18c, if the keyword-assigned dataset 1D does not have a feature 1F common to the data in the keyword-unassigned dataset 2D, the new feature F_new is added to the feature 2F of the keyword-unassigned dataset 2D.

[0063] In step S18d, the keyword 1K selected in step S18b or S18c is saved in a list L2 stored in the keyword holding unit 12d. Next, in step S18e, the keyword selecting unit 13d outputs the list L2.

[0064] (Keyword selection GUI 50 according to the first embodiment) 8 is a diagram showing an example of the GUI 50 according to embodiment 1. The GUI 50 is displayed on the display screen of the output device 1006 (FIG. 13) described below.

[0065] 8, images 1D1 to 1D30 of a keyword-assigned dataset 1D, features 1F of the images 1D1 to 1D30, and keywords 1K linked to the features 1F are displayed in a display area 51. Furthermore, images 2D1 to 2D30 of a keyword-unassigned dataset 2D are displayed in a display area 52.

[0066] Furthermore, the distribution in the latent space of the MFCVAE of each image 1D1 to 1D30 of the keyword-assigned dataset 1D and each image 2D1 to 2D30 of the keyword-unassigned dataset 2D, which was generated by the distribution generation process (FIGS. 3 and 4), is displayed in the display area 53. For example, an identification display 53D is displayed in the display area 53, indicating that image 1D11 of the keyword-assigned dataset 1D and image 2D15 of the keyword-unassigned dataset 2D are the combination of data determined to be the shortest in the keyword-assigned dataset selection process (FIGS. 3 and 5).

[0067] A user of the risk analysis system 1S references image 1D11 (road image #11 (crosswalk)) and image 2D15 (animal image #15 (zebra)) based on identification display 53D. The user then selects "stripes" as a feature 1F of image 1D11 that is common to these images.

[0068] The risk analysis system 1S receives input of the selected feature 1F of "stripes" (step S17a of the feature selection process (FIG. 6)). Then, "stripes" is selected by the risk analysis system 1S as feature 2F of the image 2D15, as indicated by arrow 54.

[0069] Furthermore, the feature 1F of "stripes" is linked to keywords 1K such as "uneven stripes," "white and brown," etc., as shown by the solid line 55. Therefore, the keyword 1K linked to the feature 1F of "stripes" is selected by the risk analysis system 1S as keyword 2K linked to "stripes," which is included in feature 2F, as shown by the arrow 56.

[0070] In this way, the keywords associated with datasets in other domains can be reused when identifying variations in the data of the target dataset. This allows for the enrichment of keywords that evoke risks in machine learning models, the expansion of variations in the target dataset, and the reduction of missed risks in the analysis of machine learning models built using the target dataset.

[0071] [Embodiment 2] In the first embodiment, a keyword 1K linked to a feature 1F common to the keyword-unassigned dataset 2D is selected as a keyword 2K to be linked to the feature 2F of the keyword-unassigned dataset 2D. In the second embodiment, the keyword-assigned dataset 1D linked to the keyword 1K selected as the keyword 2K to be linked to the feature 2F of the keyword-unassigned dataset 2D is limited to data whose feature 1F has not yet been selected as being common to the keyword-unassigned dataset 2D and whose classification result using a machine learning model is "inference failed."

[0072] This makes it possible to accurately identify data features and keywords that pose a risk of resulting in an "inference failure" classification result.

[0073] (Keyword-assigned data set selection process according to the second embodiment) 9 is a flowchart showing an example of a keyword-assigned data set selection process according to embodiment 2. The keyword-assigned data set selection process according to embodiment 2 is executed in place of the keyword-assigned data set selection process (FIG. 5) in step S16 (FIG. 3) of the main process of the risk analysis system.

[0074] First, in step S21, the dataset selection unit 13b assigns a label indicating whether the data has been selected or not and the identification result of the machine learning model to each data item in the keyword-assigned dataset 1D, whose distribution has been generated in the latent space in step S15 (FIG. 3). The machine learning model is the model trained in step S15a. The label indicating selected / not selected indicates "selected / not selected" as the feature 1F of the corresponding data item is common to the keyword-unassigned dataset 2D, and is changed from "not selected" to "selected" when the feature 1F of the corresponding data item is selected.

[0075] Next, in step S22, the data set selection unit 13b selects data whose label is "unselected" and whose classification result is "inference failed" from the keyword-assigned data set 1D.

[0076] Next, in step S23, the data set selection unit 13b executes the same processing as the keyword-assigned data set selection processing (FIG. 5) according to the first embodiment on the data set selected in step S22.

[0077] [Embodiment 3] In the first embodiment, the features 1F and keywords 1K of the keyword-assigned dataset 1D are applied only once to the keyword-unassigned dataset 2D. In the third embodiment, a process of applying the features 1F and keywords 1K of the keyword-assigned dataset 1D to the keyword-unassigned dataset 2D multiple times will be described.

[0078] (Main Processing of the Risk Analysis System According to the Third Embodiment) FIG. 10 is a flowchart illustrating an example of a main process of the risk analysis system according to the third embodiment.

[0079] The main processing of the risk analysis system of embodiment 3 differs from the main processing of the risk analysis system (Figure 3) in that step S14 is followed by step S14a, and then step S14b, after which steps S15 to S18 are repeatedly executed a predetermined number of times.

[0080] In step S14a, the distribution generation unit 13a initializes a counter variable cnt to 1. Next, in step S14b, the distribution generation unit 13a determines whether the value of the counter variable cnt is equal to or less than a predetermined number of times N_loop. If the value of the counter variable cnt is equal to or less than the predetermined number of times N_loop (YES in step S14b), the distribution generation unit 13a proceeds to step S15, and if the value of the counter variable cnt is greater than the predetermined number of times N_loop (NO in step S14b), the distribution generation unit 13a ends this main process.

[0081] However, in step S15 of the main processing of the risk analysis system according to the third embodiment, the distributions of the keyword-assigned dataset 1D and the keyword-unassigned dataset 2D are generated while changing the type of latent space for each execution of the loop indicated by the counter variable cnt. The type of latent space is changed by changing the type of training dataset and / or encoder.

[0082] In the feature selection process (FIG. 6) in step S17 of the main process of the risk analysis system according to the third embodiment, a list L1 for each loop execution indicated by the counter variable cnt is used. In the keyword selection process (FIG. 7) in step S18, a list L2 for each loop execution indicated by the counter variable cnt is used. In this way, the lists L1 and L2 for each loop execution are used to save and output different features and keywords for each loop execution.

[0083] In this embodiment, features 1F of the data in keyword-assigned dataset 1D, which are assigned to keyword-unassigned dataset 2D, and keywords 1K linked to features 1F can be extracted in multiple patterns using multiple latent spaces with different perspectives. This allows for an enrichment of keywords that evoke the risks of the machine learning model, expanding the variation in the target dataset, and reducing risk analysis omissions in the machine learning model constructed using the target dataset.

[0084] [Embodiment 4] In the fourth embodiment, the distribution of the keyword-assigned dataset 1D generated in the latent space in step S15 (FIG. 3) is output and presented.

[0085] (Distribution Generation Process According to the Fourth Embodiment) 11 is a flowchart showing an example of a distribution generation process according to the fourth embodiment. The distribution generation process according to the fourth embodiment differs from the distribution generation process according to the first embodiment (FIG. 4) in that step S15c is executed after step S15b. In step S15c, the distribution generation unit 13a outputs and presents the distribution of the keyword-assigned data set 1D generated in step S15b via a GUI 50D. The GUI 50D is displayed on the display screen of an output device 1006 (FIG. 13), which will be described later.

[0086] 12 is a diagram showing an example of a GUI 50D that displays the distribution of a keyword-assigned dataset 1D according to the fourth embodiment. The GUI 50D displays a mixture of the keyword-assigned dataset 1D and the keyword-unassigned dataset 2D in a latent space based on the MFCVAE model. The GUI 50D may also display keywords from domains different from the target dataset as the keywords 2K output by the risk analysis system 1S, in addition to keywords that have been used in the past in the same domain as the target dataset.

[0087] [Embodiment 5] In the first embodiment, the dataset to which the keyword 1K linked to the feature 1F of the keyword-assigned dataset 1D is applied is the keyword-unassigned dataset 2D. However, the present invention is not limited to this, and the dataset to which the keyword 1K linked to the feature 1F of a certain keyword-assigned dataset 1D is applied may be another keyword-assigned dataset 1D.

[0088] In embodiment 5, the main processing (Figure 3) of the risk analysis system related to embodiment 1 is executed by replacing the "keyword-assigned dataset" with the "first keyword-assigned dataset" and the "non-keyword-assigned dataset" with the "second keyword-assigned dataset."

[0089] As a result, even in the keyword-assigned dataset 1D, the features 1F and keywords 1K can be enriched using the features and keywords of other keyword-assigned datasets, which further reduces the omission of extraction of variations in the keyword-assigned dataset 1D.

[0090] (Computer 1000 hardware) 16 is a hardware diagram showing the configuration of a computer 1000. For example, the risk analysis system 1S, or each system in which the components of the risk analysis system 1S, namely the keyword-assigned dataset holding unit 12a, the keyword-unassigned dataset holding unit 112b, the feature holding unit 12c, the keyword holding unit 12d, the distribution generation unit 13a, the dataset selection unit 13b, the feature selection unit 13c, and the keyword selection unit 13d, are appropriately distributed, is realized by the computer 1000.

[0091] The computer 1000 comprises a processor 1001 including a CPU, a main memory device 1002, an auxiliary memory device 1003, a network interface 1004, an input device 1005, and an output device 1006, all of which are interconnected via an internal communication line 1009 such as a bus.

[0092] The processor 1001 controls the overall operation of the computer 1000. The main memory device 1002 is composed of, for example, a volatile semiconductor memory, and is used as a work memory for the processor 1001. The auxiliary memory device 1003 is an example of a non-transitory storage medium, and is composed of a large-capacity non-volatile storage device such as a hard disk drive, an SSD (Solid State Drive), or a flash memory, and is used to store various programs and data for a long period of time. The auxiliary memory device 1003 comprises a keyword-assigned dataset storage unit 12a, a keyword-unassigned dataset storage unit 12b, a feature storage unit 12c, and a keyword storage unit 12d.

[0093] An executable program 1100 stored in the auxiliary storage device 1003 is loaded into the main storage device 1002 when the computer 1000 is started up or when necessary, and the executable program 1100 loaded into the main storage device 1002 is executed by the processor 1001. This realizes a system that executes various processes and various functional units (a distribution generation unit 13a, a data set selection unit 13b, a feature selection unit 13c, and a keyword selection unit 13d).

[0094] The executable program 1100 may be recorded on a non-transitory recording medium, read from the non-transitory recording medium by a media reading device, and loaded into the main memory device 1002. Alternatively, the executable program 1100 may be obtained from an external computer via a network and loaded into the main memory device 1002.

[0095] The network interface 1004 is an interface device for connecting the computer 1000 to each network in the system or for communicating with other computers. The network interface 1004 is configured, for example, by a NIC (Network Interface Card) for a wired LAN (Local Area Network) or a wireless LAN.

[0096] The input device 1005 is composed of a keyboard, a pointing device such as a mouse, and the like, and is used by the user to input various instructions and information to the computer 1000. The input device 1005 corresponds to the keyword-assigned dataset input unit 11a, the keyword-unassigned dataset input unit 11b, the feature input unit 11c, and the keyword input unit 11d.

[0097] The output device 1006 is composed of a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display, and an audio output device such as a speaker, and is used to present necessary information to the user when necessary. The GUI displays 50 and 50D are displayed on the display screen of the display device of the output device 1006.

[0098] The present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the spirit of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to configurations including all of the described configurations. Furthermore, part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Furthermore, the configuration of another embodiment may be added to the configuration of one embodiment. Furthermore, part of the configuration of each embodiment may be added to, deleted from, or replaced with other configurations.

[0099] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partly or entirely realized in hardware, for example, by designing them as integrated circuits, or may be realized in software, by a processor interpreting and executing a program that realizes each function.

[0100] Information such as programs, tables, and files that realize each function can be stored in storage devices such as memory, hard disks, and SSDs (Solid State Drives), or non-temporary recording media such as IC (Integrated Circuit) cards, SD cards, and DVDs (Digital Versatile Discs).

[0101] In addition, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily represent all the control lines and information lines that are necessary for implementation. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]

[0102] 1S: Risk analysis system, 1D: Dataset with keywords, 1F, 1F1, ..., 2F: Features, 1K, 1K1, ..., 2K: Keywords, 2D: Dataset without keywords, 1000: Computer, 1001: Processor, 1002: Main memory device, 1005: Input device, 1006: Output device.

Claims

1. A risk analysis support system that supports risk analysis of machine learning models, the risk analysis support system includes a memory and a control unit, The control unit A keyword-assigned data set in which each data item is assigned a feature of the data item and a keyword associated with the feature, and a keyword-unassigned data set in which each data item is not assigned a feature of the data item and a keyword associated with the feature, are read; generating a distribution of data of the keyword-assigned dataset and the keyword-unassigned dataset in a latent space based on the feature quantities of a predetermined dataset; selecting first data of the keyword-assigned data set that is similar to second data of the keyword-unassigned data set based on a distance between a feature amount of each data of the keyword-assigned data set and a feature amount of each data of the keyword-unassigned data set; If there is a feature of the first data that is determined to be common to the second data, the feature and a keyword associated with the feature are assigned to the second data. A risk analysis support system characterized by:

2. 2. The risk analysis support system according to claim 1, the predetermined data set is a data set including either one or both of the keyword-assigned data set and the keyword-unassigned data set, The latent space is the latent space of MFCVAE (Multi-Facet Clustering Variational Auto-Encoder). A risk analysis support system characterized by:

3. 2. The risk analysis support system according to claim 1, The control unit If there is no feature of the first data that is determined to be common to the second data, adding a new feature to the second data that is determined to be common to the second data; A keyword selected from the keywords associated with the features of the data in the keyword-assigned data set is associated with the new feature. A risk analysis support system characterized by:

4. 2. The risk analysis support system according to claim 1, The control unit Among the keyword-assigned datasets, the first data, for which the distance is equal to or less than a predetermined value, is selected from among data that has not yet been selected as being similar to the second data and for which the classification result of application of the machine learning model is an inference failure. A risk analysis support system characterized by:

5. 2. The risk analysis support system according to claim 1, The control unit generating a distribution of each data of the keyword-assigned data set and the keyword-unassigned data set in the latent space for each of the different latent spaces; selecting, for each of the different latent spaces, the first data that is determined to be similar to the second data based on a distance between a feature amount of each data item in the keyword-assigned dataset and a feature amount of each data item in the keyword-unassigned dataset; If there is a feature of the first data that is determined to be common to the second data for each of the different multiple latent spaces, the feature and a keyword associated with the feature are assigned to the second data. A risk analysis support system characterized by:

6. 2. The risk analysis support system according to claim 1, The control unit The distribution of each data of the generated keyword-assigned data set and the keyword-unassigned data set is displayed on a display device. A risk analysis support system characterized by:

7. 2. The risk analysis support system according to claim 1, The control unit A first keyword-assigned data set and a second keyword-assigned data set are read, in which each data is assigned a feature of the data and a keyword linked to the feature; generating a distribution of each data item included in the first keyword-assigned data set and the second keyword-assigned data set in the latent space; selecting third data of the first keyword-assigned data set that is similar to fourth data of the second keyword-assigned data set based on a distance between each feature amount between the first keyword-assigned data set and the second keyword-assigned data set; If there is a feature of the third data that is determined to be common to the fourth data, the feature and a keyword associated with the feature are assigned to the fourth data. A risk analysis support system characterized by:

8. A risk analysis support method executed by a risk analysis support system that supports risk analysis of a machine learning model, comprising: the risk analysis support system includes a memory and a control unit, The control unit A keyword-assigned data set in which each data item is assigned a feature of the data item and a keyword associated with the feature, and a keyword-unassigned data set in which each data item is not assigned a feature of the data item and a keyword associated with the feature, are read; generating a distribution of data of the keyword-assigned dataset and the keyword-unassigned dataset in a latent space based on the feature quantities of a predetermined dataset; selecting first data of the keyword-assigned data set that is similar to second data of the keyword-unassigned data set based on a distance between a feature amount of each data of the keyword-assigned data set and a feature amount of each data of the keyword-unassigned data set; If there is a feature of the first data that is determined to be common to the second data, the feature and a keyword associated with the feature are assigned to the second data. A risk analysis support method characterized by including each process.

9. 9. The risk analysis support method according to claim 8, the predetermined data set is a data set including either one or both of the keyword-assigned data set and the keyword-unassigned data set, The latent space is the latent space of MFCVAE (Multi-Facet Clustering Variational Auto-Encoder). A risk analysis support method comprising:

10. 9. The risk analysis support method according to claim 8, The control unit If there is no feature of the first data that is determined to be common to the second data, adding a new feature to the second data that is determined to be common to the second data; A keyword selected from the keywords associated with the features of the data in the keyword-assigned data set is associated with the new feature. A risk analysis support method characterized by including each process.

11. 9. The risk analysis support method according to claim 8, The control unit Among the keyword-assigned datasets, the first data, for which the distance is equal to or less than a predetermined value, is selected from among data that has not yet been selected as being similar to the second data and for which the classification result of application of the machine learning model is an inference failure. A risk analysis support method comprising:

12. 9. The risk analysis support method according to claim 8, The control unit generating a distribution of each data of the keyword-assigned data set and the keyword-unassigned data set in the latent space for each of the different latent spaces; selecting, for each of the different latent spaces, the first data that is determined to be similar to the second data based on a distance between a feature amount of each data item in the keyword-assigned dataset and a feature amount of each data item in the keyword-unassigned dataset; If there is a feature of the first data that is determined to be common to the second data for each of the different multiple latent spaces, the feature and a keyword associated with the feature are assigned to the second data. A risk analysis support method comprising:

13. 9. The risk analysis support method according to claim 8, The control unit The distribution of each data of the generated keyword-assigned data set and the keyword-unassigned data set is displayed on a display device.

2. A risk analysis support method comprising the steps of:

14. 9. The risk analysis support method according to claim 8, The control unit A first keyword-assigned data set and a second keyword-assigned data set are read, in which each data is assigned a feature of the data and a keyword linked to the feature; generating a distribution of each data item included in the first keyword-assigned data set and the second keyword-assigned data set in the latent space; selecting third data of the first keyword-assigned data set that is similar to fourth data of the second keyword-assigned data set based on a distance between each feature amount between the first keyword-assigned data set and the second keyword-assigned data set; If there is a feature of the third data that is determined to be common to the fourth data, the feature and a keyword associated with the feature are assigned to the fourth data. A risk analysis support method characterized by including each process.

15. A risk analysis support program for causing a computer having a memory and a control unit to function as a risk analysis support system that supports risk analysis of a machine learning model, The computer, A keyword-assigned data set in which each data item is assigned a feature of the data item and a keyword associated with the feature, and a keyword-unassigned data set in which each data item is not assigned a feature of the data item and a keyword associated with the feature, are read; generating a distribution of data of the keyword-assigned dataset and the keyword-unassigned dataset in a latent space based on the feature quantities of a predetermined dataset; selecting first data of the keyword-assigned data set that is similar to second data of the keyword-unassigned data set based on a distance between a feature amount of each data of the keyword-assigned data set and a feature amount of each data of the keyword-unassigned data set; If there is a feature of the first data that is determined to be common to the second data, the feature and a keyword associated with the feature are assigned to the second data. A risk analysis support program characterized by executing each process.

Citation Information

Patent Citations

  • Information standardization method for process safety analysis

    CN110162508A

  • Method and system for establishing recommendation model of change risk keywords, method and system for recommending recommendation model of change risk keywords

    CN113822055A

  • Risk analysis assistance device, risk analysis assistance method, and risk analysis assistance program

    JP2020017238A