Information Processing Program, Information Processing Apparatus, and Information Processing Method
By determining variables for questioning based on priority and user knowledge alignment, the program efficiently selects features for interpretable machine learning models, aligning with user expectations and reducing the number of questions required.
Patent Information
- Application Number
- JP2021208217
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-07-28
- Estimated Expiration
- 2041-12-22
AI Technical Summary
Existing methods for generating interpretable machine learning models often require numerous questions to align with user knowledge, as specific statistics may not match on-site knowledge, leading to inefficiencies in feature selection.
An information processing program determines variables for questioning based on their priority and estimated conformity to user knowledge, updating these estimates based on user feedback to efficiently select features that align with user expectations.
This approach allows for more efficient feature selection that matches user knowledge, reducing the number of questions needed to generate an interpretable machine learning model.
Smart Images

Figure 0007713684000002 
Figure 0007713684000003 
Figure 0007713684000004
Abstract
Description
Technical Field
[0001] The present invention relates to information processing technology.
Background Art
[0002] As AI (Artificial Intelligence) technology has become widespread, there is an increasing demand for explainable machine learning models, such as the desire for a human-interpretable basis for judgment because humans cannot blindly accept the judgment of black-box machine learning models. For this reason, white-box models such as rule lists, decision trees, and linear models are sometimes used in advance, but simply using a white-box machine learning model does not necessarily mean a machine learning model that is interpretable and explainable to humans.
[0003] Therefore, in recent years, a dialogical approach that repeats machine learning model generation and feedback to humans has been used to generate a machine learning model that humans can accept and that has high accuracy. Here, the dialogical approach, for example, selects one feature considered important in a machine learning model, asks the user whether it is really important, and repeats the question for each feature until the user is satisfied. Note that features in a machine learning model are sometimes also called explanatory variables, simply variables, etc.
[0004] Then, according to the feedback, the parameters for optimizing the machine learning model are changed and the machine learning model is updated. By repeating such processing, a machine learning model that humans can accept and that has high accuracy is generated.
[0005] However, since the number of features of the machine learning model is large, considering the possibility that the user may interrupt the dialog midway, it is desirable to obtain as many features answered as important with as few question times as possible.
[0006] Therefore, there are methods of asking questions starting from features with large values using existing statistics such as correlation, mutual information amount, and chi-square value, and methods of measuring the influence on the prediction distribution for each feature of a machine learning model and selecting features with greater influence to ask questions, and so on.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Patent Document 5
Summary of the Invention
Problems to be Solved by the Invention
[0008] However, a specific statistic does not always match the on-site knowledge that satisfies the user, and the number of questions until the user is satisfied may increase. Also, in the case of feature selection by a machine learning model, it can be said that it depends on the question index of a specific machine learning model, and the index does not always match the user's on-site knowledge, and ultimately, the number of questions may increase.
[0009] In one aspect, it aims to more efficiently perform feature selection that matches the user's on-site knowledge in order to generate an interpretable machine learning model.
Means for Solving the Problems
[0010] In one aspect, the information processing program causes a computer to execute a process of determining a variable to be the subject of a question regarding importance based on the priority of variables based on a plurality of patterns indicating the ranking of the plurality of variables and an estimated quantity indicating the possibility of conforming to a predetermined condition for each pattern, and updating the estimated quantity based on the answer result of the question targeted at the determined variable.
Advantages of the Invention
[0011] In one aspect, in order to generate an interpretable machine learning model, it is possible to more efficiently perform feature selection that matches the user's on-site knowledge.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0013] Hereinafter, examples of the information processing program, information processing apparatus, and information processing method according to the present embodiment will be described in detail with reference to the drawings. Note that the present embodiment is not limited by these examples. Also, the respective examples can be appropriately combined within a non - conflicting range.
[0014] [Functional Configuration of Information Processing Apparatus 10] Using FIG. 1, the functional configuration of the information processing apparatus 10 according to the present embodiment will be described. FIG. 1 is a block diagram showing a configuration example of the information processing apparatus 10. As shown in FIG. 1, the information processing apparatus 10 includes a communication unit 11, a storage unit 12, and a control unit 13.
[0015] The communication unit 11 is a processing unit that controls communication with other information processing apparatuses.
[0016] The storage unit 12 is a storage device that stores various data and programs executed by the control unit 13. The storage unit 12 stores a classification model 121, index data 122, probability data 123, and the like.
[0017] The classification model 121 is, for example, a machine - learning model that classifies input data into one of two values. For example, when an image of a person is input, the classification model 121 determines whether the person is wearing a uniform and outputs the result as wearing or not wearing.
[0018] The index data 122 is data regarding an index that indicates the ranking of features, that is, the priority, in a machine - learning model. FIG. 2 is a diagram showing an example of the index data 122 according to the present embodiment. In the example of FIG. 2, the index data 122 includes three indexes X, Y, and Z, and five features a - e are shown ranked for each index.
[0019] In the example of FIG. 2, in the case of the index X, the first feature considered important in the machine learning model is considered to be a, and the user is first asked whether the feature a is important. Then, in the case of the index X, the user is next asked whether the feature b, which is the second most important feature, is important, and questions are repeatedly asked about each feature in order based on the priority of each feature until the user is satisfied. In the example of FIG. 2, for the sake of convenience, the total number of features is five, namely a to e. However, of course, in reality, the machine learning model includes more features, such as hundreds or thousands.
[0020] Also, the estimator α is the weight coefficient of each index with an initial value of 1.0. The information processing apparatus 10 asks the user whether each feature is important and updates the estimator α based on the answer result to the question. Although the details will be described later, in this way, the information processing apparatus 10 makes it easier to select features from the index on the assumption that the index including important features at the top is more in line with the on-site knowledge.
[0021] Also, each index in the index data 122 is ranked and generated using existing technologies such as statistics such as correlation, mutual information amount, and chi-square value, and the prediction distribution of each feature of the machine learning model.
[0022] The probability data 123 is data regarding the probability estimated to be followed by the user for each feature. That is, the probability data 123 is data regarding the probability for each feature that is estimated to be important and for which it is estimated that the user will answer "important (Yes)" to the question of whether it is important. In the example of FIG. 3, it is shown that the probabilities for which a Yes answer is estimated to be obtained from the user are 0.66% for feature a and 0.33% for feature e. Therefore, in the probability data 123 in the example of FIG. 3, asking the user whether feature a is important is more efficient in order to obtain as many important features as possible with as few question times as possible. The probability data 123 is generated based on the index data 122 for each question. The generation method of the probability data 123 and the details of the feature selection for asking the user will be described later.
[0023] Note that the data stored in the storage unit 12 is merely an example, and the storage unit 12 can store various data other than the above data.
[0024] The control unit 13 is a processing unit that controls the entire information processing apparatus 10, and includes a determination unit 131, an output unit 132, and an update unit 133.
[0025] The determination unit 131 selects and determines a feature to be the subject of a question regarding importance based on the priority of features in the machine learning model based on the index data 122 and the estimated quantity α indicating the possibility of conforming to a predetermined condition for each index data 122.
[0026] For example, the determination unit 131 determines a feature to be the subject of a question based on the priority of each feature determined based on at least one statistic quantity among the correlation, mutual information amount, and chi-square value for each feature of the machine learning model, and the estimated quantity α. Note that the features in the machine learning model correspond to variables in the machine learning model, and the index data 122 corresponds to each of a plurality of patterns indicating the ranking of a plurality of variables.
[0027] The output unit 132 outputs a question regarding the importance of the feature determined by the determination unit 131. For example, the output unit 132 outputs the question via an output device such as a display device connected to the information processing apparatus 10. Alternatively, the output unit 132 may output the question to an information processing terminal (not shown) communicably connected via a network.
[0028] The update unit 133 acquires the answer result of the question regarding importance output by the output unit 132, and updates the estimated quantity α based on the answer result.
[0029] For example, when the answer result of the question regarding importance indicates that the feature determined by the determination unit 131 is important, the update unit 133 updates to increase the estimated value α of the index data 122 whose priority of the feature is equal to or higher than a predetermined threshold. On the other hand, when the answer result of the question regarding importance indicates that the feature determined by the determination unit 131 is not important, the update unit 133 updates to increase the estimated value α of the index data 122 whose priority of the feature is lower than the predetermined threshold. Here, the predetermined threshold is, for example, a value indicating the top two ranks.
[0030] [Functional Details] Next, with reference to FIGS. 4 to 6, the feature determination process for the question regarding importance executed by the information processing apparatus 10 will be described in more detail. FIG. 4 is a diagram showing an example of the feature selection process for the first question according to the present embodiment. In the example of FIG. 4, in order to ask the user the first question regarding importance, the process of selecting and determining the feature to be the question target based on the index data 122 will be described.
[0031] First, as a prerequisite, assume that the total number of features in the machine learning model is five, namely a to e. Also, assume that the features that the information processing apparatus 10 does not know, that is, the features that are not preset or stored in the information processing apparatus 10 but are considered important by the user are two, namely a and b. Also, although details will be described later, assume that the user parameter δ = 2 (top two ranks) used when determining the update content of the estimated value α. The user parameter δ may be determined in advance based on, for example, a predetermined ratio of the total number of features and stored in the information processing apparatus 10. Alternatively, the user parameter δ may be set to an arbitrary value by the user.
[0032] Regarding the feature determination process, as shown on the left side of FIG. 4, the information processing apparatus 10 extracts the first-ranked feature of each index in the index data 122 as the first question candidate. In the example of FIG. 4, features a, a, and e are extracted from the indexes X, Y, and Z, respectively.
[0033] Next, as shown on the right side of FIG. 4, the information processing apparatus 10 calculates the probability of each extracted feature and generates probability data 123 for determining the feature to be the target of the first question. The probability in the probability data 123 is a probability estimated to obtain an answer from the user that it is important for a question asking about importance. The probability is the sum value for each feature of the probability calculated based on the estimator α for each index.
[0034] For example, if there are three indexes X, Y, and Z, and each is the first, second, and third index, respectively, the probability of the i-th index is calculated using the following formula (1).
[0035]
Equation
[0036] In formula (1), Σ j α j is the sum value of the estimators α for all indexes. In the example of FIG. 4, Σ j α j = 1.0 + 1.0 + 1.0 = 3.0.
[0037] The information processing apparatus 10 calculates the probability of each index using formula (1) and sums them up for each feature to obtain the probability of each feature as shown on the right side of FIG. 4. For example, in the example of FIG. 4, since feature a is extracted from two of the indexes X and Y, the probability of feature a = the probability of index X + the probability of index Y = (1.0 / 3.0) + (1.0 / 3.0) = approximately 0.66. Similarly, the probability of feature e = the probability of index Z = (1.0 / 3.0) = approximately 0.33. Note that since features b to d are not extracted as candidates for the first question, the probabilities of features b to d are 0.
[0038] Then, the information processing apparatus 10 selects and determines the feature a with the highest probability as the target of the first question. Next, the information processing apparatus 10 outputs a question to the user asking whether feature a is important as the first question. In response to the question, since feature a is a feature that the user considers important, a Yes (important) answer is obtained from the user.
[0039] Then, since it is found that feature a is important, the information processing apparatus 10 updates the estimation quantity α so as to give preference to the indicators X and Y that include feature a within the top two positions indicated by the value set in the user parameter δ, which is 2 in the example of FIG. 4. More specifically, the information processing apparatus 10 updates the estimation quantity α of the indicators X and Y to increase, for example, from 1.0 to 1.15. Note that how much the estimation quantity α is increased may be based on, for example, the number of indicators to be given preference, a preset value, and the like. Also, being within a predetermined top rank indicated by the value set in the user parameter δ corresponds to the ranking of features in the machine learning model, that is, the priority being equal to or higher than a predetermined threshold value.
[0040] Further, the information processing apparatus 10 may update the estimation quantity α of the indicator Z that does not include the important feature a within the top two positions to decrease, for example, from 1.0 to 0.85. Thereby, from the indicators X and Y that include the important feature at the top, the feature to be the question target is more likely to be selected. Note that how much the estimation quantity α is decreased may also be based on, for example, a preset value and the like.
[0041] Next, the feature selection process for the second question will be described. FIG. 5 is a diagram showing an example of the feature determination process for the second question target according to the present embodiment. The example of FIG. 5 relates to the second question after the first question is asked to the user and the estimation quantity α of each indicator is updated based on the answer result, as described with reference to FIG. 4.
[0042] First, as shown on the left side of FIG. 5, the information processing apparatus 10 extracts, as candidates for the second question, the topmost feature among the features of each indicator in the indicator data 122 that has not yet been the question target. Since the first question target was determined to be feature a as described with reference to FIG. 4, in the example of FIG. 5, features b, e, and e are extracted as question candidates from the indicators X, Y, and Z, respectively.
[0043] Next, as shown on the right side of FIG. 5, the information processing apparatus 10 calculates the probability of each extracted feature. In the example of FIG. 5, since feature e is extracted from two of the indicators Y and Z, the probability of feature e = the probability of indicator Y + the probability of indicator Z = (1.15 / 3.15) + (0.85 / 3.15) = approximately 0.63. Similarly, the probability of feature b = the probability of indicator X = (1.15 / 3.15) = approximately 0.36.
[0044] Then, the information processing apparatus 10 selects and determines the feature e with the highest probability as the target of the second question. Next, as the second question, the information processing apparatus 10 outputs a question to the user asking whether feature e is important or not. In response to this question, since feature e is not a feature that the user considers important, a No (not important) answer is obtained from the user.
[0045] And since it is found that feature e is not important, the information processing apparatus 10 updates the estimator α so as to give preference to indicator X that does not include feature e within the top two. More specifically, the information processing apparatus 10 updates the estimator α of indicator X to increase, for example, from 1.15 to 2.0.
[0046] Also, the information processing apparatus 10 may update the estimators α of indicators Y and Z that include the unimportant feature e within the top two to decrease, for example, from 1.15 to 1.1 and from 0.85 to 0.8, respectively.
[0047] Next, the feature selection process for the third question will be described. FIG. 6 is a diagram showing an example of the feature determination process for the third question target according to the present embodiment. The example of FIG. 6 relates to the third question after the second question is asked to the user and the estimator α of each indicator is updated based on the answer result, as described with reference to FIG. 5.
[0048] First, as shown on the left side of FIG. 6, the information processing apparatus 10 extracts, as the candidates for the third question, the highest-level features among the features of each index in the index data 122 that have not yet been the subject of a question. Since the first and second questions were determined to be features a and e as described with reference to FIGS. 4 and 5, in the example of FIG. 6, features b, d, and d are extracted as question candidates from indexes X, Y, and Z, respectively.
[0049] Next, as shown on the right side of FIG. 6, the information processing apparatus 10 calculates the probability of each extracted feature. In the example of FIG. 6, since feature d is extracted from two of the indexes Y and Z, the probability of feature d = the probability of index Y + the probability of index Z = (1.1 / 3.9) + (0.8 / 3.9) = approximately 0.48. Similarly, the probability of feature b = the probability of index X = (2.0 / 3.9) = approximately 0.51.
[0050] Then, the information processing apparatus 10 selects and determines the feature b with the highest probability as the third question target. Next, as the third question, the information processing apparatus 10 outputs to the user a question as to whether feature b is important. In response to this question, since feature b is a feature that the user considers important, a Yes (important) answer is obtained from the user.
[0051] Then, since it is found that feature b is important, the information processing apparatus 10 increases and updates the estimation quantity α so as to give priority to index X that includes feature b within the top two positions. Further, the information processing apparatus 10 may update the estimation quantity α of indexes Y and Z that do not include the important feature b within the top two positions so as to decrease.
[0052] In the examples of FIGS. 4 to 6, the information processing apparatus 10 was able to recognize all the features that the user considered important in three questions, but the information processing apparatus 10 does not know the total number of features that the user considers important. Therefore, after the third question, the information processing apparatus 10 attempts to ask further questions, such as the fourth and subsequent questions, until there are no more features to be questioned. Therefore, for example, the information processing apparatus 10 may confirm with the user whether the user is satisfied via a user interface or the like, and control to continue asking questions until an answer indicating satisfaction is obtained.
[0053] [Processing flow] Next, with reference to FIG. 7, the flow of the estimated value update process executed by the information processing apparatus 10 will be described. FIG. 7 is a flowchart showing an example of the flow of the estimated value update process according to the present embodiment. The estimated value update process shown in FIG. 7 is started at an arbitrary timing, such as in response to a request from the user.
[0054] First, as shown in FIG. 7, the information processing apparatus 10 acquires the index data 122 and the user parameter δ (step S101). The index data 122 ranks each feature of the target machine learning model by an existing technique using a statistic or a machine learning model, and is stored in advance in the information processing apparatus 10. The user parameter δ is determined based on, for example, the total number of features of the target machine learning model, or is set by the user, and is stored in advance in the information processing apparatus 10. In this way, the index data 122 and the user parameter δ may be generated and stored in advance, or may be generated and acquired in step S101.
[0055] Next, the information processing apparatus 10 initializes the estimated value α in the index data 122 (step S102). More specifically, in the initial stage, since the information processing apparatus 10 does not know which features the user considers important and does not know which index in the index data 122 matches the user's on-site knowledge, the estimated value α of each index is uniformly set to 1.0. Note that the initialization of the estimated value α may be performed in conjunction with the generation of the index data 122.
[0056] Next, as described with reference to FIGS. 4 to 6, the information processing apparatus 10 calculates the probability of each feature based on the index data 122, selects the feature with the highest probability from the features that have not yet been the subject of a question as the subject of the question, and asks the user (step S103).
[0057] If the answer from the user to the question in step S103 is Yes (important) (step S104: Yes), the information processing apparatus 10 updates the estimation amount α so as to give priority to the index including the subject of the question within the upper δ (step S105).
[0058] On the other hand, if the answer from the user is No (not important) (step S104: No), the information processing apparatus 10 updates the estimation amount α so as to give priority to the index not included within the upper δ (step S106).
[0059] Next, the information processing apparatus 10 confirms with the user whether the user is satisfied via a user interface or the like. If the answer is that the user is satisfied (step S107: Yes), the estimation amount update process shown in FIG. 7 ends.
[0060] On the other hand, if the answer is that the user is not satisfied (step S107: No), the process returns to step S103, and questions asking about the importance of features that have not yet been the subject of a question are repeated until the user is satisfied (steps S103 to S107).
[0061] [Effect] As described above, the information processing apparatus 10 determines the variable to be the subject of the question asking about the importance based on the priority of the variables based on a plurality of patterns indicating the ranking of the plurality of variables and the estimation amount indicating the possibility of conforming to the predetermined conditions for each pattern, and updates the estimation amount based on the answer result of the question targeting the determined variable.
[0062] In this way, based on an index and an estimated value indicating the possibility of matching the user's on-site knowledge, the information processing apparatus 10 selects and determines a feature to be a question target regarding importance, and updates the estimated value based on the answer result. Thereby, the information processing apparatus 10 can more efficiently perform feature selection that matches the user's on-site knowledge in order to generate an explainable machine learning model.
[0063] In addition, the process of updating the estimated value includes a process of updating so as to increase the estimated value of a pattern in which the priority of the determined variable is equal to or higher than a predetermined threshold when the answer result indicates that the determined variable is important.
[0064] Thereby, the information processing apparatus 10 can more easily select features that match the user's on-site knowledge.
[0065] In addition, the process of updating the estimated value includes a process of updating so as to increase the estimated value of a pattern in which the priority of the determined variable is less than a predetermined threshold when the answer result indicates that the determined variable is not important.
[0066] Thereby, the information processing apparatus 10 can more easily select features that match the user's on-site knowledge.
[0067] In addition, the process of determining a variable includes a process of determining a variable based on a priority indicating a ranking determined based on at least one of a correlation, a mutual information amount, and a chi-square value with respect to the variable, and the estimated value.
[0068] Thereby, the information processing apparatus 10 can more easily select features that match the user's on-site knowledge.
[0069] [System] The processing procedures, control procedures, specific names, information including various data and parameters shown in the above documents and drawings may be arbitrarily changed unless otherwise specified. Also, the specific examples, distributions, numerical values, etc. described in the embodiments are merely examples and may be arbitrarily changed.
[0070] Also, the specific forms of distribution and integration of the components of the information processing apparatus 10 are not limited to those shown in the drawings. For example, the determination unit 131 of the information processing apparatus 10 may be distributed to a plurality of processing units, or the determination unit 131 and the output unit 132 of the information processing apparatus 10 may be integrated into one processing unit. That is, all or part of the components may be functionally or physically distributed and integrated in any unit according to various loads, usage situations, and the like. Further, each processing function of each device may be realized in whole or in any part thereof by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or may be realized as hardware by wired logic.
[0071] FIG. 8 is a diagram showing a hardware configuration example of the information processing apparatus 10 according to the present embodiment. As shown in FIG. 8, the information processing apparatus 10 includes a communication interface 10a, an HDD (Hard Disk Drive) 10b, a memory 10c, and a processor 10d. Also, each part shown in FIG. 6 is mutually connected by a bus or the like.
[0072] The communication interface 10a is a network interface card or the like and communicates with other information processing apparatuses. The HDD 10b stores programs and data for operating the functions shown in FIG. 1 and the like.
[0073] The processor 10d is a CPU, an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), or the like. Also, the processor 10d may be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The processor 10d is a hardware circuit that executes a process of realizing each function described in FIG. 1 and the like by reading out a program for executing the same process as each processing unit shown in FIG. 1 and the like from the HDD 10b or the like and expanding it in the memory 10c.
[0074] In addition, the information processing apparatus 10 can also realize the same functions as those in the above embodiments by reading the above program from a recording medium by a medium reading apparatus and executing the read program. Note that the program in this other embodiment is not limited to being executed by the information processing apparatus 10. For example, the above embodiments may be similarly applied when other information processing apparatuses execute the program, or when these cooperate to execute the program.
[0075] This program may be distributed via a network such as the Internet. Further, this program may be recorded on a recording medium readable by an information processing apparatus such as a hard disk, a flexible disk (FD), a CD-ROM, a MO (Magneto-Optical disk), a DVD (Digital Versatile Disc), and may be executed by being read from the recording medium by the information processing apparatus.
[0076] Regarding the embodiments including the above embodiments, the following additional remarks are further disclosed.
[0077] (Supplementary Note 1) Based on the priorities of the variables based on a plurality of patterns indicating the ranking of a plurality of variables and an estimated value indicating the possibility of conforming to a predetermined condition for each pattern, determine the variable that is the subject of a question regarding importance, Update the estimated value based on the answer result of the question targeted at the determined variable An information processing program for causing a computer to execute the process. (Supplementary Note 2) The process of updating the estimated value is When the answer result indicates that the determined variable is important, update so as to increase the estimated value of the pattern in which the priority of the determined variable is equal to or higher than a predetermined threshold The information processing program according to Supplementary Note 1, characterized by including the process. (Supplementary Note 3) The process of updating the estimated value is When the determined variable indicated by the answer result is not important, update so as to increase the estimated value of the pattern in which the priority of the determined variable is not more than the predetermined threshold An information processing program according to supplementary note 2, characterized by including the processing (Supplementary note 4) The process of determining the variable is Based on the priority based on the pattern indicating the ranking determined based on at least one statistic of correlation, mutual information amount, and chi-square value for the variable, and the estimated value, determine the variable An information processing program according to any one of supplementary notes 1 to 3, characterized by including the processing (Supplementary note 5) Based on the priority of the variable based on a plurality of patterns indicating the ranking of a plurality of variables and the estimated value indicating the possibility of conforming to a predetermined condition for each pattern, determine the variable to be the subject of a question asking about importance, Update the estimated value based on the answer result of the question targeted at the determined variable An information processing apparatus including a control unit that executes the processing (Supplementary note 6) The process of updating the estimated value is When the answer result indicates that the determined variable is important, update so as to increase the estimated value of the pattern in which the priority of the determined variable is equal to or more than the predetermined threshold An information processing apparatus according to supplementary note 5, characterized by including the processing (Supplementary note 7) The process of updating the estimated value is When the answer result indicates that the determined variable is not important, update so as to increase the estimated value of the pattern in which the priority of the determined variable is not more than the predetermined threshold An information processing apparatus according to supplementary note 6, characterized by including the processing (Supplementary note 8) The process of determining the variable is Based on the priority based on the pattern indicating the ranking determined based on at least one statistic of correlation, mutual information amount, and chi-square value for the variable, and the estimated value, determine the variable An information processing apparatus according to any one of Appendices 5 to 7, characterized by including the process. (Appendix 9) Based on the priority of the variables indicated by a plurality of patterns showing the ranking of a plurality of variables and an estimated value indicating the possibility of conforming to a predetermined condition for each pattern, determine the variable that is the subject of a question asking about importance, Update the estimated value based on the answer result of the question targeting the determined variable. An information processing method in which a computer executes a process. (Appendix 10) The process of updating the estimated value is When the answer result indicates that the determined variable is important, update it so as to increase the estimated value of the pattern in which the priority of the determined variable is equal to or higher than a predetermined threshold. An information processing method according to Appendix 9, characterized by including the process. (Appendix 11) The process of updating the estimated value is When the answer result indicates that the determined variable is not important, update it so as to increase the estimated value of the pattern in which the priority of the determined variable is not equal to or higher than the predetermined threshold. An information processing method according to Appendix 10, characterized by including the process. (Appendix 12) The process of determining the variable is Based on the priority indicated by the pattern determined based on at least one statistic of correlation, mutual information amount, and chi-square value for the variable, and the estimated value, determine the variable. An information processing method according to any one of Appendices 9 to 11, characterized by including the process.
Description of Signs
[0078] 10 Information processing apparatus 10a Communication interface 10b HDD 10c Memory 10d Processor 11 Communication unit 12 Storage unit 13 Control unit 121 Classification Model 122 Index Data 123 Probability Data 131 Decision Unit 132 Output Unit 133 Update Unit
Claims
1. Determine the variable that is the subject of a question about importance based on the priority of the variable based on a plurality of patterns indicating the ranking of a plurality of variables and an estimate indicating the possibility of conforming to a predetermined condition for each pattern, Update the estimate based on the answer result of the question targeting the determined variable An information processing program for causing a computer to execute the process.
2. The process of updating the estimate is When the answer result indicates that the determined variable is important, update so as to increase the estimate of the pattern in which the priority of the determined variable is equal to or higher than a predetermined threshold The information processing program according to claim 1, characterized by including the process.
3. The process of updating the estimate is When the answer result indicates that the determined variable is not important, update so as to increase the estimate of the pattern in which the priority of the determined variable is not equal to or higher than the predetermined threshold The information processing program according to claim 2, characterized by including the process.
4. The process of determining the variable is Based on the priority based on the pattern indicating the ranking determined based on at least one of the correlation, mutual information amount, and chi-square value for the variable and the estimate, determine the variable The information processing program according to any one of claims 1 to 3, characterized by including the process.
5. Determine the variable that is the subject of a question about importance based on the priority of the variable based on a plurality of patterns indicating the ranking of a plurality of variables and an estimate indicating the possibility of conforming to a predetermined condition for each pattern, Update the estimate based on the answer result of the question targeting the determined variable An information processing apparatus including a control unit that executes the process.
6. Determine the variable that is the subject of a question about importance based on the priority of the variable based on a plurality of patterns indicating the ranking of a plurality of variables and an estimate indicating the possibility of conforming to a predetermined condition for each pattern, Update the estimate based on the answer result of the question targeting the determined variable An information processing method executed by a computer.
Citation Information
Patent Citations
AGENT SYSTEM WITH NATURAL CONVERSATION CAPABILITY WITH USER, AGENT CONTROL METHOD, AGENT CONTROL PROGRAM, AND RECORDING MEDIUM
JP2016536691A
Method and device for providing answer in question answering system
JP2017220238A
Information processing device and information processing system
JP2019169147A
Estimate program, device, and method
JP2020144592A
Method for Dynamically Assigning Question Priority Based on Question Extraction and Domain Dictionary
US20180336271A1