Model Building Device and Evaluation Device
The model construction device enhances the accuracy of evaluating respondent reliability in electronic questionnaires by using machine learning to analyze feature differences related to respondent attitude, effectively addressing the limitations of existing methods.
Patent Information
- Application Number
- JP2022568182
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-07
- Filing Date
- 2021-11-29
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2041-11-29
AI Technical Summary
Existing methods for evaluating the reliability of response content from electronic questionnaires lack accuracy, as they fail to effectively account for respondent behavior and attitude.
A model construction device that constructs a learned model for evaluating reliability by acquiring feature amounts related to respondent attitude, deriving differences between reference and target feature amounts, and using machine learning to output evaluation results.
This approach enables more accurate evaluation of respondent reliability, significantly improving the detection of Satisficing behavior compared to previous methods.
Smart Images

Figure 0007696631000001 
Figure 0007696631000002 
Figure 0007696631000003
Abstract
Description
Technical Field
[0001] One aspect of the present invention relates to a model construction device that constructs a learned model for evaluating the reliability of the response content of respondents to electronic questionnaires.
Background Art
[0002] As shown in the following Patent Documents 1 to 3 and Non-Patent Documents 1 to 3, in recent years, various methods for evaluating the reliability of the response content of respondents to electronic questionnaires (e.g., online questionnaires) have been proposed.
[0003] For example, Non-Patent Document 2, which is a paper described by the inventors of the present application (hereinafter referred to as "the inventors"), explains that the respondent's answering attitude of "Satisficing (minimization of effort)" has a great influence on the above reliability. Based on this, Non-Patent Document 2 shows the idea of obtaining a feature amount indicating the presence or absence of Satisficing of the respondent based on the respondent's answer operation log and evaluating the above reliability based on the feature amount.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0006] As described later, based on the above idea, the inventors further intensively studied specific methods for improving the evaluation accuracy of the above reliability. An object of one aspect of the present invention is to evaluate the reliability of the response content of respondents to electronic questionnaires with higher accuracy than in the past.
Means for Solving the Problems
[0007] In order to solve the above problems, a model construction device according to one aspect of the present invention is a model construction device that constructs a learned model for evaluating the reliability of the response content of respondents to electronic questionnaires, and includes a learning feature amount acquisition unit that acquires a set of feature amounts related to the response attitude of the respondents to the electronic questionnaires. The learning feature amount acquisition unit acquires, for a group of respondents composed of a plurality of respondents determined in advance, (i) a learning reference feature amount set that is a set of the feature amounts related to the reference feature amount acquisition electronic questionnaire, and (ii) a learning target feature amount set that is a set of the feature amounts related to the target feature amount acquisition electronic questionnaire. The model construction device further includes a learning difference derivation unit that derives a learning difference feature amount set that is a set indicating the difference between the learning target feature amount set and the learning reference feature amount set, and a learning unit that constructs the learned model that outputs an evaluation result regarding the reliability by performing machine learning based on the learning difference feature amount set.
[0008] Also, in order to solve the above problems, an evaluation apparatus according to an aspect of the present invention is an evaluation apparatus that evaluates the reliability of a respondent's response content to an electronic questionnaire, wherein types of feature quantities related to the respondent's response attitude to the electronic questionnaire are defined in advance, and for a respondent group composed of a plurality of predetermined respondents, (i) a learning target feature quantity set that is a set of the feature quantities related to the target feature quantity acquisition electronic questionnaire, and (ii) a learning reference feature quantity set that is a set of the feature quantities related to the reference feature quantity acquisition electronic questionnaire, a learned model that outputs an evaluation result regarding the reliability based on machine learning based on a learning differential feature quantity set that is a set indicating the difference therebetween is constructed in advance, the evaluation apparatus includes an evaluation feature quantity acquisition unit that acquires a set of the feature quantities related to the electronic questionnaire, and the evaluation feature quantity acquisition unit, for a specific respondent as the respondent, (i) an evaluation reference feature quantity set that is a set of the feature quantities related to the reference feature quantity acquisition electronic questionnaire, and (ii) an evaluation target feature quantity set that is a set of the feature quantities related to the evaluation target electronic questionnaire as the target feature quantity acquisition electronic questionnaire, and acquires, the evaluation apparatus further includes an evaluation differential derivation unit that derives an evaluation differential feature quantity set that is a set indicating the difference between the evaluation target feature quantity set and the evaluation reference feature quantity set, and an evaluation unit that causes the learned model to output the evaluation result regarding the specific respondent with respect to the evaluation target electronic questionnaire by inputting the evaluation differential feature quantity set into the learned model.
Effect of the Invention
[0009] According to one aspect of the present invention, it is possible to evaluate the reliability of a respondent's response content to an electronic questionnaire with higher accuracy than in the past.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Mode for Carrying Out the Invention
[0011] 〔Embodiment 1〕 The information processing apparatus 1 according to Embodiment 1 will be described below. For convenience of explanation, members having the same functions as those described in Embodiment 1 will be denoted by the same reference numerals in the following embodiments, and the description thereof will not be repeated. For the sake of simplicity, descriptions of matters similar to known techniques will also be omitted as appropriate. Note that each configuration and each numerical value described in this specification are merely examples. In the following description, an electronic questionnaire is simply abbreviated as a "questionnaire".
[0012] As an example of the questionnaire, an online questionnaire (typically, a WEB questionnaire) can be cited. However, the form of the questionnaire according to one aspect of the present invention is not particularly limited as long as it is possible to obtain the answers of the respondents by receiving the input operations of the respondents with respect to an electronic device (e.g., the terminal device 900 described below).
[0013] (Overview of the information processing apparatus 1) FIG. 1 is a block diagram showing the configuration of the information processing apparatus 1 and its main surrounding parts. The information processing apparatus 1 is connected to the terminal device 900 via the communication network 800. As an example, the information processing apparatus 1 is a server device (e.g., a cloud server) of a questionnaire provider.
[0014] The terminal device 900 is an example of an electronic device for causing a questionnaire respondent (hereinafter referred to as a "respondent") to give an answer. As an example, the terminal device 900 is a smartphone owned by the respondent. The terminal device 900 in the example of FIG. 1 includes a touch panel 910 (a member in which an input unit and a display unit are integrated). However, of course, in the terminal device 900, the input unit and the display unit may be provided as separate members.
[0015] The terminal device 900 communicates with the information processing apparatus 1 via the communication network 800 to obtain data of a questionnaire page (e.g., a WEB page) provided by a questionnaire provider. In Embodiment 1, the data of each questionnaire page is stored in advance in the storage unit 90 of the information processing apparatus 1.
[0016] Then, the terminal device 900 presents a questionnaire screen to the respondent by displaying a questionnaire page on the touch panel 910 using a predetermined application (e.g., a web browser). As a result, the respondent can be made to answer the questionnaire by performing an input operation (e.g., a touch operation) on the touch panel 910. Hereinafter, the input operation for answering the questionnaire is referred to as an answer operation.
[0017] In FIG. 1, for simplicity, only one terminal device 900 is shown. However, of course, the information processing device 1 can communicate with a plurality of terminal devices 900 via the communication network 800. For this reason, the information processing device 1 can acquire questionnaire answers from each of a plurality of respondents.
[0018] The information processing device 1 includes a control device 10 and a storage unit 90. The control device 10 comprehensively controls each part of the information processing device 1. The control device 10 includes a model construction device 11 and an evaluation device 12. As will be described later, the model construction device 11 constructs a model (evaluation model) for evaluating the reliability of the respondent's answer content to the questionnaire by machine learning. The evaluation model in this specification means a learned model constructed by the model construction device 11. Then, the evaluation device 12 evaluates the reliability of the respondent's answer content to the questionnaire using the evaluation model pre-constructed by the model construction device 11.
[0019] The storage unit 90 stores various types of data and programs used in the processing of the control device 10. In this specification, the log of the respondent's response operations to the questionnaire is referred to as the response operation log. Also, the DB (Database) for the response operation log is referred to as the response operation log DB. As shown in FIG. 1, the response operation log DB 91 is stored in the storage unit 90. The response operation log DB 91 in the example of FIG. 1 generically represents the learning log DB 91a and the evaluation log DB 91b. The learning log DB 91a means the response operation log DB used in a series of processes (learning phase) by the model construction device 11. In contrast, the evaluation log DB 91b means the response operation log DB used in a series of processes (evaluation phase) by the evaluation device 12.
[0020] In this specification, the functional unit that acquires the response operation log is generically referred to as the "log acquisition unit". In Embodiment 1, the log acquisition unit in the model construction device 11 is referred to as the learning log acquisition unit 110a. In contrast, the log acquisition unit in the evaluation device 12 is referred to as the evaluation log acquisition unit 110b. When there is no need to distinguish between the learning log acquisition unit 110a and the evaluation log acquisition unit 110b, both are also generically referred to as the log acquisition unit 110.
[0021] The log acquisition unit 110 acquires the response operation log from the terminal device 900 via the communication network 800. In Embodiment 1, as shown in Non-Patent Document 2, the log acquisition unit 110 detects a predetermined touch event (more specifically, touch start, touch move, and touch end) on the touch panel 910, and based on this, acquires the response operation log related to each touch event. Then, the log acquisition unit 110 updates the response operation log DB 91 by registering the acquired response operation log in the response operation log DB 91.
[0022] In the example of FIG. 1, for convenience of explanation, the learning log acquisition unit 110a and the evaluation log acquisition unit 110b are illustrated as separate functional units. However, of course, in the information processing apparatus 1, the learning log acquisition unit 110a and the evaluation log acquisition unit 110b may be implemented as an integrated functional unit. This also applies to the following: (i) the learning feature quantity acquisition unit 111a and the evaluation feature quantity acquisition unit 111b, and (ii) the learning difference derivation unit 112a and the evaluation difference derivation unit 112b.
[0023] In this specification, a functional unit that acquires feature quantities related to the answering attitude of respondents to questionnaires is generically referred to as the "feature quantity acquisition unit". In Embodiment 1, the feature quantity acquisition unit in the model construction apparatus 11 is referred to as the learning feature quantity acquisition unit 111a. On the other hand, the feature quantity acquisition unit in the evaluation apparatus 12 is referred to as the evaluation feature quantity acquisition unit 111b. When there is no need to distinguish between the learning feature quantity acquisition unit 111a and the evaluation feature quantity acquisition unit 111b, both are also generically referred to as the feature quantity acquisition unit 111.
[0024] In Embodiment 1, as shown in Non-Patent Document 2, the feature quantity acquisition unit 111 acquires feature quantities based on the answer operation logs recorded in the answer operation log DB 91. Specifically, the feature quantity acquisition unit 111 acquires feature quantities by analyzing the answer operation logs. The types of feature quantities acquired by the feature quantity acquisition unit 111 are defined in advance. Main examples of feature quantities include, at the time of answering a questionnaire: (i) the scroll length of the questionnaire screen, (ii) the scroll speed of the questionnaire screen, (iii) the number of text changes, and (iv) the answering time, etc.
[0025] More specifically, in Embodiment 1, the feature quantity acquisition unit 111 acquires reference feature quantities (baseline feature quantities) by analyzing the answer operation logs at the time of answering a reference feature quantity acquisition questionnaire (baseline acquisition questionnaire) described later. In addition, the feature quantity acquisition unit 111 acquires target feature quantities by analyzing the answer operation logs at the time of answering a target questionnaire described later.
[0026] In this specification, the difference (subtraction) between the target feature amount and the reference feature amount is referred to as the differential feature amount. And the functional unit that derives the differential feature amount is generically referred to as the "differential derivation unit". In Embodiment 1, the differential derivation unit in the model construction device 11 is referred to as the learning differential derivation unit 112a. On the other hand, the differential derivation unit in the evaluation device 12 is referred to as the evaluation differential derivation unit 112b. When there is no need to distinguish between the learning differential derivation unit 112a and the evaluation differential derivation unit 112b, both are also generically referred to as the differential derivation unit 112.
[0027] As shown in FIG. 1, the model construction device 11 includes a learning log acquisition unit 110a, a learning feature amount acquisition unit 111a, a learning differential derivation unit 112a, and a learning unit 113. The evaluation device 12 includes an evaluation log acquisition unit 110b, an evaluation feature amount acquisition unit 111b, an evaluation differential derivation unit 112b, and an evaluation unit 123. Hereinafter, prior to the specific description of the operation of the information processing device 1, new findings discovered through the inventors' intensive studies will be described.
[0028] (Inventors' Considerations Regarding Feature Amounts) As shown in Non-Patent Document 2, the behavior of the operation of the terminal device 900 by the respondent at the time of questionnaire answering reflects the respondent's answering attitude. Therefore, the inventors further studied the types of feature amounts that are preferable for the detection of Satisficing.
[0029] At the time of questionnaire answering, it is generally considered that the respondent goes through a series of answering processes of "Process 1: Read the question content" → "Process 2: Think about the answer content for the question content by oneself" → "Process 3: Input an answer based on the answer content thought by oneself".
[0030] The inventors considered that "in a series of response processes, the presence or absence of satisficing by the respondents will be significantly manifested in Process 2." This is because it is expected that respondents with satisficing (hereinafter referred to as "satisficing respondents") will not pay much cognitive cost in considering the response content. Satisficing respondents can be rephrased as respondents with low reliability regarding the response content (inappropriate respondents). On the other hand, for respondents without satisficing (hereinafter referred to as "non-satisficing respondents"), it is expected that they will pay more cognitive cost in considering the response content compared to satisficing respondents. Non-satisficing respondents can be rephrased as respondents with high reliability regarding the response content (appropriate respondents).
[0031] Based on the above considerations, the inventors designed a dummy questionnaire to eliminate the influence of Process 2. Specifically, the inventors designed the above dummy questionnaire based on a new idea of using the feature amount regarding the dummy questionnaire as a reference (baseline) for the feature amount regarding the actual questionnaire (target questionnaire).
[0032] Therefore, in this specification, the above dummy questionnaire is referred to as a baseline acquisition questionnaire. Hereinafter, the baseline is abbreviated as "BL". Accordingly, for example, the baseline acquisition questionnaire is abbreviated as the "BL acquisition questionnaire". In the following description, the feature amount regarding the BL acquisition questionnaire (baseline feature amount) is referred to as the BL feature amount. The BL feature amount may also be referred to as the reference feature amount. For this reason, the BL acquisition questionnaire may also be referred to as a reference feature amount acquisition questionnaire (more specifically, a reference feature amount acquisition electronic questionnaire).
[0033] The questionnaire 290 for BL acquisition in FIG. 2 is an example of the questionnaire for BL acquisition designed by the inventors. The questionnaire 290 for BL acquisition includes a Likert-type questionnaire 291 for BL acquisition and a free description-type questionnaire 292 for BL acquisition. The questionnaire for BL acquisition is designed so that, as BL feature quantities, feature quantities of the same type as the feature quantities (target feature quantities) in the actual questionnaire can be acquired. Specifically, the questionnaire 290 for BL acquisition is designed in the same format as the actual questionnaire.
[0034] Therefore, the questionnaire 290 for BL acquisition is designed in the same format as the target questionnaire 390 described later. For example, the questionnaire 290 for BL acquisition is designed to have the same question format and the same number of questions as the target questionnaire 390. Also, the questionnaire 290 for BL acquisition is designed to have question frames and answer columns of the same size as the target questionnaire 390.
[0035] Note that, in order to surely acquire the BL feature quantities from the respondents, it is preferable that the questionnaire for BL acquisition be presented to the respondents prior to the target questionnaire. This is because, if the questionnaire for BL acquisition is presented to the Satisficing respondents after the target questionnaire, there is a possibility that the Satisficing respondents will close the questionnaire page without answering the questionnaire for BL acquisition.
[0036] Thus, for example, it is preferable that the questionnaire page be designed such that the target questionnaire is presented to the respondents upon completion of answering the questionnaire for BL acquisition. In the experiment described below, using the questionnaire page designed in such a manner, after completion of answering the questionnaire 290 for BL acquisition, the target questionnaire 390 was presented to each respondent, whereby the BL feature quantities and the target feature quantities were respectively acquired.
[0037] As shown in FIG. 2, in the Likert-form BL acquisition questionnaire 291, the response content (options) for each question includes a description (instruction description) that instructs the respondent. Inst11 and Inst12 in FIG. 2 are each an example of the instruction description in the Likert-form BL acquisition questionnaire 291. Inst11 is an instruction description of "Please select 'Does not apply well at all'" for a certain question. Inst12 is an instruction description of "Please select 'Somewhat applicable'" for another question. In this way, the Likert-form BL acquisition questionnaire 291 is designed so that in a Likert-form questionnaire, without going through Process 2, the respondent can be made to answer each question as per the instruction description.
[0038] Similarly, in the free description-form BL acquisition questionnaire 292, an instruction description regarding the response content (description content) for each question is included. Inst2 in FIG. 2 is an example of the instruction description in the free description-form BL acquisition questionnaire 292. In the example of FIG. 2, Inst2 is an instruction description of "Please enter manually in the answer column for a certain question without copying & pasting the following text". Sptxt in FIG. 2 is an example of the text specified for manual entry. In this way, the free description-form BL acquisition questionnaire 292 is designed so that in a free description-form questionnaire, without going through Process 2, the respondent can be made to answer each question as per the instruction description.
[0039] As described above, in the BL acquisition questionnaire, unlike an actual questionnaire, the influence of Process 2 can be excluded, and the respondent can be made to answer each question. Therefore, the feature quantity (BL feature quantity) derived based on the response operation log for the BL acquisition questionnaire (hereinafter referred to as the BL log) is a feature quantity from which the influence of Process 2 has been excluded. The BL feature quantity can also be expressed as a feature quantity (pure feature quantity) that does not reflect the presence or absence of Satisficing of the respondent.
[0040] On the other hand, in the actual questionnaire, unlike the questionnaire for BL acquisition, no instruction description for each question is included (see also FIG. 3 described below). Therefore, in the actual questionnaire, Process 2 affects the response content. In this specification, a questionnaire in which the influence of Process 2 is not excluded is referred to as a true questionnaire. The actual questionnaire is an example of a true questionnaire. Note that the true questionnaire may also be referred to as the original questionnaire.
[0041] The inventors are considering performing a reliability evaluation on the true questionnaire by using the feature amounts related to the true questionnaire. Therefore, in this specification, the feature amounts related to the true questionnaire are also referred to as target feature amounts. Also, the true questionnaire that is the target for acquiring the target feature amounts is also referred to as the target questionnaire. Since the target questionnaire can also be said to be a questionnaire for acquiring the target feature amounts, it may be referred to as a target feature amount acquisition questionnaire (more specifically, a target feature amount acquisition electronic questionnaire).
[0042] The inventors designed a target questionnaire for the experiment to acquire the target feature amounts. The target questionnaire 390 in FIG. 3 shows an example of the target questionnaire designed by the inventors. FIG. 3 is a figure paired with FIG. 2. The target questionnaire 390 includes a Likert-type target questionnaire 391 and a free description-type target questionnaire 392.
[0043] Note that, for the verification of the presence or absence of Satisficing by the inventors, the target questionnaire 390 is designed to include (i) questions for acquiring the Directed Questions scale (hereinafter referred to as DQS questions) and (ii) questions for acquiring the Attentive Responding Scale (hereinafter referred to as ARS questions). Both DQS and ARS are known indices indicating the presence or absence of Satisficing (see Non-Patent Document 3).
[0044] As shown in FIG. 3, in the Likert scale target questionnaire 391, unlike the Likert scale BL acquisition questionnaire 291, there is no instruction description regarding the response content (options) for each question. Similarly, in the free description format target questionnaire 392, unlike the free description format BL acquisition questionnaire 292, there is no instruction description regarding the response content (description content) for each question.
[0045] Thus, unlike the BL acquisition questionnaire 290, the target questionnaire 390 is designed to have the respondents answer each question without eliminating the influence of Process 2. Therefore, the feature quantity (target feature quantity) derived based on the response operation log for the target questionnaire (hereinafter referred to as the target log) is a feature quantity that includes the influence of Process 2, unlike the BL feature quantity. That is, the target feature quantity is a feature quantity in which the presence or absence of the respondent's Satisficing is reflected, unlike the BL feature quantity.
[0046] Based on the above considerations regarding the BL feature quantity and the target feature quantity, the inventors further focused on "the difference between the target feature quantity and the BL feature quantity". As described above, the difference feature quantity is a feature quantity defined as the difference between the target feature quantity (a feature quantity in which the presence or absence of the respondent's Satisficing is reflected) and the BL feature quantity (a feature quantity in which the presence or absence of the respondent's Satisficing is not reflected). By subtracting the BL feature quantity from the target feature quantity, it is considered that the BL feature component (offset value) included in the target feature quantity can be eliminated. From this, the inventors hypothesized that "in the difference feature quantity, the presence or absence of the respondent's Satisficing is reflected more strongly compared to the target feature quantity".
[0047] Therefore, the inventors used the information processing apparatus 1 to conduct an experiment to derive each feature amount based on the above-described BL acquisition questionnaire 290 and the target questionnaire 390. Specifically, the inventors acquired the BL log and the target log for a group of respondents composed of a plurality of predetermined respondents by the log acquisition unit 110. Subsequently, the inventors used the feature amount acquisition unit 111 to obtain a set of BL feature amounts (BL feature amount set) composed of a plurality of types of BL feature amounts and (ii) a set of target feature amounts (target feature amount set) composed of a plurality of types of target feature amounts based on the BL log and the target log, respectively. Subsequently, the inventors used the difference derivation unit 112 to derive a set of difference feature amounts (difference feature amount set) composed of a plurality of types of difference feature amounts from the BL feature amount set and the target feature amount set.
[0048] The inventors conducted a statistical study on each feature amount obtained in the above experiment. Hereinafter, referring to FIG. 4, the main study results by the inventors will be described. FIG. 4 is a graph showing the p-value (significance probability) of each feature amount obtained in the above experiment. The horizontal axis in the graph of FIG. 4 indicates each feature amount (variable). The asterisk (*) in FIG. 4 is a symbol indicating that there is a significant difference. Note that the significant difference in the following description means the significant difference between the normal group and the Satisficing group described later. Further, "there is a significant difference" in the following description means that for a certain feature amount, p < α. α is the significance level. The inventors set α = 0.05 in the above experiment.
[0049] Each feature amount in the graph of FIG. 4 is as follows in order from the left side.
[0050] · longIntervalNum: The number of times the non-operation time is too long; · straightLiningNum: The maximum value of the number of consecutive identical answers; · reselectNum: The number of times of changing the option; · refocusNum: The number of times of refocusing on the text box; ·deleteTextNum: Number of times text is deleted (delete); ·textNum: Average number of characters in the text; ·scrollLength: Average value of the scroll length; ·scrollDuration: Average value of the scroll time; ·scrollSpeed: Average value of the scroll speed; ·reverceScroll: Number of reverse scrolls; ·noAnswer: Number of unanswered questions (number of questions not answered); ·middleAnswer: Number of middle answers; ·deleteTextBaseRate: Deletion rate of text (BL feature); ·deleteTextNum_selfDev: Difference between the target feature and the BL feature for the number of times text is deleted by the same respondent (difference feature); ·deleteRate_selfDev: Difference between the target feature and the BL feature for the deletion rate of text by the same respondent (difference feature); ·textNumDev: Difference between the number of characters in the text of a single respondent and the average value of respondents; ·scrollLengthCov: Coefficient of variation of the scroll length; ·scrollDurationCov: Coefficient of variation of the scroll time; ·scrollSpeedCov: Coefficient of variation of the scroll speed; ·scrollLengthBase: Scroll length (BL feature); ·scrollDurationBase: Scroll time (BL feature); ·scrollSpeedBase: Scroll speed (BL feature); ·scrollLength_selfDev: Difference between the target feature and the BL feature for the scroll length (difference feature); ·scrollDuration_selfDev: Difference between the target feature and the BL feature for the scroll time (difference feature); ·scrollSpeed_selfDev: The difference (differential feature) between the target feature quantity and the BL feature quantity regarding the scroll speed; ·ansTimeLikertFiltered: The average response time in Likert-scale questions; ·ansTimeLikertFilteredCov: The coefficient of variation of the average response time in Likert-scale questions; ·ansTimeTextFiltered: The average response time in free-form questions; ·reselectNumDev: The difference between an individual respondent and the respondent average value regarding the number of times of changing options; ·deleteTextNumDev: The difference between an individual respondent and the respondent average value regarding the number of times of deleting text; ·scrollSpeedDev: The difference between an individual respondent and the respondent average value regarding the scroll speed; ·reverceScrollDev: The difference between an individual respondent and the respondent average value regarding the number of reverse scrolls.
[0051] In the example of Figure 4, the feature quantity with the character "Base" in the variable name is the BL feature quantity. Also, the feature quantity with the character "selfDev" in the variable name is the differential feature quantity. Other feature quantities are the target feature quantities.
[0052] In the example of FIG. 4, the "number of times of excessive non-operation time" means the number of times the time when the respondent is not operating the terminal exceeds a predetermined threshold value. The number of consecutive identical answers means the number of times the same numbered option is continuously selected in a Likert-type question. The scroll length means the amount of movement of the screen (questionnaire screen) in one scroll operation. The scroll speed means the value obtained by dividing the "scroll length" by the "time required for one scroll operation". The number of neutral answers means the number of answers of "neither" in the Likert type. The number of non-responses means the number of questions that have not been answered. The number of characters in the text means the number of characters per free-form question. The text deletion rate means the value obtained by dividing the "number of characters in the text" by the "number of deletion operations for the same text". The coefficient of variation refers to Cov (Coefficient of variation).
[0053] Hereinafter, three target feature quantities, namely, (i) the number of text deletion times, the scroll length (e.g., the average value of the scroll length), and the scroll speed (e.g., the average value of the scroll speed), and (ii) three differential feature quantities corresponding to each of the three target feature quantities will be described.
[0054] In the following description, the differential feature quantity for (i) the number of text deletion times is referred to as "differential number of text deletion times", the differential feature quantity for (ii) the scroll length is referred to as "differential scroll length", and the differential feature quantity for (iii) the scroll speed is referred to as "differential scroll speed", respectively.
[0055] In this specification, "text change" is defined by two elements: "text addition" and "text deletion". Therefore, for example, "the number of text changes" is represented by the sum of "the number of text insertions" and "the number of text deletions". However, in the above experiment, for simplicity, only "text deletion" is handled as an element of "text change". From this, the number of text deletions described in this specification may be understood as an example of "the number of text changes". Therefore, the number of differential text deletions described in this specification can be read as "the number of differential text changes".
[0056] As shown in Figure 4, for the above three target feature quantities "the number of text deletions (deleteTextNum), scroll length (scrollLength), and scroll speed (scrollSpeed)", it was confirmed that there is a significant difference only in scrollSpeed among the three target feature quantities.
[0057] On the other hand, for the above three differential feature quantities "the number of differential text deletions (deleteTextNum_selfDev), differential scroll length (dscrollLength_selfDev), and differential scroll speed (scrollSpeed_selfDev)", it was confirmed that there are significant differences in all of the three differential feature quantities.
[0058] The results shown in Figure 4 support the above hypothesis that "in the differential feature quantities, the presence or absence of Satisficing of the respondents is more strongly emphasized and reflected compared to the target feature quantities". Furthermore, the results also confirm that for the number of text deletions and scroll length, the presence or absence of Satisficing that could not be appropriately standardized by the target feature quantities can be appropriately standardized by the differential feature quantities.
[0059] Subsequently, referring to the box-and-whisker plots shown in Figures 5 to 10, the above three target feature quantities and the above three differential feature quantities will be further examined.
[0060] First, the inventors evaluated whether each respondent belonging to the respondent group has Satisficing based on the DQS obtained by the DQS question. When the DQS of a respondent was below a predetermined value, the inventors attached a "no DQS" label to the respondent. Note that the fact that the DQS is below the predetermined value suggests that the respondent does not have Satisficing. On the other hand, when the DQS was greater than the predetermined value, the inventors attached a "DQS" label to the respondent. Subsequently, the inventors determined that (i) the respondents with the "no DQS" label are non-Satisficing respondents and (ii) the respondents with the "DQS" label are Satisficing respondents, respectively.
[0061] Furthermore, the inventors evaluated whether each respondent belonging to the respondent group has Satisficing based on the ARS obtained by the ARS question. When both Inconsistency and Infrequency among the ARSs of a respondent were below a predetermined value, the inventors attached a "no ARS" label to the respondent. Note that the fact that both Inconsistency and Infrequency are below the predetermined value suggests that the respondent does not have Satisficing.
[0062] On the other hand, when either Inconsistency or Infrequency was greater than the predetermined value, the inventors attached an "OR" label to the respondent. Furthermore, when both Inconsistency and Infrequency were greater than the predetermined value, the inventors attached an "AND" label to the respondent. Subsequently, the inventors determined that (i) the respondents with the "no ARS" label are non-Satisficing respondents and (ii) the respondents with the "OR" label or the "AND" label are non-Satisficing respondents, respectively.
[0063] In the following description, the group of Satisficing respondents is referred to as the Satisficing group, and the group of non-Satisficing respondents is referred to as the normal group, respectively. As described above, in the evaluation based on DQS, the group of respondents with the no DQS label (no DQS group) is the normal group, and the group of respondents with the DQS label (DQS group) is the Satisficing group. On the other hand, in the evaluation based on ARS, the group of respondents with the no ARS label (no ARS group) is the normal group, and the union of the group of respondents with the OR label (OR group) and the group of respondents with the AND label (AND group) is the Satisficing group.
[0064] Figures 5, 7, and 9 each show box plots indicating the distributions of deleteTextNum, scrollLength, and scrollSpeed in the respondent group. In contrast, Figures 6, 8, and 10 each show box plots indicating the distributions of deleteTextNum_selfDev, scrollLength_selfDev, and scrollSpeed_selfDev in the respondent group. Figures 6, 8, and 10 are paired with Figures 5, 7, and 9, respectively.
[0065] Graph 591 in Figure 5 is a box plot showing the distribution of deleteTextNum for each of the no DQS group and the DQS group. Graph 592 is a box plot showing the distribution of deleteTextNum for each of the no ARS group, the OR group, and the AND group. Graph 691 in Figure 6 is a box plot showing the distribution of deleteTextNum_selfDev for each of the no DQS group and the DQS group. Graph 692 is a box plot showing the distribution of deleteTextNum_selfDev for each of the no ARS group, the OR group, and the AND group.
[0066] Graph 791 in Figure 7 is a box plot showing the distribution of scrollLength for each of the no DQS group and the DQS group. Graph 792 is a box plot showing the distribution of scrollLength for each of the no ARS group, the OR group, and the AND group. Graph 891 in Figure 8 is a box plot showing the distribution of scrollLength_selfDe for each of the no DQS group and the DQS group. Graph 892 is a box plot showing the distribution of scrollLength_selfDe for each of the no ARS group, the OR group, and the AND group.
[0067] Graph 991 in Figure 9 is a box plot showing the distribution of scrollSpeed for each of the no DQS group and the DQS group. Graph 992 is a box plot showing the distribution of scrollSpeed for each of the no ARS group, the OR group, and the AND group. Graph 1091 in Figure 10 is a box plot showing the distribution of scrollSpeed_selfDev for each of the no DQS group and the DQS group. Graph 1092 is a box plot showing the distribution of scrollSpeed_selfDev for each of the no ARS group, the OR group, and the AND group.
[0068] As shown in Figures 5, 7, and 9, for the above three target feature quantities (deleteTextNum, scrollLength, and scrollSpeed), it was confirmed that there was a significant difference only in scrollSpeed among the three target feature quantities. In contrast, as shown in Figures 6, 8, and 10, for the above three differential feature quantities (deleteTextNum_selfDev, scrollLength_selfDev, and scrollSpeed_selfDev), it was confirmed that there was a significant difference in any of the differential feature quantities.
[0069] These results shown in FIGS. 5 to 10 also support the above hypothesis that "in the differential feature amounts, the presence or absence of Satisficing of the respondents is more strongly emphasized and reflected compared to the target feature amounts". Further, each of the results shown in FIGS. 6, 8, and 10 suggests that the above three differential feature amounts are particularly effective as feature amounts for discriminating the presence or absence of Satisficing.
[0070] (An example of the operation of the information processing apparatus 1) As described above, the inventors have found a new finding that "according to the differential feature amount (the difference between the target feature amount and the BL feature amount), the presence or absence of Satisficing of the respondents can be further emphasized and indexed compared to the target feature amount". And based on this finding, the inventors created the information processing apparatus 1. Hereinafter, an example of the operation of the information processing apparatus 1 will be described. The operation of the information processing apparatus 1 is roughly classified into the operation in the learning phase (the operation of the model construction apparatus 11) and the operation in the evaluation phase (the operation of the evaluation apparatus 12).
[0071] (1: Learning phase) Hereinafter, with reference to FIG. 11, the learning phase will be described. FIG. 11 is a flowchart exemplifying the flow of processes S1 to S8 in the learning phase. First, the model construction apparatus 11 causes a respondent group composed of a plurality of predetermined respondents to answer a BL acquisition questionnaire. Hereinafter, the number of respondents belonging to the respondent group is represented as N. As an example, N = 2000. Specifically, the model construction apparatus 11 causes the terminal device 900 owned by each of the plurality of respondents belonging to the respondent group (that is, each of the N terminal devices 900) to display the BL acquisition questionnaire (S1). As the BL acquisition questionnaire, for example, the above-described BL acquisition questionnaire 290 may be used.
[0072] The learning log acquisition unit 110a acquires, from each of the N terminal devices 900, an answer operation log for the BL acquisition questionnaire (hereinafter referred to as the learning BL log) for the respondent group (S2). Then, the learning log acquisition unit 110a updates the learning log DB 91a by registering the acquired learning BL log in the learning log DB 91a.
[0073] The learning feature quantity acquisition unit 111a acquires, for the respondent group, a set of feature quantities related to the BL acquisition questionnaire (hereinafter referred to as the learning BL feature quantity set) by analyzing the learning BL log recorded in the learning log DB 91a (S3). The learning BL feature quantity set may also be referred to as the learning reference feature quantity set.
[0074] Subsequently, the model construction device 11 causes the respondent group to answer the learning target questionnaire (the target questionnaire for machine learning). Specifically, the model construction device 11 causes each of the N terminal devices 900 to display the learning target questionnaire (S4). The learning target questionnaire may also be referred to as the learning target feature quantity acquisition questionnaire (more specifically, the learning target feature quantity acquisition electronic questionnaire). As the learning target questionnaire, for example, the above-described target questionnaire 390 may be used.
[0075] The learning log acquisition unit 110a acquires, from each of the N terminal devices 900, an answer operation log for the learning target questionnaire (hereinafter referred to as the learning target log) for the respondent group (S5). Then, the learning log acquisition unit 110a updates the learning log DB 91a by registering the acquired learning target log in the learning log DB 91a.
[0076] The learning feature quantity acquisition unit 111a acquires, for the respondent group, a set of feature quantities related to the learning target questionnaire (hereinafter referred to as the learning target feature quantity set) by analyzing the learning target log recorded in the learning log DB 91a (S6).
[0077] As described above, in the learning phase, by having the same group of respondents answer the BL acquisition questionnaire and the learning target questionnaire, N learning BL feature sets and N learning target feature sets are respectively obtained. Also, as described above, the BL acquisition questionnaire is designed so that it can obtain the same type of features as the learning target features as the learning BL features. Therefore, the learning BL feature set and the learning target feature set contain the same type of features. Accordingly, the learning differential feature set described below also contains the same type of features as the learning BL feature set and the learning target feature set.
[0078] Subsequently, the learning differential derivation unit 112a derives a set showing the difference between the learning target feature set and the learning BL feature set (hereinafter referred to as the learning differential feature set) (S7). Specifically, the learning differential derivation unit 112a derives N learning differential feature sets.
[0079] As described above, the three differential features, namely the differential text deletion count, the differential scroll length, and the differential scroll speed, are expected to be effective as features for determining the presence or absence of Satisficing. Therefore, in the learning differential derivation unit 112a, it is preferable that the BL acquisition questionnaire and the learning target questionnaire are designed so that these three differential features can be derived.
[0080] Specifically, the BL acquisition questionnaire is preferably designed so that it can obtain, as BL features, (i) the number of text deletions (BL text deletion count) at the time of answering the BL acquisition questionnaire, (ii) the scroll length (BL scroll length) at the time of answering the BL acquisition questionnaire, and (iii) the scroll speed (BL scroll speed) at the time of answering the BL acquisition questionnaire. Note that the BL text deletion count is the "reference text deletion count" and the BL scroll length is the "reference scroll length" and the BL scroll speed is the "reference scroll speed" and Each may be referred to as such. As described above, the number of BL text deletions can be read as the number of BL text changes. Therefore, the number of reference text deletions can be read as the number of reference text changes.
[0081] Similarly, the learning target questionnaire is preferably designed so that (i) the number of text deletions (target text deletion count) when answering the learning target questionnaire, (ii) the scroll length (target scroll length) when answering the learning target questionnaire, and (iii) the scroll speed (target scroll speed) when answering the learning target questionnaire can be obtained as target feature amounts. The target text deletion count can be read as the target text change count.
[0082] When the BL acquisition questionnaire is designed as described above, the learning feature amount acquisition unit 111a can obtain a set of learning BL feature amounts including the BL text deletion count, the BL scroll length, and the BL scroll speed from the BL acquisition questionnaire. Similarly, when the learning target questionnaire is designed as described above, the learning feature amount acquisition unit 111a can obtain a set of learning target feature amounts including the target text change count, the target scroll length, and the target scroll speed from the learning target questionnaire.
[0083] Therefore, the learning difference derivation unit 112a Derives the difference between the target text deletion count and the BL text deletion count as the difference text deletion count, Derives the difference between the target scroll length and the BL scroll length as the difference scroll length, Derives the difference between the target scroll speed and the BL scroll speed as the difference scroll speed, respectively. The difference text deletion count can be read as the difference text change count.
[0084] As described above, when the questionnaire for BL acquisition and the target questionnaire for learning are designed as described above, the learning differential derivation unit 112a can derive a set of learning differential feature quantities including the number of differential text changes, the differential scroll length, and the differential scroll speed. Based on the set of learning differential feature quantities, by causing the learning unit 113 described below to construct an evaluation model (learned model), the accuracy of the reliability evaluation by the evaluation model can be further improved.
[0085] The learning unit 113 constructs an evaluation model by performing machine learning based on the set of learning differential feature quantities (S8). The learning unit 113 constructs an evaluation model by machine learning, which outputs an evaluation result regarding the reliability of the answer content of the respondent to the questionnaire. In Embodiment 1, the learning unit 113 constructs an evaluation model by performing supervised learning.
[0086] Known algorithms may be used for supervised learning. Main examples of the supervised learning algorithm in one aspect of the present invention include LightGBM (a decision tree-based gradient boosting algorithm), random forest, XGBoost, and logistic regression. However, of course, the supervised learning algorithms that can be used in one aspect of the present invention are not limited to these.
[0087] In Embodiment 1, the learning unit 113 is supplied with teacher data including (i) N pieces of learning data and (ii) N pieces of label data (correct answer data) corresponding to each of the N pieces of learning data. The learning unit 113 constructs an evaluation model by performing learning using the teacher data. The label data is created in advance based on the evaluation results by any known method for evaluating the reliability (e.g., the evaluation results based on DQS or ARS described above).
[0088] In the label data in Embodiment 1, 0 or 1 is assigned as the label value. The label value 1 indicates that the response content of a certain respondent's (one respondent out of N respondents) learning target questionnaire is "inappropriate". That is, the label value 1 indicates that the above-mentioned one respondent is a Satisficing respondent. In contrast, the label value 0 indicates that the response content of the above-mentioned one respondent's learning target questionnaire is "appropriate". That is, the label value 0 indicates that the above-mentioned one respondent is a non-Satisficing respondent.
[0089] In Embodiment 1, the learning data includes learning differential feature amounts. Specifically, each of the N learning data includes each of the N sets of learning differential feature amounts. Therefore, the learning unit 113 constructs an evaluation model by performing machine learning using the N sets of learning differential feature amounts as input variables (explanatory variables). The evaluation model outputs an evaluation result as an output variable (objective variable).
[0090] Note that it is preferable that the learning data further includes a set of learning target feature amounts. Specifically, it is preferable that each of the N learning data further includes each of the N sets of learning target feature amounts. In this case, the learning unit 113 constructs an evaluation model by performing machine learning using the set of learning differential feature amounts and the set of learning target feature amounts as input variables.
[0091] In this way, it is preferable that the learning unit 113 constructs an evaluation model by performing machine learning using the set of learning differential feature amounts and the set of learning target feature amounts as input variables. By constructing an evaluation model using learning data including more input variables, the accuracy of the reliability evaluation by the evaluation model can be further improved.
[0092] In addition, the learning data may further include a learning BL feature set in addition to the learning differential feature set and the learning target feature set. Specifically, each of the N learning data preferably further includes each of the N learning BL feature sets. In this case, the learning unit 113 constructs an evaluation model by performing machine learning using the learning differential feature set, the learning target feature set, and the learning BL feature set as input variables. By constructing an evaluation model using learning data that includes more input variables, it is also expected to further improve the accuracy of the reliability evaluation by the evaluation model.
[0093] (2: Evaluation Phase) The evaluation phase is a series of processes following the learning phase. Hereinafter, with reference to FIG. 12, the evaluation phase will be described. FIG. 12 is a flowchart illustrating the flow of processes S11 to S18 in the evaluation phase. FIG. 12 is a figure paired with FIG. 11. In the evaluation phase, the evaluation device 12 uses the evaluation model pre-constructed by the model construction device 11 to evaluate the reliability of the response content of a specific respondent (hereinafter referred to as the specific respondent) with respect to the target questionnaire as the evaluation target. Hereinafter, the target questionnaire as the evaluation target in the evaluation phase is referred to as the evaluation target questionnaire (more specifically, the evaluation target electronic questionnaire). The specific respondent may (i) belong to the respondent group in the learning phase, or (ii) may not belong to the respondent group.
[0094] First, the evaluation device 12 causes the specific respondent to answer the BL questionnaire. Specifically, the evaluation device 12 causes the terminal device 900 owned by the specific respondent to display the BL questionnaire (S11). The same BL questionnaire as in the learning phase is used also in the evaluation phase.
[0095] The evaluation log acquisition unit 110b acquires, from the terminal device 900, an answer operation log for the BL questionnaire (hereinafter referred to as the evaluation BL log) for a specific respondent (S12). Then, the evaluation log acquisition unit 110b updates the evaluation log DB 91b by registering the acquired evaluation BL log in the evaluation log DB 91b.
[0096] The evaluation feature quantity acquisition unit 111b acquires, for a specific respondent, a set of feature quantities related to the BL questionnaire (hereinafter referred to as the evaluation BL feature quantity set) by analyzing the evaluation BL log recorded in the evaluation log DB 91b (S13). The evaluation BL feature quantity set may also be referred to as the evaluation reference feature quantity set.
[0097] Subsequently, the evaluation device 12 causes the specific respondent to answer the evaluation target questionnaire. Specifically, the evaluation device 12 causes the terminal device 900 owned by the specific respondent to display the evaluation target questionnaire (S14). The evaluation target questionnaire is designed in the same format as the learning target questionnaire.
[0098] The evaluation log acquisition unit 110b acquires, from the terminal device 900, an answer operation log for the evaluation target questionnaire (hereinafter referred to as the evaluation target log) for a specific respondent (S15). Then, the evaluation log acquisition unit 110b updates the evaluation log DB 91b by registering the acquired evaluation target log in the evaluation log DB 91b.
[0099] The evaluation feature quantity acquisition unit 111b acquires, for a specific respondent, a set of feature quantities related to the evaluation target questionnaire (hereinafter referred to as the evaluation target feature quantity set) by analyzing the evaluation target log recorded in the evaluation log DB 91b (S16).
[0100] The evaluation difference derivation unit 112b derives a set showing the difference between the evaluation target feature quantity set and the evaluation BL feature quantity set (hereinafter referred to as the evaluation difference feature quantity set) (S17). As described above, in the evaluation phase, by having a specific respondent answer the BL questionnaire and the evaluation target questionnaire, the evaluation BL feature quantity set and the evaluation target feature quantity set are respectively obtained. And according to the BL questionnaire, as the evaluation BL feature quantity, a feature quantity of the same type as the evaluation target feature quantity can be obtained. For this reason, the evaluation BL feature quantity set and the evaluation target feature quantity set include the same type of feature quantities. Therefore, the evaluation difference feature quantity set also includes the same type of feature quantities as the evaluation BL feature quantity set and the evaluation target feature quantity set.
[0101] The evaluation unit 123 acquires the evaluation model from the learning unit 113. Then, the evaluation unit 123 inputs the evaluation difference feature quantity set into the evaluation model, causing the evaluation model to output an evaluation result for the specific respondent regarding the evaluation target questionnaire (S18). In other words, the evaluation unit 123 inputs the evaluation difference feature quantity set into the evaluation model as an input variable, and obtains the evaluation result as an output variable from the evaluation model.
[0102] In Embodiment 1, the output value (evaluation result) output by the evaluation model is binary data similar to the above-mentioned label value. The output value 1 indicates that the answer content of the specific respondent to the evaluation target questionnaire is "inappropriate". That is, the output value 1 indicates that the specific respondent is a Satisficing respondent. In contrast, the output value 0 indicates that the answer content of the specific respondent to the evaluation target questionnaire is "appropriate". That is, the output value 0 indicates that the specific respondent is a non-Satisficing respondent.
[0103] The evaluation unit 123 may perform a predetermined process according to the evaluation result obtained from the evaluation model. For example, when the output value 1 is output from the evaluation model, the evaluation unit 123 may cause the terminal device 900 to display a warning message. Thereby, it is possible to prompt a specific respondent evaluated as a Satisficing respondent to improve the answering attitude.
[0104] Note that, similar to the above description of the learning differential feature amount set, the evaluation differential feature amount set preferably includes three differential feature amounts: the differential text change count (e.g., the differential text deletion count), the differential scroll length, and the differential scroll speed. By inputting the evaluation differential feature amount set including the three differential feature amounts into the evaluation model, the accuracy of the evaluation result output by the evaluation model can be further improved.
[0105] Therefore, the evaluation BL feature amount set preferably includes three BL feature amounts: the BL text change count (e.g., the BL text deletion count), the BL scroll length, and the BL scroll speed. Similarly, the evaluation target feature amount set preferably includes three target feature amounts: the target text change count (e.g., the target text deletion count), the target scroll length, and the target scroll speed. In this case, the evaluation differential derivation unit 112b can derive the evaluation differential feature amount set including the three differential feature amounts.
[0106] Note that in the learning phase, when the evaluation model is constructed by machine learning using the learning differential feature amount set and the learning target feature amount set as input variables, the evaluation unit 123 inputs the evaluation differential feature amount set and the evaluation target feature amount set as input variables into the evaluation model, thereby causing the evaluation model to output an evaluation result.
[0107] Also, in the learning phase, when an evaluation model is constructed by machine learning using the learning differential feature set, the learning target feature set, and the learning BL feature set as input variables, the evaluation unit 123 inputs the evaluation differential feature set, the evaluation target feature set, and the evaluation BL feature set as input variables into the evaluation model, thereby causing the evaluation model to output an evaluation result.
[0108] (Effect) As described above, the inventors obtained a new idea of "acquiring BL features (features from which the influence of Process 2 has been eliminated)", which is not disclosed or suggested in any prior art documents. On this basis, the inventors created the information processing apparatus 1 based on a further new idea of "using differential features to further emphasize and index the presence or absence of Satisficing of respondents compared to target features".
[0109] As described above, in the learning phase, an evaluation model is constructed based on N learning differential feature sets (differential feature sets obtained from N respondents belonging to the respondent group). The evaluation model thus constructed generally reflects the presence or absence of Satisficing of N respondents. Therefore, according to the evaluation model, in the evaluation phase, the reliability of the response content of a respondent (specific respondent) regardless of belonging to the respondent group can be evaluated with high accuracy.
[0110] In order to verify the effectiveness of the evaluation model constructed by the information processing apparatus 1, the inventors conducted an evaluation experiment on 48 test data using the evaluation model. In the evaluation experiment, the inventors used LightGBM as the algorithm for supervised learning to construct the evaluation model. Also, the types of each feature used in the learning phase and the evaluation phase are as shown in the example of FIG. 4.
[0111] In the evaluation experiment by the inventors, ·TP (True Positive, number of true positives) = 18 ·TN (True Negative, number of true negatives) = 21 ·FP (False Positive, number of false positives) = 4 ·FN (False Negative, number of false negatives) = 5 The following results were obtained. In the above results, "positive" means that the respondent is a Satisficing respondent. In contrast, "negative" means that the respondent is a non-Satisficing respondent.
[0112] Based on the above results, when calculating the "recall", Recall = TP / (TP + FN) = 18 / 23 ≈ 0.783 It was. Recall is one of the indicators showing the detection accuracy of inappropriate answers by machine learning.
[0113] Thus, as a result of the evaluation experiment, it was confirmed that a high detection accuracy of 78.3% was achieved by the evaluation model constructed by the information processing apparatus 1. The detection accuracy is sufficiently high compared to the detection accuracy of inappropriate answers by conventional machine learning. For example, in an experimental example disclosed in Non-Patent Document 3, the detection accuracy of inappropriate answers is 55.6%. As described above, according to the information processing apparatus 1, the reliability of the answer content of the respondent to the questionnaire can be evaluated with higher accuracy than before.
[0114] [Embodiment 2] In Embodiment 1, the case of constructing an evaluation model by supervised learning was exemplified. However, the machine learning algorithm in one aspect of the present invention is not limited to supervised learning. As long as the learning unit 113 can construct an evaluation model by machine learning. Therefore, the learning unit 113 may construct an evaluation model by unsupervised learning.
[0115] Unsupervised learning may use known algorithms. As the main examples of unsupervised learning algorithms in one aspect of the present invention, the k-nearest neighbor method, principal component analysis, support vector machines, and k-means clustering can be mentioned. However, of course, the unsupervised learning algorithms that can be used in one aspect of the present invention are not limited to these.
[0116] When unsupervised learning is adopted, an evaluation model can be constructed without preparing correct answer data in advance. Therefore, when unsupervised learning is adopted, an evaluation model can be constructed more easily than when supervised learning is adopted. However, when supervised learning is adopted, it is also expected that a more reliable evaluation model can be constructed by performing learning using correct answer data compared to the case where unsupervised learning is adopted. Whether to adopt supervised learning or unsupervised learning as the machine learning algorithm may be appropriately selected by the user of the information processing apparatus 1.
[0117] 〔Modification Example〕 (1) In each of the above-described embodiments, the case where a plurality of respondents answer the questionnaire using their respective individual terminal devices 900 has been exemplified. However, the terminal device according to one aspect of the present invention may be a terminal device that is commonly used among a plurality of respondents. For example, the terminal device may be one terminal device (for convenience, referred to as terminal device 900C) fixedly installed in a facility (e.g., a commercial facility or a public facility). The terminal device 900C only needs to be connected to the information processing apparatus 1 via the communication network 800. As the main examples of the terminal device 900C, digital signage, kiosk terminals, and information search PCs (Personal Computers) can be mentioned.
[0118] In this case, by having each respondent who visits the facility operate the terminal device 900C, each respondent can be made to answer the questionnaire. The information processing apparatus 1 can acquire the response operation log of each respondent from the terminal device 900C, and can acquire the feature amount of each respondent based on the response operation log. As described above, the terminal device according to one aspect of the present invention may be a mobile terminal or a stationary terminal device.
[0119] (2) In each of the above-described embodiments, the case where the feature amount is acquired by analyzing the response operation log has been exemplified. However, the method for acquiring the feature amount in the feature amount acquisition unit 111 is arbitrary and is not limited to the above example. Therefore, the information processing apparatus 1 does not necessarily have to have the log acquisition unit 110.
[0120] As an example, the information processing apparatus 1 may acquire a moving image obtained by photographing the state of the respondent's questionnaire response from an imaging device (for example, a surveillance camera installed around the terminal device 900C in the facility) via the communication network 800. The information processing apparatus 1 stores the moving image in the storage unit 90. Then, the feature amount acquisition unit 111 acquires the moving image stored in the storage unit 90, and acquires the feature amount by analyzing the moving image.
[0121] As described above, the feature amount acquisition unit 111 only needs to be configured to be able to acquire the feature amount. However, as described in Embodiment 1, it is preferable that the feature amount acquisition unit 111 acquires the feature amount by analyzing the response operation log. This is because it is not necessary to provide additional hardware (for example, an imaging device) when acquiring the feature amount based on the response operation log. In addition, since the data amount of the response operation log is smaller than the data amount of the moving image, it is not necessary to provide a large-capacity storage unit 90. Thus, according to the feature amount acquisition unit 111 of Embodiment 1, the feature amount can be acquired inexpensively and simply. Therefore, the information processing apparatus 1 (particularly, the model construction apparatus 11 and the evaluation apparatus 12) can be provided inexpensively and simply.
[0122] 〔Example of Realization by Software〕 The control blocks of the information processing apparatus 1 (particularly, the model construction apparatus 11 and the evaluation apparatus 12) may be implemented by a logic circuit (hardware) formed on an integrated circuit (IC chip) or the like, or may be implemented by software.
[0123] In the latter case, the information processing apparatus 1 includes a computer that executes instructions of a program, which is software for realizing each function. This computer includes, for example, one or more processors and a computer-readable recording medium storing the above program. Then, in the above computer, when the above processor reads and executes the above program from the above recording medium, the object of one aspect of the present invention is achieved. As the above processor, for example, a CPU (Central Processing Unit) can be used. As the above recording medium, a "non-transitory tangible medium", for example, in addition to a ROM (Read Only Memory), a tape, a disk, a card, a semiconductor memory, a programmable logic circuit, etc. can be used. Further, it may further include a RAM (Random Access Memory) for expanding the above program. Also, the above program may be supplied to the above computer via any transmission medium (communication network, broadcast wave, etc.) capable of transmitting the program. Note that one aspect of the present invention can also be realized in the form of a data signal embedded in a carrier wave, in which the above program is embodied by electronic transmission.
[0124] 〔Supplementary Notes〕 One aspect of the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
Explanation of Reference Numerals
[0125] 1 Information processing apparatus 10 Control apparatus 11 Model construction apparatus 12 Evaluation apparatus 90 Memory unit 91a Learning log DB 91b Evaluation log DB 110a Learning log acquisition unit 110b Evaluation log acquisition unit 111a Learning feature quantity acquisition unit 111b Evaluation feature quantity acquisition unit 112a Learning difference derivation unit 112b Evaluation difference derivation unit 113 Learning unit 123 Evaluation unit 290 BL acquisition questionnaire (reference feature quantity acquisition electronic questionnaire) 390 Target questionnaire (target feature quantity acquisition electronic questionnaire)
Claims
1. A model construction device for constructing a learned model for evaluating the reliability of respondents' answers to an electronic questionnaire, comprising a learning feature quantity acquisition unit that acquires a set of feature quantities related to the answering attitude of the respondent to the electronic questionnaire, The set of feature quantities includes, at the time of answering the electronic questionnaire, (i) the number of text changes, and (ii) the scroll length of the screen of the electronic questionnaire, and (iii) the scroll speed of the screen of the electronic questionnaire, The learning feature quantity acquisition unit, for a group of respondents composed of a plurality of respondents determined in advance, (i) a learning reference feature quantity set that is a set of the feature quantities related to the reference feature quantity acquisition electronic questionnaire, and (ii) a learning target feature quantity set that is a set of the feature quantities related to the target feature quantity acquisition electronic questionnaire, and acquires The reference feature quantity acquisition electronic questionnaire includes an instruction description to the respondent regarding the answer content to each question in the reference feature quantity acquisition electronic questionnaire, The learning reference feature quantity set includes, at the time of answering the reference feature quantity acquisition electronic questionnaire, (i) a reference text change number that is the number of text changes, and (ii) a reference scroll length that is the scroll length of the screen of the reference feature quantity acquisition electronic questionnaire, and (iii) a reference scroll speed that is the scroll speed of the screen of the reference feature quantity acquisition electronic questionnaire, The target feature quantity acquisition electronic questionnaire does not include an instruction description to the respondent regarding the answer content to each question in the target feature quantity acquisition electronic questionnaire, The learning target feature quantity set includes, at the time of answering the target feature quantity acquisition electronic questionnaire, (i) the number of changes to the text, i.e., the target text change count, and (ii) the target scroll length, which is the scroll length of the screen of the electronic questionnaire for obtaining the target feature amount, and (iii) the target scroll speed, which is the scroll speed of the screen of the electronic questionnaire for obtaining the target feature amount, and The model construction device a learning difference derivation unit that derives a learning difference feature amount set, which is a set showing the difference between the learning target feature amount set and the learning reference feature amount set; a learning unit that constructs the learned model that outputs the evaluation result regarding the reliability by performing machine learning based on the learning difference feature amount set; The learning difference feature amount set (i) the difference in the number of text changes, i.e., the difference text change count, which is the difference between the target text change count and the reference text change count, and (ii) the difference in the scroll length, i.e., the difference scroll length, which is the difference between the target scroll length and the reference scroll length, and (iii) the difference in the scroll speed, i.e., the difference scroll speed, which is the difference between the target scroll speed and the reference scroll speed; Model construction device.
2. The learning feature amount acquisition unit according to claim 1, wherein the learning feature amount acquisition unit acquires the set of feature amounts by analyzing a log of the answer operation of the respondent to the electronic questionnaire.
3. The learning unit according to claim 1 or 2, wherein the learning unit constructs the learned model by performing the machine learning with the learning target feature amount set as an input variable together with the learning difference feature amount set.
4. The learning unit according to any one of claims 1 to 3, wherein the learning unit constructs the learned model by performing supervised learning as the machine learning.
5. The learning unit constructs the learned model by performing unsupervised learning as the machine learning, according to any one of claims 1 to 3.
6. An evaluation device for evaluating the reliability of the respondent's answer content for an electronic questionnaire, The types of feature quantities related to the respondent's answering attitude for the electronic questionnaire are predefined, The types of the feature quantities are, at the time of answering the electronic questionnaire, (i) the number of text changes, and (ii) the scroll length of the screen of the electronic questionnaire, and (iii) the scroll speed of the screen of the electronic questionnaire, and include, For a group of respondents composed of a plurality of predefined respondents, (i) a learning target feature quantity set which is a set of the feature quantities for the target feature quantity acquisition electronic questionnaire, and (ii) a learning reference feature quantity set which is a set of the feature quantities for the reference feature quantity acquisition electronic questionnaire, a learned model that outputs an evaluation result for the reliability based on a learning differential feature quantity set which is a set showing the difference between them is pre-constructed in the learning phase, The reference feature quantity acquisition electronic questionnaire includes an instruction description for the respondent regarding the answer content for each question in the reference feature quantity acquisition electronic questionnaire, The learning reference feature quantity set is, at the time of answering the reference feature quantity acquisition electronic questionnaire, (i) a reference text change number which is the number of text changes, and (ii) a reference scroll length which is the scroll length of the screen of the reference feature quantity acquisition electronic questionnaire, and (iii) a reference scroll speed which is the scroll speed of the screen of the reference feature quantity acquisition electronic questionnaire, and include, The electronic questionnaire for acquiring the target feature amount does not include an instruction description for the respondent regarding the response content for each question in the electronic questionnaire for acquiring the target feature amount. The target feature amount set for learning includes (i) the target text change count, which is the number of text changes, and (ii) the target scroll length, which is the scroll length of the screen of the electronic questionnaire for acquiring the target feature amount, and (iii) the target scroll speed, which is the scroll speed of the screen of the electronic questionnaire for acquiring the target feature amount. The differential feature amount set for learning includes (i) the differential text change count in the learning phase, which is the difference between the target text change count indicated by the target feature amount set for learning and the reference text change count indicated by the reference feature amount set for learning, and (ii) the differential scroll length in the learning phase, which is the difference between the target scroll length indicated by the target feature amount set for learning and the reference scroll length indicated by the reference feature amount set for learning, and (iii) the differential scroll speed in the learning phase, which is the difference between the target scroll speed indicated by the target feature amount set for learning and the reference scroll speed indicated by the reference feature amount set for learning. The evaluation device includes an evaluation feature amount acquisition unit that acquires the set of the feature amounts regarding the electronic questionnaire. In an evaluation phase following the learning phase, the evaluation feature amount acquisition unit acquires, for a specific respondent as the respondent, (i) an evaluation reference feature amount set, which is the set of the feature amounts regarding the reference feature amount acquisition electronic questionnaire, and (ii) an evaluation target feature amount set, which is the set of the feature amounts regarding the evaluation target electronic questionnaire as the electronic questionnaire for acquiring the target feature amount. The above evaluation reference feature amount set is for the above specific respondent when answering the above reference feature amount acquisition electronic questionnaire, (i) the reference text change count, which is the number of text changes, and (ii) the reference scroll length, which is the scroll length of the screen of the above reference feature amount acquisition electronic questionnaire, and (iii) the reference scroll speed, which is the scroll speed of the screen of the above reference feature amount acquisition electronic questionnaire, and The above evaluation target feature amount set is for the above specific respondent when answering the above evaluation target electronic questionnaire, (i) the target text change count, which is the number of text changes, and (ii) the target scroll length, which is the scroll length of the screen of the above evaluation target electronic questionnaire, and (iii) the target scroll speed, which is the scroll speed of the screen of the above evaluation target electronic questionnaire, and The above evaluation device In the above evaluation phase, an evaluation difference derivation unit that derives an evaluation difference feature amount set, which is a set showing the difference between the above evaluation target feature amount set and the above evaluation reference feature amount set, In the above evaluation phase, an evaluation unit that inputs the above evaluation difference feature amount set into the above learned model to cause the above learned model to output the above evaluation result for the above specific respondent with respect to the above evaluation target electronic questionnaire, and The above evaluation difference feature amount set (i) the difference text change count in the above evaluation phase, which is the difference between the target text change count indicated by the above evaluation target feature amount set and the reference text change count indicated by the above evaluation reference feature amount set, and (ii) the difference scroll length in the above evaluation phase, which is the difference between the target scroll length indicated by the above evaluation target feature amount set and the reference scroll length indicated by the above evaluation reference feature amount set, and (iii) the differential scroll speed in the evaluation phase, which is the difference between the target scroll speed indicated by the target feature quantity set for evaluation and the reference scroll speed indicated by the reference feature quantity set for evaluation, An evaluation device.
Citation Information
Patent Citations
Method for investigating questionnaire, questionnaire system and recording medium
JP2002092291A
Method and system for automatically removing reliability-doubtful answer of questionnaire research using communication terminal
JP2002342531A
Questionnaire tabulation system
JP2013012120A
Reply quality determination device and reply quality determination method
JP2015114971A
Questionnaire data checkup device and its program
JP2017191517A