Information processing device, method, program, and system

By calculating question importance and creating a second classification model that excludes less critical questions, the program ensures accurate clustering with fewer questions, improving response rates and enabling effective customer analysis.

JP7680505B2Active Publication Date: 2025-05-20LTD ADK マ ー ケ TE ィ ン グ · ソ リ ュ ー シ ョ ン ズ
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023141895
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2025-05-20
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

The challenge lies in balancing the number of questions in a questionnaire survey to maintain response quality and sincerity while ensuring accurate clustering, as too many questions lead to decreased response rates and too few questions compromise clustering accuracy.

Method used

A program that calculates the importance of questions using a first classification model based on supervised learning, excludes less important questions to create a second classification model, allowing for accurate clustering with a reduced number of questions.

Benefits of technology

This approach enhances response rates by reducing the number of questions needed for accurate clustering, facilitating effective customer analysis and targeted promotional activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007680505000001
    Figure 0007680505000001
  • Figure 0007680505000002
    Figure 0007680505000002
  • Figure 0007680505000003
    Figure 0007680505000003
Patent Text Reader

Abstract

To properly determine a cluster to which a respondent belongs, while reducing the number of questions requiring a reply.SOLUTION: A program makes a computer function as: means for calculating a degree of importance of a plurality of questions with respect to classification conducted in accordance with a first classification model created according to supervised learning using first learning data that is based on reply contents of a plurality of respondents with respect to a questionnaire survey composed of multiple questions, and on a respective clustering result of the plurality of respondents; means for acquiring second learning data in which data related to some questions has been removed from the first learning data on the basis of the degree of importance; and means for creating a second classification model by conducting supervised learning using the second learning data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to an information processing device, a method, a program, and a system. [Background technology]

[0002] For example, by conducting a questionnaire survey on a small group (sample) extracted from a population, such as users of a specific website, and clustering the responses, it is possible to visualize the characteristic clusters (groups) that make up that population.

[0003] Patent Document 1 discloses a technical idea that aims to efficiently and accurately classify responses to open-ended questions in a questionnaire by applying machine learning techniques. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2023-037406 A Summary of the Invention [Problem to be solved by the invention]

[0005] In general, the more questions there are, the more complicated the task of answering them becomes, which raises concerns about problems such as a decline in the response rate to the questionnaire survey and an increase in those who do not answer sincerely. On the other hand, the fewer the number of questions, the more likely it is that the quality of clustering will decline, making it impossible to define appropriate clusters or to properly determine (classify) which cluster a respondent belongs to.

[0006] An objective of the present disclosure is to appropriately determine the cluster to which a respondent belongs while reducing the number of questions that need to be answered. [Means for solving the problem]

[0007] A program according to one embodiment of the present disclosure causes a computer to function as: means for calculating the importance of a plurality of questions with respect to classification performed by a first classification model created by supervised learning using first learning data based on answers given by a plurality of respondents to a questionnaire survey consisting of a plurality of questions and clustering results for each of the plurality of respondents; means for acquiring second learning data by excluding data relating to some of the questions from the first learning data based on the importance; and means for performing supervised learning using the second learning data to create a second classification model. [Brief description of the drawings]

[0008] [Figure 1] 1 is a block diagram showing a configuration of an information processing system according to an embodiment of the present invention; [Diagram 2] FIG. 2 is a block diagram showing a configuration of a client device according to the present embodiment. [Diagram 3] FIG. 2 is a block diagram showing a configuration of a server according to the present embodiment. [Figure 4] FIG. 1 is an explanatory diagram of one aspect of the present embodiment. [Diagram 5] FIG. 4 is a diagram showing a data structure of first learning data in the present embodiment. [Figure 6] FIG. 4 is a diagram showing a data structure of second learning data in the present embodiment. [Figure 7] 11 is a flowchart of a first classification model creation process according to the present embodiment. [Figure 8] 13 is a flowchart of a second classification model creation process according to the present embodiment. [Figure 9] 11A to 11C are diagrams illustrating example screens displayed in the second classification model creation process of the present embodiment. [Figure 10] 13 is a flowchart of a second classification model editing process according to the present embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. In the drawings for explaining the embodiment, the same components are generally designated by the same reference numerals, and the repeated description will be omitted.

[0010] (1) Information Processing System Configuration The configuration of the information processing system will be described below. Fig. 1 is a block diagram showing the configuration of the information processing system according to the present embodiment.

[0011] As shown in FIG. 1, the information processing system 1 includes a client device 10 and a server 30. The client device 10 and the server 30 are connected via a network (eg, the Internet or an intranet) NW.

[0012] The client device 10 is an example of an information processing device that transmits a request to the server 30. The client device 10 is, for example, a smartphone, a tablet terminal, or a personal computer. A user of the client device 10 is a person who receives a second classification model creation service, which will be described later.

[0013] The server 30 is an example of an information processing device that provides the client device 10 with a response in response to a request transmitted from the client device 10. The server 30 is, for example, a server computer.

[0014] (1-1) Client device configuration The configuration of the client device will now be described with reference to Fig. 2, which is a block diagram showing the configuration of the client device of this embodiment.

[0015] 2, the client device 10 includes a storage device 11, a processor 12, an input / output interface 13, and a communication interface 14. The client device 10 is connected to a display 21.

[0016] The storage device 11 is configured to store programs and data. The storage device 11 is, for example, a combination of a Read Only Memory (ROM), a Random Access Memory (RAM), and a storage (for example, a flash memory or a hard disk).

[0017] The programs include, for example, the following programs: ·OS (Operating System) programs · Programs for applications that process information (e.g. web browsers)

[0018] The data includes, for example, the following data: Databases referenced in information processing Data obtained by performing information processing (i.e., the results of performing information processing)

[0019] The processor 12 is a computer that realizes the functions of the client device 10 by running a program stored in the storage device 11. The processor 12 is, for example, at least one of the following: ·CPU(Central Processing Unit) ·GPU(Graphic Processing Unit) ·ASIC(Application Specific Integrated Circuit) ·FPGA(Field Programmable Gate Array)

[0020] The input / output interface 13 is configured to obtain information (eg, a user's instruction) from an input device connected to the client device 10, and to output information (eg, an image) to an output device connected to the client device 10. The input device is, for example, a keyboard, a pointing device, a touch panel, or a combination thereof. The output device is, for example, a display 21, a speaker, or a combination thereof.

[0021] The communication interface 14 is configured to control communications between the client device 10 and an external device (eg, a server 30).

[0022] The display 21 is configured to display an image (a still image or a moving image). The display 21 is, for example, a liquid crystal display or an organic EL display.

[0023] (1-2) Server configuration The configuration of the server will now be described with reference to Fig. 3, which is a block diagram showing the configuration of the server according to this embodiment.

[0024] As shown in FIG. 3, the server 30 includes a storage device 31, a processor 32, an input / output interface 33, and a communication interface .

[0025] The storage device 31 is configured to store programs and data, and is, for example, a combination of a ROM, a RAM, and a storage (for example, a flash memory or a hard disk).

[0026] The programs include, for example, the following programs: -OS programs Application programs that perform information processing

[0027] The data includes, for example, the following data: Databases referenced in information processing Results of information processing

[0028] The processor 32 is a computer that realizes the functions of the server 30 by starting a program stored in the storage device 31. The processor 32 is, for example, at least one of the following. ·CPU GPU ·ASIC FPGA

[0029] The input / output interface 33 is configured to obtain information (eg, a user's instruction) from an input device connected to the server 30, and to output information (eg, an image) to an output device connected to the server 30. The input device is, for example, a keyboard, a pointing device, a touch panel, or a combination thereof. The output device is, for example, a display.

[0030] The communication interface 34 is configured to control communications between the server 30 and external devices (eg, the client device 10).

[0031] (2) One aspect of the embodiment An embodiment of the present invention will now be described with reference to Fig. 4, which is an explanatory diagram of an embodiment of the present invention.

[0032] As shown in FIG. 4, a questionnaire survey including a plurality of questions is conducted on a large number of respondents RS1. The answer result AR2 includes the answer content to each of the plurality of questions for each respondent. In other words, the answer result AR2 of each respondent can be expressed as a vector whose elements are the answer content to each question. The questionnaire survey typically uses a web questionnaire consisting of multiple choice questions, but is not limited to this example.

[0033] The analyst AN3 performs clustering of the respondents RS1 based on the answer results AR2. That is, the analyst AN3 classifies the respondents RS1 into multiple clusters according to the similarity of the answer results AR2. As a clustering method, for example, an algorithm such as the k-means method can be used. The analyst AN3 may also narrow down the question (answer) items to be referenced according to the purpose of clustering, or integrate items by utilizing factor analysis. That is, the number of questions to be referenced in clustering can be reduced compared to the number of questions included in the original questionnaire survey. The analyst AN3 is typically a human being, but clustering can also be performed by such machine learning (unsupervised learning).

[0034] The server 30 acquires first learning data based on the answer results of each respondent RS1 and the clustering results (i.e., information indicating a combination of answers given by each respondent RS1 to multiple questions (questions referred to in clustering) and a cluster into which the respondent RS1 is classified based on the combination of answers). The server 30 creates a first classification model LM4 by learning (supervised learning using the first learning data) about the correlation between the combination of answers given by each respondent to multiple questions and the cluster into which the respondent is classified based on the combination of answers. According to the first classification model LM4, even if a new respondent who was not included in the respondent RS1 answers multiple questions (the same number as the number of questions referred to in clustering, for example 30), the cluster to which the new respondent belongs can be determined (classified). Note that, instead of the server 30, another information processing device may perform the process for creating the first classification model LM4.

[0035] The server 30 individually calculates the importance of the above-mentioned multiple questions (multiple questions referred to in clustering) for the classification performed by the first classification model LM4. The server 30 acquires second learning data in which data (feature values) related to some questions are excluded from the first learning data based on the calculated importance. The server 30 creates a second classification model by learning (supervised learning using the second learning data) about the combination of the answers to the multiple questions (questions not subject to exclusion) by each respondent and the correlation with the cluster into which the respondent is classified. According to the second classification model, even if a new respondent who is not included in the respondent RS1 answers the above-mentioned multiple questions (for example, 10 questions, which is less than the number of questions referred to in clustering), the cluster to which the new respondent belongs can be appropriately determined (classified). In other words, the number of questions necessary to classify a respondent into a cluster can be reduced, thereby increasing the response rate. For example, by applying the second classification model to customer analysis, classification can be performed for a larger number of customers, making it easier to carry out effective promotional activities (for example, advertising) according to the classification results.

[0036] (3) Training data The learning data of this embodiment will be described.

[0037] (3-1) First training data The first learning data of this embodiment will be described below. Fig. 5 is a diagram showing the data structure of the first learning data of this embodiment.

[0038] As shown in FIG. 5, the first learning data consists of a plurality of sets of data. Each data consists of a feature consisting of a plurality of elements (N elements in this example) and correct answer data corresponding to the feature. N is an integer equal to or greater than 2. Each element of the feature represents the content of the answer given by the corresponding respondent to N questions (questions referenced in clustering) included in the questionnaire. For example, the element of answer #1 is the content of the answer to question #1. The correct answer data corresponding to the feature represents the cluster to which the corresponding respondent was assigned in clustering.

[0039] (3-2) Second training data The second learning data of this embodiment will be described below with reference to Fig. 6, which is a diagram showing the data structure of the second learning data of this embodiment.

[0040] As shown in FIG. 6, the second learning data consists of a plurality of sets of data. Each data consists of a feature consisting of a plurality of elements (Ni in this example) and correct answer data corresponding to the feature. i is an integer equal to or greater than 1 and smaller than N. The second learning data is obtained by excluding specific i elements from the feature of the first learning data. That is, each element of the feature represents the content of the answer given by the corresponding respondent to N questions included in the questionnaire (Ni questions obtained by excluding i questions from the questions referred to in the clustering). For example, the element of answer #4 is the content of the answer to question #4. The correct answer data corresponding to the feature represents the cluster to which the corresponding respondent is assigned in the clustering.

[0041] (4) Information processing The information processing of this embodiment will be described.

[0042] (4-1) First classification model creation process The first classification model creation process of this embodiment will be described with reference to Fig. 7, which is a flowchart of the first classification model creation process of this embodiment.

[0043] The process in FIG. 7 may be started in response to an instruction (e.g., an instruction to create a first classification model) from a user of the client device 10 or another information processing device. Alternatively, the process in FIG. 7 may be started with the completion of clustering (i.e., completion of creation of the first learning data) as a trigger. However, the start conditions of the process in FIG. 7 are not limited to these examples. The first classification model creation process may be executed by an information processing device different from the server 30 (i.e., the entity executing the second classification model creation process).

[0044] As shown in FIG. 7, the server 30 acquires the first learning data (S130). Specifically, the server 30 acquires the first learning data (FIG. 5) from the storage device 31 or an external information processing device.

[0045] As a first example of acquiring the first learning data (S130), the client device 10 (or another information processing device) transmits a file (e.g., a file in CSV format) that stores input / output data of the clustering performed in advance (i.e., combinations of answers to multiple questions for each respondent and information indicating the cluster to which the respondent is determined to belong) to the server 30. The server 30 acquires the file as the first learning data.

[0046] As a second example of acquiring the first learning data (S130), a user (who may be a person different from the person who receives the second classification model creation service) of the client device 10 (or another information processing device) selects a file corresponding to any one of the clusterings from among files that store input / output data of multiple clusterings performed in the past, which are stored in the storage device 31 or an external storage device (which may include a cloud server, etc.). The client device 10 transmits information indicating the selected file to the server 30. The server 30 acquires the selected file as the first learning data.

[0047] As a third example of acquiring the first learning data (S130), the server 30 acquires, as the first learning data, a file that stores input and output data of the clustering executed immediately before the first classification model creation process.

[0048] After step S130, the server 30 executes creation of a first classification model (S131). Specifically, the server 30 constructs a first classification model by supervised learning using the first learning data acquired in step S130. The first classification model is learned (trained) so as to infer corresponding correct answer data (corresponding respondent cluster) for each input of feature amounts included in the first learning data.

[0049] The first classification model may have a decision tree structure, but is not limited thereto and may be any model capable of solving a classification model. As the machine learning algorithm of the first classification model, for example, LightGBM can be used. However, the machine learning algorithm of the first classification model is not limited to LightGBM, and other gradient boosting decision tree algorithms (for example, XGBoost, RandomForest, or Catboos, etc.), or random forests, etc. can be used.

[0050] The server 30 may use a part of the first learning data acquired in step S130 for validation (i.e., adjusting hyperparameters) or testing (model evaluation) rather than training the first classification model. The ratio or number of data for validation or testing may be predetermined, may be specified by a user of the client device 10 (or another information processing device), or may be determined by an algorithm.

[0051] (4-2) Second classification model creation process The second classification model creation process of this embodiment will be described. Fig. 8 is a flowchart of the second classification model creation process of this embodiment. Fig. 9 is a diagram showing an example of a screen displayed in the second classification model creation process of this embodiment.

[0052] The process of Fig. 8 may be started in response to an instruction from the user of the client device 10 (e.g., an instruction to create a second classification model). Alternatively, the process of Fig. 8 may be started with the completion of the first classification model creation process (Fig. 7) as a trigger. However, the start conditions of the process of Fig. 8 are not limited to these examples.

[0053] As shown in FIG. 8, the server 30 executes acquisition of model data (S230). Specifically, the server 30 acquires the first classification model, which is the basis of the second classification model, and the first learning data from the storage device 31 or an external information processing device.

[0054] As a first example of acquiring model data (S230), the client device 10 transmits the first classification model and the first training data to the server 30. The server 30 acquires the first classification model and the first training data transmitted by the client device 10.

[0055] As a second example of acquiring model data (S230), the user of the client device 10 selects one of the first classification models created in the past and stored in the storage device 31 or an external storage device (which may include a cloud server, etc.). Alternatively, the user may select the first learning data instead of the first classification model. The client device 10 transmits information indicating the selected first classification model to the server 30. The server 30 acquires the selected first classification model and the first learning data used to train the first classification model.

[0056] As a third example of model data acquisition (S230), the server 30 acquires the first classification model created by the first classification model creation process (Figure 7) executed immediately before the second classification model creation process, and the first training data used to train the first classification model.

[0057] After step S230, the server 30 executes calculation of the importance of each question (S231). Specifically, the server 30 calculates an index of the influence of each feature on the classification process by the first classification model as the importance of the question for that feature. As an example, when the first classification model has a decision tree structure, the server 30 calculates a value based on at least one of the gain or frequency of each feature as the importance of the question corresponding to that feature. The gain of a feature indicates how close the feature can be to a leaf (objective variable) by using the feature in the decision tree structure represented by the first classification model. The frequency of a feature indicates the number of times that feature is adopted as a node in the decision tree structure represented by the first classification model.

[0058] After step S231, the server 30 executes output of the importance (S232). Specifically, the server 30 outputs the importance of each question calculated in step S231. As an example, the server 30 transmits information for displaying the importance of each question to the client device 10. This information may be, for example, information required to generate a screen showing the importance of each question, or information about the screen. The client device 10 displays, for example, the screen of FIG. 9 on the display 21.

[0059] The screen in FIG. 9 includes objects J20 to J23. Object J20 displays information to prompt the user of client device 10 to select questions to be retained in the simple questionnaire survey (hereinafter referred to as "target questions") or questions to be excluded from the simple questionnaire survey (hereinafter referred to as "non-target questions") from among the multiple questions referenced in clustering. Here, the simple questionnaire survey refers to a questionnaire survey that is answered by subjects who are to be classified using the second classification model.

[0060] The object J21 displays the importance of each question calculated in step S231. Although the object J21 shows the questions in descending order of importance, the object J21 may be configured to show the questions in ascending order of importance. Alternatively, the object J21 may be configured to change the sorting method in response to a user instruction.

[0061] A corresponding object J21a is placed for each question. The object J21a receives a user instruction to specify the corresponding question as a target question or a non-target question (that is, a user instruction to select a target question or a non-target question). Object J22 displays the total number of selected target questions.

[0062] The object J23 accepts a user instruction to complete the selection of the target question / non-target question. When the object J23 is selected, the client device 10 generates a selection result according to the user instruction based on the input state of the object J21a corresponding to each question, and transmits the selection result to the server 30.

[0063] After step S232, the server 30 executes acquisition of the selection result (S233). Specifically, the server 30 receives from the client device 10 the selection result of the question for the importance information output in step S232.

[0064] After step S233, the server 30 acquires the second learning data (S234). Specifically, first, the server 30 identifies a target question (i.e., a question to be maintained in the simple questionnaire survey) from among a plurality of questions corresponding to the feature quantities of the first learning data acquired in step S230. As a first example, the server 30 identifies the target question according to the selection result acquired in step S233. That is, the server 30 identifies the target question according to whether each of the plurality of questions is designated as a target question or a non-target question. As a second example, the server 30 identifies the target question according to a predetermined algorithm (e.g., identifying target questions in descending order of importance of the corresponding feature quantities within a range not exceeding a predetermined upper limit number, or identifying a question whose corresponding importance is equal to or greater than a threshold value as a target question). According to the second example, steps S232 to S233 can be omitted. The third example is a combination of the first and second examples. For example, the server 30 may set, for example, an upper limit number of target questions or a threshold value for distinguishing target questions in response to a user instruction, and identify the target question according to a predetermined algorithm. Alternatively, the server 30 may determine some of the target questions in response to a user instruction and determine the remainder of the target questions according to an algorithm. Then, the server 30 acquires second learning data by excluding elements (feature amounts) corresponding to non-target questions that do not correspond to the target questions, from the feature amounts of the first learning data acquired in step S230.

[0065] After step S234, the server 30 executes creation of a second classification model (S235). Specifically, the server 30 constructs a second classification model by supervised learning using the second learning data acquired in step S234. The second classification model is learned (trained) so as to infer corresponding correct answer data (corresponding respondent cluster) for each input of feature amounts included in the second learning data.

[0066] The second classification model may have a decision tree structure, but is not limited thereto and may be any model capable of solving a classification model. As the machine learning algorithm of the second classification model, for example, LightGBM can be used. However, the machine learning algorithm of the second classification model is not limited to LightGBM, and other gradient boosting decision tree algorithms (for example, XGBoost, RandomForest, or Catboos, etc.), or random forests, etc. can be used.

[0067] The server 30 may use a part of the second learning data acquired in step S234 for validation or testing rather than training the second classification model. The ratio or number of pieces of validation or testing data may be predetermined, may be specified by the user of the client device 10, or may be determined by an algorithm.

[0068] (4-3) Second classification model editing process The second classification model editing process of this embodiment will be described with reference to the flowchart of FIG 10.

[0069] The process of FIG. 10 may be started in response to an instruction (e.g., an instruction to edit the second classification model) from a user of the client device 10 or another information processing device. Alternatively, the process of FIG. 10 may be started with the completion of the second classification model creation process (FIG. 8) as a trigger. However, the start conditions of the process of FIG. 10 are not limited to these examples. The second classification model editing process may be executed by an information processing device other than the server 30 (i.e., the executing entity of the second classification model creation process).

[0070] As shown in FIG. 10, the server 30 executes acquisition of model data (S330). Specifically, the server 30 acquires the second classification model to be edited and the second learning data from the storage device 31 or an external information processing device.

[0071] As a first example of acquiring model data (S330), the client device 10 or another information processing device transmits the second classification model to be edited and the second learning data used to train the second classification model to the server 30. The server 30 acquires the second classification model and the second learning data transmitted by the client device 10.

[0072] As a second example of acquiring model data (S330), a user of the client device 10 or another information processing device (which may be a person different from the person who receives the provision of the second classification model creation service) selects a desired second classification model as a second classification model to be edited from among second classification models created in the past and stored in the storage device 31 or an external storage device (which may include a cloud server, etc.). Alternatively, instead of the second classification model, the second learning data used to train the second classification model may be selected. The client device 10 transmits information indicating the selected second classification model to the server 30. The server 30 acquires the selected second classification model and the second learning data used to train the second classification model.

[0073] As a third example of model data acquisition (S330), the server 30 acquires the second classification model (the second classification model to be edited) created by the second classification model creation process (Figure 8) executed immediately before the second classification model editing process, and the second learning data used to learn the second classification model.

[0074] Furthermore, in acquiring model data (S330), the server 30 may acquire a first classification model that is a base of the second classification model to be edited, and first learning data used to learn the first classification model. The server 30 may also acquire information on the classification accuracy of the first classification model that is a base of the second classification model to be edited. In addition, if the second classification model to be edited has been edited in the past, the server 30 may acquire at least one of the second classification model before the edit, the second learning data used to learn the second classification model before the edit, or information on the classification accuracy of the second classification model before the edit.

[0075] After step S330, the server 30 executes calculation of classification accuracy (S331). Specifically, the server 30 calculates the classification accuracy of the second classification model to be edited, acquired in step S330. When the process transitions from step S334 to step S331, the server 30 calculates the classification accuracy of the second classification model recreated in step S334. As an index of classification accuracy, for example, an accuracy rate, a precision rate, a recall rate, or a combination thereof (for example, an F value), etc. can be adopted. For example, test data that is not used in training and verifying the second classification model, among the second learning data acquired in step S330, can be used to calculate the classification accuracy.

[0076] Furthermore, the server 30 may calculate the classification accuracy of any of the following models. The first classification model that is the basis for the second classification model to be edited - The second classification model before editing (the second classification model at any point in time if one or more editing processes have been performed on the second classification model)

[0077] After step S331, the server 30 executes output of the classification accuracy (S332). Specifically, the server 30 outputs the classification accuracy calculated in step S331. As a first example, the server 30 outputs only the classification accuracy of the second classification model to be edited. As a second example, the server 30 outputs the classification accuracy of the comparison target in addition to the classification accuracy of the second classification model to be edited. The classification accuracy of the comparison target may be the classification accuracy of at least one of the first classification model that is the base of the second classification model to be edited or the second classification model before editing, or may be a value determined independently of the second classification model to be edited.

[0078] As an example, the server 30 transmits information for displaying the classification accuracy to the client device 10 or another information processing device. This information may be, for example, information required for generating a screen showing the classification accuracy, or information about the screen. The client device 10 displays the screen showing the classification accuracy on the display 21. Although not shown, such a screen may include, in addition to an object for displaying the classification accuracy, an object for accepting a user instruction for changing the target question or a user instruction for ending the second classification model editing process.

[0079] After step S332, the server 30 accepts a user instruction to change the target question or a user instruction to end the second classification model editing process. The user instruction to change the target question may include at least one of the following: An instruction to exclude features included in the current second learning data (in other words, an instruction to change a target question to a non-target question) An instruction to add (restore) features that were included in the first training data but were excluded from the current second training data (in other words, an instruction to change a non-target question to a target question)

[0080] When a user instruction to end the second classification model editing process is received, the server 30 ends the second classification model editing process.

[0081] On the other hand, when a user instruction to change the target question is received, the server 30 executes a change of the second learning data (S333). Specifically, the server 30 changes the second learning data acquired in step S330 in response to the received user instruction. As a first example, the server 30 removes the feature quantity instructed to be removed from the feature quantities of the second learning data. As a second example, the server 30 adds the feature quantity instructed to be added (returned) from the feature quantities of the first learning data to the second learning data. The third example is a combination of the first and second examples.

[0082] After step S333, the server 30 executes re-creation of the second classification model (S334). Specifically, the server 30 constructs a second classification model by supervised learning using the second learning data changed in step S333. As in the creation of the second classification model (S225) described above, the second classification model is learned (trained) so as to infer corresponding correct answer data (corresponding respondent cluster) for each input of feature amounts included in the changed second learning data.

[0083] After step S334, the server 30 re-executes the steps of calculating the classification accuracy (S331) through outputting the classification accuracy (S332). However, in calculating the classification accuracy (S331), the server 30 calculates the classification accuracy of the second classification model created in the immediately preceding re-creation of the second classification model (S334), but can omit the calculation of the classification accuracy of the first classification model and the second classification model before editing.

[0084] (5) Summary As described above, the server 30 of the present embodiment calculates the importance of a plurality of questions for classification by the first classification model created by supervised learning using the first learning data based on the contents of answers of a plurality of respondents to a questionnaire survey consisting of a plurality of questions and the clustering results of each of the plurality of respondents. The server 30 acquires second learning data in which data related to some questions are excluded from the first learning data based on the importance, and performs supervised learning using the second learning data to create a second classification model. This makes it possible to appropriately classify a cluster to which a new respondent belongs by simply conducting a simplified questionnaire survey consisting of fewer questions than the plurality of questions referred to in the clustering. In other words, the number of questions required to classify a respondent into a cluster can be reduced, thereby increasing the response rate. For example, by applying the second classification model to customer analysis, classification can be performed for a larger number of customers, making it easier to carry out effective promotional activities (e.g., advertising) according to the classification results.

[0085] The server 30 may output the importance of a plurality of questions, so that the user (a person receiving the second classification model creation service) can recognize the importance of each question.

[0086] The server 30 may output information for displaying a plurality of questions in ascending or descending order of importance, thereby enabling the user to easily recognize the importance of each question.

[0087] The server 30 may specify a target question from among the multiple questions in response to a user instruction. The second learning data may be data obtained by removing from the first learning data feature quantities corresponding to non-target questions that do not fall under the target questions from among the multiple questions. This allows the user to adjust questions to be included in the simple questionnaire survey or questions not to be included in the simple questionnaire survey.

[0088] The user instruction may indicate an upper limit number of target questions. The server 30 may identify target questions in descending order of importance, up to a limit of the number of target questions. This allows the user to control the upper limit number of questions to be included in the simple questionnaire survey.

[0089] The user instruction may specify whether at least some of the questions are target questions or non-target questions. The server 30 may identify target questions depending on whether each of the questions is specified as a target question or a non-target question. This allows the user to freely select questions to be included in the quick questionnaire survey or questions not to be included in the quick questionnaire survey.

[0090] The importance of each of the plurality of questions may be based on at least one of a gain or a frequency of a feature corresponding to the answer to the question in the first classification model, thereby enabling the importance of each question to be reasonably quantified.

[0091] The server 30 may calculate the classification accuracy of the first classification model and the classification accuracy of the second classification model, and output the classification accuracy of the first classification model and the classification accuracy of the second classification model, allowing the user to easily compare the classification accuracy of the second classification model with the classification accuracy of the first classification model on which it is based.

[0092] The server 30 may identify a target question among the multiple questions. The second learning data may be data obtained by removing from the first learning data a feature corresponding to a non-target question that does not correspond to the target question among the multiple questions. The server 30 may be configured to receive a user instruction to change the target question after outputting the classification accuracy of the first classification model and the classification accuracy of the second classification model, and when a user instruction to change the target question is received, the server 30 may change the second learning data in response to the change in the target question, perform supervised learning using the changed second learning data, and recreate the second classification model. This allows the user to improve the classification accuracy of the second classification model by trial and error in determining questions to be included in the simple questionnaire survey or questions not to be included in the simple questionnaire survey.

[0093] (6) Other modifications The storage device 11 may be connected to the client device 10 via a network NW. The display 21 may be integrated with the client device 10. The storage device 31 may be connected to the server 30 via the network NW.

[0094] Each step of the above information processing can be executed by either the client device 10 or the server 30. For example, a process described as being performed by one device may be executed by another device, and a process described as being performed by the interaction of multiple devices may be executed by a single device. In addition, although the above description shows an example in which each step is executed in a specific order in each process, the execution order of each step is not limited to the described example as long as there is no dependency between the steps.

[0095] The plurality of questions may all be multiple choice questions, which makes it easier to learn the correspondence between the answers and the feature values, and is expected to improve the validity of the classification model.

[0096] Although the embodiment of the present invention has been described in detail above, the scope of the present invention is not limited to the above embodiment. Furthermore, the above embodiment can be improved or modified in various ways without departing from the spirit of the present invention. Furthermore, the above embodiment and the modified examples can be combined. [Explanation of symbols]

[0097] 1: Information processing system 10: Client device 11:Storage device 12: Processor 13: Input / Output Interface 14: Communication interface 21: Display 30: Server 31:Storage device 32 : Processor 33: Input / Output Interface 34: Communication interface

Claims

1. Computer, a means for calculating the importance of features, which are data related to the plurality of questions, for classification of respondents, which features are generated by supervised learning using first learning data based on answers of a plurality of respondents to a questionnaire survey consisting of a plurality of questions and on respective clustering results obtained by classifying the plurality of respondents based on combinations of the answers, and which is performed by a first classification model having a decision tree structure, the importance being a value based on at least one of a gain or a frequency of the features; a means for acquiring second learning data by excluding data relating to some questions from the first learning data based on the importance; A means for performing supervised learning using the second learning data to create a second classification model; Functioning as a the means for acquiring the second learning data accepts a selection input of a target question to be retained in a simple questionnaire survey to be answered by a subject to be classified using the second classification model, or a non-target question to be excluded from the simple questionnaire survey, among the questions referred to in the clustering, and acquires the second learning data by excluding a feature corresponding to the non-target question; The means for creating the second classification model creates the second classification model by learning a correlation between a combination of the contents of the answers to the target questions and a cluster into which the answerer is classified. program.

2. causing the computer to function as a means for outputting the importance of the plurality of questions; The program according to claim 1.

3. the output means outputs information for displaying the plurality of questions in ascending or descending order of importance. The program according to claim 2.

4. causing the computer to function as a means for identifying a target question from among the plurality of questions in response to a user instruction; the second learning data is data obtained by removing from the first learning data feature quantities corresponding to non-target questions that do not correspond to the target questions among the plurality of questions. The program according to claim 1.

5. the user instruction indicates an upper limit number of the target questions; the specifying means specifies the target questions in descending order of importance, without exceeding an upper limit number of the target questions. The program according to claim 4.

6. the user instruction designates whether at least some of the plurality of questions are the target questions or the non-target questions; the specifying means specifies the target question depending on whether each of the plurality of questions is designated as the target question or the non-target question; The program according to claim 4.

7. The computer, A means for calculating a classification accuracy of the first classification model and a classification accuracy of the second classification model; a means for outputting the classification accuracy of the first classification model and the classification accuracy of the second classification model; The program according to claim 1 , further functioning as:

8. causing the computer to function as a means for identifying a target question from among the plurality of questions; the second learning data is data obtained by removing, from the first learning data, feature quantities corresponding to non-target questions that do not correspond to target questions among the plurality of questions; the specifying means is configured to receive a user instruction to change the target question after outputting the classification accuracy of the first classification model and the classification accuracy of the second classification model; the means for creating the second classification model, when a user instruction to change the target question is accepted, changes the second learning data in response to the change in the target question, and performs supervised learning using the changed second learning data to recreate the second classification model. The program according to claim 7.

9. The plurality of questions are all multiple choice questions. The program according to claim 1.

10. The computer a step of calculating the importance of features, which are data related to the plurality of questions, for classification of respondents, using a first classification model having a decision tree structure, the features being created by supervised learning using first learning data based on answers of a plurality of respondents to a questionnaire survey consisting of a plurality of questions and respective clustering results in which the plurality of respondents are classified based on combinations of the answers, the importance being a value based on at least one of a gain or a frequency of the features; A step of calculating the importance of the plurality of questions with respect to classification performed by a first classification model created by supervised learning using first learning data based on the contents of answers of a plurality of respondents to a questionnaire survey consisting of a plurality of questions and the results of the clustering of each of the plurality of respondents; acquiring second learning data by excluding data related to some questions from the first learning data based on the importance; performing supervised learning using the second training data to generate a second classification model; Run In the step of acquiring the second learning data, a selection input of a target question to be retained in a simple questionnaire survey to be answered by a subject to be classified using the second classification model or a non-target question to be excluded from the simple questionnaire survey is accepted from among the questions referred to in the clustering, and a feature corresponding to the non-target question is excluded to acquire the second learning data; In the step of creating the second classification model, the second classification model is created by learning about a correlation between a combination of the answer contents to the target question and a cluster into which the answerer is classified. method.

11. a means for calculating the importance of features, which are data related to the plurality of questions, for classification of respondents, using a first classification model having a decision tree structure, the features being created by supervised learning using first learning data based on answers of a plurality of respondents to a questionnaire survey consisting of a plurality of questions and respective clustering results obtained by classifying the plurality of respondents based on combinations of the answers, the importance being a value based on at least one of a gain or a frequency of the features; a means for acquiring second learning data by excluding data related to some questions from the first learning data based on the importance; A means for performing supervised learning using the second learning data to create a second classification model; Equipped with the means for acquiring the second learning data accepts a selection input of a target question to be retained in a simple questionnaire survey to be answered by a subject to be classified using the second classification model, or a non-target question to be excluded from the simple questionnaire survey, among the questions referred to in the clustering, and acquires the second learning data by excluding a feature corresponding to the non-target question; The means for creating the second classification model creates the second classification model by learning a correlation between a combination of the contents of the answers to the target questions and a cluster into which the answerer is classified. Information processing device.

12. A system including a plurality of information processing devices, a means for calculating the importance of features, which are data related to the plurality of questions, for classification of respondents, using a first classification model having a decision tree structure, the features being created by supervised learning using first learning data based on answers of a plurality of respondents to a questionnaire survey consisting of a plurality of questions and respective clustering results obtained by classifying the plurality of respondents based on combinations of the answers, the importance being a value based on at least one of a gain or a frequency of the features; a means for acquiring second learning data by excluding data related to some questions from the first learning data based on the importance; A means for performing supervised learning using the second learning data to create a second classification model; Equipped with the means for acquiring the second learning data accepts a selection input of a target question to be retained in a simple questionnaire survey to be answered by a subject to be classified using the second classification model, or a non-target question to be excluded from the simple questionnaire survey, among the questions referred to in the clustering, and acquires the second learning data by excluding a feature corresponding to the non-target question; The means for creating the second classification model creates the second classification model by learning a correlation between a combination of the contents of the answers to the target questions and a cluster into which the answerer is classified. system.

Citation Information

Patent Citations

  • Information processing method and information processing apparatus

    JP2023037406A