Survey result analysis program, survey result analysis method, and information processing device
The survey result analysis program addresses the challenge of inconsistent grouping by generating causal relationship candidates and using optimization techniques to ensure accurate and consistent analysis of survey respondent behaviors.
Patent Information
- Application Number
- JP2021174678
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2041-10-26
AI Technical Summary
Conventional techniques fail to accurately group survey respondents based on their attributes and behaviors, leading to inconsistent analysis of causal relationships, which hinders the understanding of overall trends in group behavior.
A survey result analysis program that generates multiple causal relationship candidates and uses combinatorial optimization to find combinations that minimize or maximize an objective function, ensuring a predetermined percentage of respondents are correctly explained by the causal relationships.
Enables accurate analysis of the causal relationship between attributes and behavior of survey respondents, maintaining consistent groupings and improving the reliability of analysis results.
Smart Images

Figure 0007733300000004 
Figure 0007733300000005 
Figure 0007733300000006
Abstract
Description
[Technical Field]
[0001] The present invention relates to a survey result analysis program, a survey result analysis method, and an information processing device. [Background technology]
[0002] In order to understand the overall trends in the behavior of people with various attributes, surveys are widely conducted and data analysis is performed using the responses of the survey respondents. For example, a large number of people are asked to fill out a survey about their attributes and food waste behavior trends, and the data from the responses is then analyzed using a computer. This makes it possible to find causal relationships between the attributes of respondent groups and the types of food waste behavior they exhibit.
[0003] In terms of data analysis technologies, a technique using decision tree analysis to analyze ticket satisfaction has been proposed. It has also been proposed to use decision tree analysis to generate classifiers in a adoption forecasting system that predicts when each customer will begin using a product or service. A technology that uses data responses to multiple questions about each project has also been proposed for a computer system that detects signs of risk in projects. Furthermore, it has been proposed to use customer attribute data, including customer survey information, in an information processing device that extracts variables that explain predictions made using deep learning. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] US Patent Application Publication No. 2017 / 308903 [Patent Document 2] Japanese Patent Application Laid-Open No. 2009-238193 [Patent Document 3] Japanese Patent Application Laid-Open No. 2010-108404 [Patent Document 4] International Publication No. 2018 / 142753 Summary of the Invention [Problem to be solved by the invention]
[0005] When analyzing the causal relationship between the attributes and behavior of the overall tendency of a group of survey respondents using survey results, it is possible to group respondents with similar attributes and analyze each group. If each respondent belongs to a group and the causal relationship between the attributes and behavior of the respondents in each group can be found, the causal relationship between the attributes and behavior of the respondents for all respondents can be understood.
[0006] However, conventional techniques have not been able to group respondents in a way that can properly explain the causal relationship between their attributes and behavior for the entire group of respondents (whole group or the majority group).As a result, conventional techniques perform analysis based on inappropriate groupings, making it difficult to correctly analyze the causal relationship between the attributes and behavior of survey respondents.
[0007] In one aspect, the present invention aims to enable a correct analysis of the causal relationship between the attributes and behavior of survey respondents. [Means for solving the problem]
[0008] In one proposal, a survey result analysis program is provided that causes a computer to perform the following processes. The computer generates a plurality of causal relationship candidates, each including a pair of a first answer candidate to at least some of the one or more first questions and a second answer candidate to one of the one or more second questions, based on survey result data indicating the answers of each of a plurality of respondents to a survey including one or more first questions regarding the attributes of the respondent and one or more second questions regarding the behavior of the respondent.The computer then searches, based on the survey result data, for a solution to a combinatorial optimization problem that minimizes or maximizes the value of an objective function, the value of which varies depending on the causal relationship candidates to be combined, under a first constraint that a predetermined percentage or more of the plurality of respondents give the same answer as any pair of the first answer candidate and the second answer candidate in the causal relationship candidates to be combined. [Effects of the Invention]
[0009] According to one aspect, it is possible to correctly analyze the causal relationship between the attributes and behavior of a survey respondent. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 2 is a diagram illustrating an example of a method for analyzing survey results according to the first embodiment. [Figure 2] FIG. 2 illustrates an example of computer hardware. [Figure 3] FIG. 1 is a block diagram showing the functions of a computer for data analysis. [Figure 4] FIG. 10 is a diagram showing an example of questionnaire result data. [Figure 5] FIG. 10 is a diagram illustrating an example of processing in a data analysis unit. [Figure 6] FIG. 10 is a diagram illustrating an example of analysis result data. [Figure 7] 10 is a flowchart illustrating an example of a procedure for a data analysis process. [Figure 8] 10 is a flowchart showing a first example of a causal relationship candidate generation process. [Figure 9] FIG. 10 is a diagram illustrating an example of generated causal relationship candidates. [Figure 10] 10 is a flowchart showing a second example of the causal relationship candidate generating process. [Figure 11] FIG. 10 is a diagram illustrating an example of generating candidates for causal relationships using decision tree analysis. [Figure 12] 10 is a flowchart illustrating an example of a procedure for combinatorial optimization processing that prioritizes a small number of causal relationships. [Figure 13] FIG. 10 is a schematic diagram of a procedure for determining candidates for causal relationships to be retained in the analysis results obtained by combinatorial optimization processing that prioritizes a small number of causal relationships. [Figure 14] 10 is a flowchart illustrating an example of a procedure for combinatorial optimization processing that prioritizes accuracy of causal relationships. [Figure 15] FIG. 10 is a schematic diagram of a procedure for determining candidates for causal relationships to be retained in the analysis results obtained by combinatorial optimization processing that prioritizes the accuracy of causal relationships. DETAILED DESCRIPTION OF THE INVENTION
[0011] The present embodiment will be described below with reference to the drawings. Note that each embodiment can be implemented in combination with a plurality of other embodiments within a range that does not contradict each other. [First embodiment] First, a first embodiment will be described. The first embodiment is a survey result analysis method that generates a large number of candidates for causal relationships between the attributes and behaviors of respondents to a survey, and searches for an optimal combination of causal relationships using an objective function according to the purpose of the analysis while covering all respondents. The purpose of the analysis is, for example, to ensure that the causal relationships between the attributes and behaviors of the respondents are expressed as accurately as possible by the combined candidates for causal relationships, or to combine as few candidates for causal relationships as possible.
[0012] Fig. 1 is a diagram showing an example of a survey result analysis method according to a first embodiment. Fig. 1 shows an information processing device 10 that implements the survey result analysis method. The information processing device 10 can implement the survey result analysis method by, for example, executing a survey result analysis program.
[0013] The information processing device 10 includes a storage unit 11 and a processing unit 12. The storage unit 11 is, for example, a memory or a storage device included in the information processing device 10. The processing unit 12 is, for example, a processor or an arithmetic circuit included in the information processing device 10.
[0014] The storage unit 11 stores survey result data 11a. The survey result data 11a shows the answers of each of a plurality of respondents to a survey that includes one or more first questions about the attributes of the respondent and one or more second questions about the behavior of the respondent. For example, the survey result data 11a shows that the respondent with respondent number "1" answered "1" to the question "How many people are there in your household?", answered "No" to the question "Do you cook for yourself?", and answered "Not much" to the question "Do you leave a lot of food uneaten?"
[0015] The processing unit 12 performs data analysis on the questionnaire result data 11a and analyzes the causal relationships between the attributes and behaviors of a group of respondents. Here, the difficulty of determining the causal relationships that cover all respondents will be explained.
[0016] In data analysis, for example, cluster analysis is used to group multiple respondents who give similar answers. Decision tree analysis is then used to analyze the explanations for the groups generated by cluster analysis (what attributes the respondent groups have). In decision tree analysis, a decision tree is generated that branches out according to the answers to questions (respondent attributes), and each node in the decision tree contains respondents who gave the same answers up to that node. Because groups of respondents belonging to the same node share common attributes, by referring to the decision tree, it is possible to explain the attributes that groups of respondents belonging to each node have and what behaviors they tend to exhibit.
[0017] For example, if multiple respondents belonging to a certain group have the same food waste behavior, it is possible that the group's description (attributes such as whether they live alone or cook at home) is the cause of their food waste behavior. In other words, it is possible to understand the causal relationship between the respondent's attributes and their food waste behavior.
[0018] However, when data analysis is performed by combining cluster analysis and decision tree analysis, the groupings analyzed using cluster analysis may not match the groups (explained objects) explained by decision tree analysis. For example, a group of respondents who are the explained objects belonging to a node in the decision tree may belong to a different group in cluster analysis. If the groups obtained by cluster analysis do not match the explained objects in decision tree analysis, the results of the decision tree analysis cannot be used to explain the groups generated by cluster analysis.
[0019] Although it is possible to achieve consistent groupings by dividing the data into detailed groups in the decision tree analysis, if the decision tree analysis becomes too detailed, the groups generated by the cluster analysis will contain multiple groups with different causal relationships, which will reduce the reliability of the causal relationships of the groups generated by the cluster analysis.
[0020] Furthermore, even if cluster analysis alone is used, it is possible to analyze the co-occurrence trends of responses to survey questions by group, but the co-occurrence trends do not necessarily indicate causal relationships.
[0021] As described above, in the past, there was a possibility that the grouping would be inconsistent and useful analysis results would not be obtained. Conversely, if analysis could be performed while maintaining the consistency of the grouping, it would be possible to prevent failure in the analysis of causal relationships due to inconsistency in the grouping. Therefore, the processing unit 12 generates a large number of candidates for causal relationships (groups of respondents that can be explained by the causal relationships) and, from among them, finds a combination of candidates for causal relationships that can explain all respondents. This allows the analysis of causal relationships and grouping to be performed as an integrated process, making it possible to prevent inconsistency in the grouping.
[0022] For example, the processing unit 12 generates a plurality of causal relationship candidates 1a, 1b,... based on the survey result data 11a. Each of the causal relationship candidates 1a, 1b,... includes a pair of a first answer candidate to at least some of the one or more first questions and a second answer candidate to one of the one or more second questions. For example, the first answer candidate of the causal relationship candidate 1a is "How many people are there in your household? = 1 person," and the second answer candidate is "Do you have a lot of leftovers? = Not many."
[0023] Each of the causal relationship candidates 1a, 1b,... represents a group candidate that includes one or more respondents. In other words, the set of respondents who gave the same answer as the pair of the first and second answer candidates for each of the causal relationship candidates 1a, 1b,... constitutes the group corresponding to that causal relationship. For example, a candidate group corresponding to causal relationship candidate 1a would include a respondent who answered "1" to the question "How many people are in your household?" and "Not many" to the question "Do you have a lot of leftovers?" Respondents who belong to a group corresponding to one of the causal relationship candidates 1a, 1b,... can be called respondents who can be correctly explained by that causal relationship candidate.
[0024] Next, the processing unit 12 searches for a solution to the combinatorial optimization problem that minimizes or maximizes the value of an objective function whose value changes depending on the candidate causal relationships to be combined, based on the survey result data 11a. As a constraint condition for the combinatorial optimization problem at this time, a first constraint condition is used, which requires that a predetermined percentage or more of the multiple respondents give the same answer as a pair of a first answer candidate and a second answer candidate for one of the candidate causal relationships to be combined. The first constraint condition means that a predetermined percentage or more of the multiple respondents can be correctly explained by one of the candidate causal relationships to be combined. If the condition of a predetermined percentage or more of the respondents applies to all respondents, then all respondents can be explained by the candidate causal relationships to be combined.
[0025] The solution to the combinatorial optimization problem obtained in this way indicates candidate causal relationships to be combined. The existence of the first constraint guarantees that the combination of candidate causal relationships indicated in the solution will correctly explain at least a specified percentage of the multiple respondents. Furthermore, each of the multiple candidate causal relationships 1a, 1b, ... corresponds to a group of respondents who can be correctly explained by that candidate causal relationship. In other words, the causal relationships and the groups are associated in advance, and their consistency is maintained. Therefore, it is possible to correctly analyze the causal relationships between the attributes and behavior of survey respondents.
[0026] The processing unit 12 uses an appropriate objective function according to the purpose of the analysis as the objective function of the combinatorial optimization problem. For example, when the objective is to identify as few causal relationships as possible that can explain all respondents, the processing unit 12 searches for a solution to the first combinatorial optimization problem that minimizes the number of candidate causal relationships to be combined. This makes it possible to obtain as few causal relationships as possible that can explain all respondents (or at least a predetermined percentage).
[0027] In other cases, the purpose is to identify causal relationships as accurately as possible that can explain all respondents. In this case, the processing unit 12 searches for a solution to the second combinatorial optimization problem that maximizes the accuracy of the causal relationship between the attribute indicated by the first answer candidate and the behavior indicated by the second answer candidate for each of the causal relationship candidates to be combined. This allows the most accurate causal relationship that can explain all respondents (or at least a predetermined percentage) to be obtained.
[0028] The processing unit 12 may search both the first and second combinatorial optimization problems. In this case, the processing unit 12 first solves the combinatorial optimization problem corresponding to the objective with the higher importance, and then uses that solution to solve the other combinatorial optimization problem.
[0029] For example, when it is more important to reduce the number of causal relationships obtained as a solution, the processing unit 12 first searches for a solution to the first combinatorial optimization problem, and then searches for a solution to the second combinatorial optimization problem using the solution to the first combinatorial optimization problem. For example, in addition to the first constraint, the processing unit 12 uses a second constraint that the number of candidate causal relationships to be combined is equal to the number of candidate causal relationships included in the combination obtained as the solution to the first combinatorial optimization problem as a constraint for the search for a solution to the second combinatorial optimization problem. By adding the second constraint, when there are multiple combination patterns of the minimum number of candidate causal relationships that can explain all respondents (or at least a predetermined percentage), it is possible to obtain the most accurate combination pattern among them as a solution.
[0030] Furthermore, when maximizing the accuracy of the causal relationships obtained as a solution is more important, the processing unit 12 first searches for a solution to the second combinatorial optimization problem, and then searches for a solution to the first combinatorial optimization problem using the solution to the second combinatorial optimization problem. For example, the processing unit 12 uses the first constraint and the third constraint as constraints for the search for a solution to the first combinatorial optimization problem. The third constraint is a condition that the accuracy of each candidate causal relationship to be combined must be equal to the accuracy of each candidate causal relationship included in the combination obtained as the solution to the second combinatorial optimization problem. The accuracy of the candidate causal relationship refers to the accuracy of the causal relationship between the attribute and behavior indicated by the candidate causal relationship. By adding the third constraint, when there are multiple combination patterns of candidate causal relationships that are most accurate and can explain all respondents (or a predetermined percentage or more), a combination pattern that combines the fewest candidate causal relationships among them can be obtained as a solution.
[0031] The search for a solution to the second combinatorial optimization problem is a process of finding a combination of candidate causal relationships that maximizes the accuracy of the causal relationship between the attributes and actions indicated in each candidate causal relationship to be combined. The accuracy of the causal relationship can be expressed numerically using the following calculation.
[0032] For example, the processing unit 12 counts the number of respondents who cannot be explained by the candidate causal relationship for each candidate causal relationship to be combined. The number of respondents who cannot be explained by the candidate causal relationship is a subtraction value obtained by subtracting the number of second respondents who gave the same answer as the second answer candidate indicated in the candidate causal relationship from the number of first respondents who gave the same answer as the first answer candidate indicated in the candidate causal relationship. The processing unit 12 then uses the sum of the subtraction values (the number of respondents who cannot be explained by the candidate causal relationship) for each candidate causal relationship to be combined as an index of the accuracy of the causal relationship. In this case, minimizing the sum results in maximizing the accuracy of the causal relationship.
[0033] By expressing the accuracy of the combination of candidate causal relationships in this way using the total number of respondents who cannot be explained by the candidate causal relationships, the fewer the number of candidate causal relationships included in the solution, the more value is added to the total value, reducing the accuracy. Conversely, the more candidate causal relationships included in the solution, the higher the accuracy. Therefore, using this method of calculating accuracy prevents the number of candidate causal relationships included in the solution from becoming too small when searching for a solution to the second optimization problem.
[0034] Even if the generated causal relationship candidates 1a, 1b,... have low accuracy, the number of respondents who can be explained by the causal relationship candidate is small, and the candidate is unlikely to be included in the solution to the first optimization problem. Therefore, the processing unit 12 can generate as many causal relationship candidates 1a, 1b,... as possible, including those with low accuracy. However, if the number of causal relationship candidates 1a, 1b,... is too large, the amount of calculation required to solve the combinatorial optimization problem becomes enormous. Therefore, the processing unit 12 generates causal relationship candidates that are as accurate as possible.
[0035] For example, the processing unit 12 selects at least a portion of one or more first questions and one of one or more second questions, and selects one first answer candidate for each of the selected first questions. The processing unit 12 also selects a second answer candidate (e.g., the answer given by the most respondents) for the selected second question based on answers to the selected second question from respondents who gave the same answer as the selected first answer candidate. The processing unit 12 then generates candidates for causal relationships including the selected first answer candidate and the selected second answer candidate.
[0036] This makes it possible to reduce the number of candidate causal relationships 1a, 1b,... without reducing the accuracy of the combinations of candidate causal relationships included in the solution to the combinatorial optimization problem. Reducing the number of candidate causal relationships 1a, 1b,... also shortens the time required to find a solution to the combinatorial optimization problem.
[0037] Second Embodiment Next, a second embodiment will be described. In the second embodiment, the results of a questionnaire survey are used to analyze, by a computer, the causal relationship between the attributes of the respondents to the questionnaire and their food waste behavior.
[0038] 2 is a diagram illustrating an example of computer hardware. The entire computer 100 is controlled by a processor 101. A memory 102 and multiple peripheral devices are connected to the processor 101 via a bus 109. The processor 101 may be a multiprocessor. The processor 101 is, for example, a central processing unit (CPU), a micro processing unit (MPU), or a digital signal processor (DSP). At least some of the functions realized by the processor 101 executing a program may be realized by an electronic circuit such as an application specific integrated circuit (ASIC) or a programmable logic device (PLD).
[0039] The memory 102 is used as a main storage device of the computer 100. The memory 102 temporarily stores at least a part of the OS (Operating System) programs and application programs to be executed by the processor 101. The memory 102 also stores various data used in processing by the processor 101. As the memory 102, for example, a volatile semiconductor storage device such as a RAM (Random Access Memory) is used.
[0040] The peripheral devices connected to the bus 109 include a storage device 103, a GPU (Graphics Processing Unit) 104, an input interface 105, an optical drive device 106, a device connection interface 107, and a network interface 108.
[0041] The storage device 103 electrically or magnetically writes and reads data to and from a built-in recording medium. The storage device 103 is used as an auxiliary storage device for the computer 100. The storage device 103 stores an OS program, application programs, and various data. Note that the storage device 103 may be, for example, an HDD (Hard Disk Drive) or an SSD (Solid State Drive).
[0042] The GPU 104 is an arithmetic unit that performs image processing and is also called a graphics controller. The GPU 104 is connected to a monitor 21. The GPU 104 displays an image on the screen of the monitor 21 in accordance with an instruction from the processor 101. The monitor 21 may be a display device using organic EL (Electro Luminescence) or a liquid crystal display device.
[0043] The input interface 105 is connected to a keyboard 22 and a mouse 23. The input interface 105 transmits signals sent from the keyboard 22 and the mouse 23 to the processor 101. The mouse 23 is an example of a pointing device, and other pointing devices can also be used. Examples of other pointing devices include a touch panel, a tablet, a touch pad, and a trackball.
[0044] The optical drive device 106 uses a laser beam or the like to read data recorded on an optical disc 24 or write data to the optical disc 24. The optical disc 24 is a portable recording medium on which data is recorded so that it can be read by reflected light. The optical disc 24 includes a DVD (Digital Versatile Disc), a DVD-RAM, a CD-ROM (Compact Disc Read Only Memory), a CD-R (Recordable) / RW (Rewritable), and the like.
[0045] The device connection interface 107 is a communication interface for connecting peripheral devices to the computer 100. For example, a memory device 25 or a memory reader / writer 26 can be connected to the device connection interface 107. The memory device 25 is a recording medium equipped with a function for communicating with the device connection interface 107. The memory reader / writer 26 is a device for writing data to the memory card 27 or reading data from the memory card 27. The memory card 27 is a card-type recording medium.
[0046] The network interface 108 is connected to the network 20. The network interface 108 transmits and receives data to and from other computers or communication devices via the network 20. The network interface 108 is a wired communication interface connected by a cable to a wired communication device such as a switch or a router. The network interface 108 may also be a wireless communication interface connected by radio waves to a wireless communication device such as a base station or an access point.
[0047] The computer 100 can realize the processing functions of the second embodiment by using the above-described hardware. The device shown in the first embodiment can also be realized by using the same hardware as the computer 100 shown in FIG.
[0048] The computer 100 realizes the processing functions of the second embodiment by executing a program recorded on, for example, a computer-readable recording medium. The program describing the processing to be executed by the computer 100 can be recorded on various recording media. For example, the program to be executed by the computer 100 can be stored in a storage device 103. The processor 101 loads at least a portion of the program in the storage device 103 into the memory 102 and executes the program. The program to be executed by the computer 100 can also be recorded on a portable recording medium such as an optical disk 24, a memory device 25, or a memory card 27. The program stored on the portable recording medium becomes executable after being installed on the storage device 103, for example, under the control of the processor 101. The processor 101 can also read and execute the program directly from the portable recording medium.
[0049] Such a computer 100 can perform data analysis using the survey results. 3 is a block diagram showing the functions of a computer for data analysis. The computer 100 has a memory unit 110 and a data analysis unit 120. The memory unit 110 stores survey result data 111. The survey result data 111 is data indicating answers to a plurality of questions included in the survey for each respondent who answered the survey. The memory unit 110 is realized using, for example, a part of the storage area of the memory 102 or the storage device 103.
[0050] The data analysis unit 120 analyzes the causal relationships between the attributes of respondents and food waste behavior based on the questionnaire result data 111. For example, the data analysis unit 120 prepares multiple candidates for the causal relationships between the attributes of respondents to the questionnaire and food waste behavior. The data analysis unit 120 then calculates and outputs as few combinations of candidate causal relationships as possible as accurately as possible so as to cover all respondents. The data analysis unit 120 can solve the combinations of candidate causal relationships found by calculation as a combinatorial optimization problem.
[0051] 4 is a diagram showing an example of survey result data. In the survey result data 111, the answers of the respondents to the questions included in the survey are set in association with the respondent numbers for identifying the respondents. The questions are divided into questions for the condition part of the candidate causal relationship and questions for the conclusion part of the causal relationship.
[0052] Condition questions are questions about the respondent's attributes. For example, "How many people are in your household?" and "Do you cook for yourself?" are examples of condition questions. In response to the question "How many people are in your household?", the respondent answers the number of people in the household to which the respondent belongs. In response to the question "Do you cook for yourself?", the respondent will answer "Yes" if they are in the habit of cooking for themselves, and "No" if they are not in the habit of cooking for themselves.
[0053] Conclusion questions are questions about the respondent's food waste behavior. Examples of conclusion questions include "Do you leave a lot of food uneaten?" and "Do you often exceed the expiration date?" In response to the question "Do you leave a lot of food uneaten?", respondents will answer "a lot" if they leave a lot of food uneaten from meals, and "a little" if they leave a little uneaten from meals. In response to the question "Do you often exceed the expiration date?", respondents will answer "a lot" if they often let purchased food exceed the expiration date without consuming it, and will answer "a little" if they rarely let purchased food exceed the expiration date without consuming it.
[0054] In the example in Figure 4, only the parts of the answers that change in content, such as "1 person," "no," and "few," are shown, but each answer also contains information about the corresponding question. For example, each answer contains information about the corresponding question, such as "How many people are in your household? = 1 person," "Do you cook at home? = No," and "Do you leave a lot of food? = Few."
[0055] The question for the condition part in the second embodiment is an example of the first question shown in the first embodiment, and the question for the conclusion part in the second embodiment is an example of the second question shown in the first embodiment.
[0056] The results of many respondents answering the questionnaire are stored in the storage unit 110 as questionnaire result data 111. The data analysis unit 120 then performs data analysis based on the questionnaire result data 111.
[0057] FIG. 5 is a diagram showing an example of processing in the data analysis unit. For example, the data analysis unit 120 generates multiple candidate causal relationships based on the survey result data 111. The candidate causal relationships include one or more questions for the condition part and one question for the conclusion part. The data analysis unit 120 determines, from the multiple candidate causal relationships, a combination of causal relationships that can correctly explain all respondents to the survey result data 111, using a solution search method using combinatorial optimization. At this time, the data analysis unit 120 sets constraints for the combinatorial optimization so as to reduce the number of causal relationships to be combined. The data analysis unit 120 then outputs analysis result data 30.
[0058] The analysis result data 30 includes multiple pieces of causal relationship data 31, 32,... that indicate causal relationships that correctly explain the respondents. Each piece of causal relationship data 31, 32,... includes the causal relationship, the number of respondents who fall under the condition part of the causal relationship, the number of respondents who fall under the conclusion part of the causal relationship, the accuracy of the causal relationship, and the numbers of the respondents who fall under the causal relationship.
[0059] The analysis result data 30 also includes the accuracy of the causal combination. The accuracy of the causal combination is expressed, for example, as the total number of respondents who cannot be correctly explained by the causal relationship. In this case, the smaller the value, the more accurate the causal combination.
[0060] For example, the data analysis unit 120 calculates the total number of respondents corresponding to the condition part for each combination of candidate causal relationships remaining in the analysis results. The data analysis unit 120 also calculates the total number of respondents corresponding to the conclusion part for each combination of candidate causal relationships remaining in the analysis results. The data analysis unit 120 then calculates the accuracy of the combination of candidate causal relationships by subtracting the total number of respondents corresponding to the conclusion part from the total number of respondents corresponding to the condition part. The smaller the calculated subtraction value, the more accurate the combination of candidate causal relationships.
[0061] 6 is a diagram showing an example of analysis result data. For example, the condition part of the causal relationship shown in the causal relationship data 31 is answering "1" to the question "How many people in your household?" and answering "No" to the question "Do you cook at home?". Also, the conclusion part of the causal relationship shown in the causal relationship data 31 is answering "Not often" to the question "Do you often exceed the expiration date?"
[0062] The number of respondents corresponding to the condition part of the causal relationship data 31 is "10 people." Of the respondents corresponding to the condition part of the causal relationship data 31, the number of respondents corresponding to the conclusion part is "6 people." The accuracy of each causal relationship is shown, for example, by the percentage of respondents who fall into the conclusion part among those who fall into the condition part. In the case of causal relationship data 31, the accuracy of the causal relationship is "0.6." In this case, the higher the accuracy value of the causal relationship, the more accurate the causal relationship.
[0063] In the causal relationship data 31, the respondent numbers of the respondents corresponding to the causal relationship are "1, 3, 4, . . . ". The causal relationship data 32 also shows the same type of information as the causal relationship data 31. The analysis result data 30 further shows the accuracy of the combination of causal relationships. In the example of Figure 6, the number of respondents who fall into the condition part of causal relationship 1 is 10, the number of respondents who fall into the conclusion part is 6, the number of respondents who fall into the condition part of causal relationship 2 is 20, and the number of respondents who fall into the conclusion part is 16. Therefore, the "accuracy of the combination of candidate causal relationships = 10 + 20 - (6 + 16) = 8". This is the sum of the number of respondents who cannot be explained by the causal relationships shown in the causal relationship data 31 and the number of respondents who cannot be explained by the causal relationships shown in the causal relationship data 32.
[0064] Next, we will explain in detail the data analysis process for analyzing the causal relationship between the attributes of survey respondents and food waste behavior. 7 is a flowchart showing an example of the procedure for data analysis processing. The processing shown in FIG. 7 will be explained below in order of step number.
[0065] [Step S101] The data analysis unit 120 reads the questionnaire result data 111 from the storage unit 110. [Step S102] The data analysis unit 120 performs a causal relationship candidate generation process based on the questionnaire result data 111. The causal relationship candidate generation process will be described in detail later (see FIGS. 8 and 10). The causal relationship candidate generation process generates multiple causal relationship candidates, each of which has answers to one or more questions for the condition part of the questionnaire as condition parts and an answer to one question for the conclusion part of the questionnaire as a conclusion part.
[0066] [Step S103] The data analysis unit 120 executes a combinatorial optimization process to determine whether or not to retain each of the multiple causal relationship candidates in the analysis results. The combinatorial optimization process will be described in detail later (see FIGS. 12 and 14).
[0067] [Step S104] The data analysis unit 120 generates analysis result data 30 in which the candidate causal relationships determined by the combinatorial optimization process to be retained in the analysis results are treated as causal relationships between the respondent's attributes and food waste behavior. The respondent's attributes are indicated in the condition part of the causal relationship. The food waste behavior is indicated in the conclusion part of the causal relationship. The data analysis unit 120 displays the contents of the generated analysis result data 30 on, for example, the monitor 21. The data analysis unit 120 also stores the generated analysis result data 30 in the storage device 103.
[0068] In this way, the causal relationship between respondents' attributes and food waste behavior is analyzed. Next, the causal relationship candidate generation process will be described in detail. 8 is a flowchart showing a first example of the process of generating causal relationship candidates. The process shown in FIG. 8 will be explained below in order of step number.
[0069] [Step S111] The data analysis unit 120 initializes the number i of combinations of questions to be included in the condition part to "1" (i=1). [Step S112] The data analysis unit 120 generates candidates for the condition part based on a combination of i questions for the condition part.
[0070] For example, when i=1, the data analysis unit 120 generates condition part candidates for each answer to each condition part question, such as "How many people in your household?" and "Do you cook at home?". For example, for the question "How many people in your household?", condition part candidates are generated for each household size, such as "How many people in your household? = 1 person" and "How many people in your household? = 2 people." Furthermore, for the question "Do you cook at home?", two condition part candidates are generated, "Do you cook at home? = Yes" and "Do you cook at home? = No."
[0071] When i=2, the data analysis unit 120 generates condition part candidates for each combination of answers to a combination of two questions in the condition part. For example, for the question combination of "How many people in your household?" and "Do you cook for yourself?", the data analysis unit 120 generates condition part candidates for each combination of answers to the questions. For example, the data analysis unit 120 generates condition part candidates such as "How many people in your household?=1 AND do you cook for yourself?=Yes", "How many people in your household?=1 AND do you cook for yourself?=No", "How many people in your household?=2 AND do you cook for yourself?=Yes", and "How many people in your household?=2 AND do you cook for yourself?=No".
[0072] Similarly, when the value of i is 3 or greater, the data analysis unit 120 generates candidates for the condition part for each combination of answers to the questions for each combination of i questions in the condition part. <Step S113> The data analysis unit 120 selects one unselected condition part candidate from among the generated condition part candidates.
[0073] [Step S114] The data analysis unit 120 counts the number of respondents who answered the corresponding question among the selected condition part candidates, for each answer to the question for the conclusion part, based on the survey result data 111. A respondent who answered the selected condition part candidate is a respondent who gave the same answer as the condition part candidate to the question for the selected condition part candidate.
[0074] [Step S115] The data analysis unit 120 determines the answer to the question for the conclusion part that has the most relevant respondents (respondents corresponding to the selected candidate condition part) as the conclusion part.The data analysis unit 120 then generates a candidate causal relationship in which the selected candidate condition part is the condition part of the causal relationship and the determined conclusion part is the conclusion part of the causal relationship.
[0075] [Step S116] The data analysis unit 120 determines whether there are any unselected condition section candidates. If there are any unselected condition section candidates, the data analysis unit 120 proceeds to step S113. If there are no unselected condition section candidates, the data analysis unit 120 proceeds to step S117.
[0076] [Step S117] The data analysis unit 120 determines whether the number of generated causal relationship candidates has reached a predetermined upper limit. If the number of causal relationship candidates has reached the upper limit, the data analysis unit 120 terminates the causal relationship candidate generation process. If the number of causal relationship candidates has not reached the upper limit, the data analysis unit 120 proceeds to step S118.
[0077] [Step S118] The data analysis unit 120 increments the number i of combinations of questions in the condition part by "1" (i=i+1), and proceeds to step S112. In this way, the data analysis unit 120 repeats the generation of candidates for causal relationships while increasing the number of combinations of questions in the condition parts until the number of candidates for causal relationships reaches a predetermined upper limit.
[0078] Fig. 9 is a diagram showing an example of a generated candidate causal relationship. The example in Fig. 9 shows an example of a generated candidate causal relationship when the answer "1 person" to the question "How many people in your household?" is selected as a candidate for the condition part to be processed. In the example in Fig. 9, the number of respondents who fit the candidate for the condition part (respondents who answered "How many people in your household? = 1 person") is 100.
[0079] The data analysis unit 120 generates conclusion candidates for each answer to the question "Do they leave a lot of leftovers?" for the conclusion part ("Do they leave a lot of leftovers? = many" and "Do they leave a lot of leftovers? = few"). The data analysis unit 120 then counts the number of respondents who fall under each of the condition part candidates. The data analysis unit 120 determines the conclusion part candidate with the most respondents as the conclusion part of the candidate causal relationship. In the example of FIG. 9, 40 out of 100 people answered "Do they leave a lot of leftovers? = many," and 60 out of 100 people answered "Do they leave a lot of leftovers? = few." Therefore, the data analysis unit 120 generates a causal relationship candidate 41 whose condition part is "How many people in your household? = 1" and whose conclusion part is "Do they leave a lot of leftovers? = few."
[0080] Furthermore, the data analysis unit 120 generates conclusion candidates for each answer to the question for the conclusion, "Do they often exceed their expiration date?" ("Do they often exceed their expiration date? = many" and "Do they often exceed their expiration date? = few"). In the example of FIG. 9, 70 out of 100 people answered "Do they often exceed their expiration date? = many," and 30 out of 100 people answered "Do they often exceed their expiration date? = few." Therefore, the data analysis unit 120 generates candidate 42 for a causal relationship whose condition part is "How many people in a household? = 1" and whose conclusion part is "Do they often exceed their expiration date? = many."
[0081] This type of causal relationship candidate is generated for each condition part candidate. If the number of combinations of condition part questions is two, the combination of the two condition part questions becomes a condition part candidate. For example, suppose that the number of respondents who fit the condition part candidate "How many people in your household? = 4 AND Do you cook at home? = Yes" is 40. In this case, suppose that the number of respondents who fit the conclusion part candidate "Does it often exceed its expiration date? = Many" is 30, and the number of respondents who fit the conclusion part candidate "Does it often exceed its expiration date? = Few" is 10. In this case, the conclusion part of the causal relationship candidate whose condition part is "How many people in your household? = 4 AND Do you cook at home? = Yes" is "Does it often exceed its expiration date? = Many."
[0082] The information set in the condition part of the candidate causal relationship in the second embodiment is an example of the first answer candidate in the first embodiment. Also, the information set in the conclusion part of the candidate causal relationship in the second embodiment is an example of the second answer candidate in the first embodiment.
[0083] In the examples shown in Figures 8 and 9, the number of questions included in the causal relationship candidates is increased, and condition part candidates are generated comprehensively for each combination of questions. However, it is also possible to generate causal relationship candidates efficiently using decision tree analysis.
[0084] 10 is a flowchart showing a second example of the causal relationship candidate generation process. The process shown in FIG. 10 will be explained below in order of step number. [Step S131] The data analysis unit 120 selects one question for the conclusion from among unselected questions for the conclusion.
[0085] [Step S132] The data analysis unit 120 generates a decision tree corresponding to the selected question for the conclusion part. Each node of the decision tree other than the leaves is associated with a question for the condition part. Upper nodes are connected to their lower nodes by branches corresponding to the answers to the questions at the upper nodes. The data analysis unit 120 then follows the branches from the root node according to the answers to the questions for the condition part for each answerer. For each node of the decision tree, the data analysis unit 120 then sets, for each answer to the selected question for the conclusion part, the number of answerers who gave that answer among the answerers who reached the corresponding node.
[0086] [Step S133] The data analysis unit 120 determines whether there are any unselected questions for the conclusion. If there are any unselected questions for the conclusion, the data analysis unit 120 proceeds to step S131. If there are no unselected questions for the conclusion, the data analysis unit 120 proceeds to step S134.
[0087] [Step S134] The data analysis unit 120 generates candidates for causal relationships corresponding to nodes other than the root node of each decision tree. The condition part of the generated candidate causal relationship is the logical product of answers to one or more questions for the condition part from the root node of the decision tree to the relevant node. The conclusion part of the generated candidate causal relationship is set to the answer given by the most respondents to the question for the conclusion part corresponding to the decision tree.
[0088] FIG. 11 is a diagram illustrating an example of generating candidates for causal relationships using decision tree analysis. FIG. 11 shows a decision tree 50 corresponding to the question for the conclusion part, "Do you have a lot of leftovers?". The decision tree 50 is an example in which there are two questions for the condition part, "How many people in your household?" and "Do you cook at home?". The question for the condition part, "How many people in your household?", is associated with node 51, which is the root node of the decision tree 50. The question for the condition part, "Do you cook at home?", is associated with nodes 52 and 53 below node 52. The branch from node 51 to node 52 is associated with the answer "1 person" to "How many people in your household?". The branch from node 51 to node 53 is associated with the answer "2 or more people" (including answers such as 2, 3, ...) to "How many people in your household?"
[0089] Each of nodes 52 to 57 other than the root node has the number of respondents who reach that node for each answer to the conclusion question "Do you have a lot of leftovers?" For example, the respondent who reaches node 52 is the respondent who answered "1" to the condition question "How many people are in your household?" In the example of FIG. 11, there are 100 respondents who reach node 52. Of these, 40 respondents answered "a lot" and 60 answered "not many" to the conclusion question "Do you have a lot of leftovers?"
[0090] The data analysis unit 120 generates candidates for causal relationships corresponding to each of the nodes 52 to 57 other than the root node. The condition part of the generated candidate for causal relationship is set to a common answer to the question for the condition part by respondents who reach the corresponding node. Furthermore, the conclusion part of the generated candidate for causal relationship is set to the most common answer by respondents who reach each of the nodes 52 to 57 to the question for the conclusion part corresponding to the decision tree 50, "Is there a lot of leftover food?"
[0091] For example, the condition part of candidate causal relationship 61 corresponding to node 52 is set to "How many people in the household? = 1 person." Of the 100 respondents who reach node 52, 60 answer "few" to the question for the conclusion part, "Are there many food leftovers?" Therefore, the conclusion part of candidate causal relationship 61 is set to "Are there many food leftovers? = Few."
[0092] Additionally, the condition part of candidate causal relationship 62 corresponding to node 53 is set to "How many people in a household? = 2 or more." Of the 140 respondents who reach node 53, 80 answer "a lot" to the question for the conclusion part, "Do you have a lot of leftovers?". The conclusion part of candidate causal relationship 62 is set to "Do you have a lot of leftovers? = A lot."
[0093] Similarly, candidate causal relationships corresponding to other nodes 54 to 57 of the decision tree 50 are also generated. For example, the candidate causal relationships corresponding to node 54 are the condition part "How many people in a household? = 1 AND Do you cook at home? = Yes" (number of applicable respondents: 50) and the conclusion part "Do you have a lot of leftovers? = Low" (number of applicable respondents: 40). The candidate causal relationships corresponding to node 55 are the condition part "How many people in a household? = 1 AND Do you cook at home? = No" (number of applicable respondents: 50) and the conclusion part "Do you have a lot of leftovers? = High" (number of applicable respondents: 30). The candidate causal relationships corresponding to node 56 are the condition part "How many people in a household? = 2 or more AND Do you cook at home? = Yes" (number of applicable respondents: 70) and the conclusion part "Do you have a lot of leftovers? = High" (number of applicable respondents: 50). The candidates for the causal relationship corresponding to node 57 are the condition part "How many people are in your household? = 2 or more AND Do you cook at home? = No" (number of applicable respondents: 70) and the conclusion part "Do you leave a lot of food uneaten? = Not much" (number of applicable respondents: 40).
[0094] Furthermore, for the nodes of the decision tree generated in response to other questions for the decision section (for example, "Does the expiration date often pass?"), corresponding candidates for causal relationships are generated. Once the causal relationship candidates are generated, the data analysis unit 120 uses a combinatorial optimization technique to calculate a combination of causal relationship candidates that can accurately explain all respondents with as few causal relationships as possible. The process for optimizing the combination of causal relationship candidates differs depending on whether priority is given to a small number of causal relationships or to the accuracy of the causal relationships.
[0095] 12 is a flowchart showing an example of a procedure for combinatorial optimization processing that prioritizes a small number of causal relationships. The processing shown in FIG. 12 will be described below in order of step number. [Step S141] The data analysis unit 120 sets the arrays "A(i,k),B(k),C(k),D(i)" to be used in the combinatorial optimization calculation. i is the respondent number. k is the number of the candidate causal relationship.
[0096] A(i, k) is an array indicating respondents who can be correctly explained by the candidate causal relationships. If the kth candidate causal relationship can correctly explain the i-th respondent, the data analysis unit 120 sets the value of A(i, k) to "1." If the kth candidate causal relationship cannot correctly explain the i-th respondent, the data analysis unit 120 sets the value of A(i, k) to "0."
[0097] B(k) is an array indicating the number of respondents whose answers cannot be correctly explained by the candidate causal relationships. The data analysis unit 120 sets, for example, the number of respondents who fall under the condition part but not the conclusion part of the k-th candidate causal relationship to B(k) as the number of respondents whose answers cannot be correctly explained by the candidate causal relationships.
[0098] C(k) is an array indicating candidates for causal relationships to be retained in the analysis results. The data analysis unit 120 sets C(k) to "1" if the kth candidate for causal relationships is to be retained in the analysis results, and sets C(k) to "0" if the kth candidate for causal relationships is not to be retained in the analysis results.
[0099] D(i) is an array indicating whether the respondent can be correctly explained by any of the candidate causal relationships remaining in the analysis results. The data analysis unit 120 sets D(i)≧1 if the i-th respondent can be correctly explained by at least one of the candidate causal relationships remaining in the analysis results. For example, D(i) can be expressed by the following formula.
[0100]
number
[0101] [Step S142] The data analysis unit 120 calculates the variable “N s ,N r ,N f ,N t ". N s is a variable indicating the total number of respondents. The data analysis unit 120 calculates the total number of respondents shown in the questionnaire result data 111 as N s Set to Nr is the total number of causal relationship candidates. The data analysis unit 120 calculates the total number of causal relationship candidates generated in the causal relationship candidate generation process as N r Set to N f is the total number of candidates for causal relationships to be retained in the analysis results. f The value of is expressed by the following formula:
[0102]
number
[0103] N t is a variable that indicates the total number of respondents who cannot be explained by the candidate causal relationships that remain in the analysis results. t The value of is expressed by the following formula:
[0104]
number
[0105] [Step S143] The data analysis unit 120 calculates the minimum number of candidates for causal relationships (N f,minimized ) is calculated by combinatorial optimization. The objective function to be minimized in the combinatorial optimization problem is N f The constraint is "D(i) ≥ 1" (i = 1, 2, , N s )
[0106] [Step S144] When the number of causal relationships is the smallest, the data analysis unit 120 performs combinatorial optimization to find a combination of candidate causal relationships that can most accurately explain all respondents. The objective function to be minimized in this combinatorial optimization problem is N t The constraint is "D(i) ≥ 1" (i = 1, 2, , N s ), and "N f =N f,minimized As a result of this combinatorial optimization calculation, the candidate causal relationship for which "C(k)=1" is obtained becomes the causal relationship to be output as the analysis result.
[0107] In this way, the combination that can most accurately represent the causal relationship among the minimum number of combinations of candidate causal relationships that can correctly explain all of the respondents can be determined as the analysis result.
[0108] Figure 13 is a schematic diagram of the procedure for determining candidates for causal relationships to be left in the analysis results by combinatorial optimization processing that prioritizes the small number of causal relationships. As shown in Figure 13, in the first stage of combinatorial optimization problem 71, all respondents (i=1, 2, . . . , N s ) under the constraint that "D(i) ≥ 1", the total number of candidates for causal relationships to be retained in the analysis results is N f The combination of candidate causal relationships that minimizes D(i) is calculated. The constraint "D(i) ≥ 1" ensures that all respondents are correctly explained. By solving this combinatorial optimization problem 71, the minimum number of candidate causal relationships that can explain all respondents (N f,minimized ) is obtained.
[0109] In the second stage of combinatorial optimization problem 72, the minimum number of candidates for causality (N f,minimized ) is the total number of candidates for causal relationships that remain in the analysis results, N f Under the constraint that t In this case, the combination of candidate causal relationships that minimizes the s ) is constrained to be "D(i) ≥ 1", which ensures that the candidates for causal relationships remaining in the analysis correctly describe all respondents.
[0110] In this way, "N f =N f,minimized " By solving combinatorial optimization problem 72 under the constraint, only the combination pattern that combines the minimum number of candidate causal relationships that can correctly explain all of the respondents can be the solution to combinatorial optimization problem 72. If there are multiple such combination patterns, the combination pattern that combines the candidate causal relationships that most accurately express the causal relationships among those combination patterns is obtained as the solution.
[0111] Fig. 14 is a flowchart showing an example of the procedure for combinatorial optimization processing that prioritizes the accuracy of causal relationships. The processing of steps S151 and S152 shown in Fig. 14 is the same as the processing of steps S141 and S142 shown in Fig. 12. Below, the processing of steps S153 and S154 that differ from the processing of Fig. 12 will be explained in order of step number.
[0112] [Step S153] The data analysis unit 120 selects a combination of candidate causal relationships (N t,minimized ) is calculated by combinatorial optimization. The objective function to be minimized in the combinatorial optimization problem is N t The constraint is "D(i) ≥ 1" (i = 1, 2, , N s )
[0113] [Step S154] The data analysis unit 120 performs combinatorial optimization to find a combination of candidate causal relationships that can explain all respondents with the fewest number of causal relationships when accuracy is optimal. The objective function to be minimized in this combinatorial optimization problem is N f The constraint is "D(i) ≥ 1" (i = 1, 2, , N s ), and "N t =N t,minimized As a result of this combinatorial optimization calculation, the candidate causal relationship for which "C(k)=1" is obtained becomes the causal relationship to be output as the analysis result.
[0114] In this way, the analysis result can be the smallest number of combinations of candidate causal relationships that can correctly explain all respondents, out of the combinations of candidate causal relationships that can most accurately represent the causal relationships.
[0115] Figure 15 is a schematic diagram of the procedure for determining candidates for causal relationships to be retained in the analysis results of combinatorial optimization processing that prioritizes the accuracy of causal relationships. As shown in Figure 15, in the first stage of combinatorial optimization problem 73, all respondents (i=1, 2,...N s) is constrained to be "D(i) ≥ 1". Under this constraint, the total number of respondents N that cannot be accurately explained by the candidate causal relationships remaining in the analysis results is t The combination of candidate causal relationships that minimizes D(i) is calculated. The constraint "D(i) ≥ 1" ensures that all respondents are correctly explained. By solving this combinatorial optimization problem 73, the minimum value of the total number of respondents who cannot be accurately explained (N t,minimized ) is obtained.
[0116] In the second stage of combinatorial optimization problem 74, the minimum total number of respondents who cannot be accurately explained (N t,minimized ) is the total number of candidates for causal relationships that remain in the analysis results, N t Under the constraint that f In this case, the combination of candidate causal relationships that minimizes the s ) is constrained to be "D(i) ≥ 1", which ensures that the candidates for causal relationships remaining in the analysis correctly describe all respondents.
[0117] In this way, "N t =N t,minimized By solving combinatorial optimization problem 74 under the constraint ", only the combination pattern of the most accurate combination of candidate causal relationships can be the solution to combinatorial optimization problem 74. When there are multiple such combination patterns, the combination pattern of the combination that has the smallest total number of causal relationships among those combination patterns is obtained as the solution.
[0118] Each of the one or more causal relationships obtained as a solution can be considered to represent a group of respondents that can be explained by that causal relationship. Therefore, grouping is performed by the process of generating candidate causal relationships. The data analysis unit 120 determines the causal relationship to be used as the analysis result from among the multiple candidate causal relationships that represent the group, and consistency in the grouping is maintained.
[0119] Furthermore, in combinatorial optimization problems, the constraint "D(i) ≥ 1" ensures that each respondent is included in one of the groups (a group of respondents that can be explained by the causal relationships included in the analysis results). Note that there are multiple candidate causal relationships that can explain a single respondent. Therefore, a single respondent is allowed to belong to multiple groups.
[0120] Furthermore, the analysis result data 30 indicates the accuracy of each causal relationship as well as the accuracy of the combination of causal relationships. The accuracy of the combination of causal relationships is expressed, for example, as the total number of respondents who cannot be explained by each of the candidate causal relationships remaining in the analysis results. In this case, the greater the number of candidate causal relationships remaining in the analysis results, the greater the value of the accuracy of the combination of causal relationships. In other words, the accuracy of the combination of causal relationships reflects the accuracy of each causal relationship and the small number of candidate causal relationships remaining in the analysis results. This allows for a comprehensive evaluation of the accuracy of the causal relationships and the small number of groups generated (which is the same as the number of candidate causal relationships remaining in the analysis results).
[0121] Other Embodiments In the second embodiment, the causal relationship between the attributes of respondents to the questionnaire and food waste behavior is analyzed, but the data analysis process shown in the second embodiment can be applied to various analyses other than food waste behavior.
[0122] In the second embodiment, respondents are included in one of the groups, but it is also possible to allow a certain percentage of respondents not to be included in any group. In this case, instead of the constraint that "D(i) ≥ 1" is satisfied for all i, a constraint that "D(i) ≥ 1" is satisfied for a predetermined percentage or more (for example, 90% or more) of i is used.
[0123] Furthermore, by converting a combinatorial optimization problem into an Ising model, it is also possible to solve it using an Ising machine. An Ising machine is a computer specialized for optimization problems using the Ising model, which is one of the magnetic material models in physics. Using an Ising machine makes it possible to efficiently search for solutions to combinatorial optimization problems. Ising machines include quantum annealing machines that use superconducting qubits, coherent Ising machines that use the properties of light as artificial spin, and machines that solve combinatorial optimization problems using digital circuits inspired by quantum phenomena.
[0124] Although the embodiments have been described above, the configuration of each part shown in the embodiments can be replaced with other parts having similar functions. Also, any other components or processes may be added. Furthermore, any two or more configurations (features) of the above-described embodiments may be combined. [Explanation of symbols]
[0125] 1a, 1b,... Candidates for causal relationships 10. Information processing equipment 11 Storage section 11a Survey result data 12 Processing section
Claims
1. Based on survey result data showing the answers of each of a plurality of respondents to a survey including one or more first questions regarding the attributes of the respondent and one or more second questions regarding the behavior of the respondent, a plurality of candidate causal relationships are generated, each candidate causal relationship including a pair of a first candidate answer to at least some of the one or more first questions and a second candidate answer to one of the one or more second questions; determining, for each of the plurality of respondents, whether or not the respondent has given the same answer as the pair of the first answer candidate and the second answer candidate in the candidate causal relationship based on the survey result data; a search is made for a combination of the candidate causal relationships that minimizes the number of candidates for the causal relationships to be combined, among combinations of the candidate causal relationships that satisfy a first constraint that a predetermined percentage or more of the plurality of respondents give the same answer as a pair of the first answer candidate and the second answer candidate of any of the candidates for the causal relationships to be combined; A survey result analysis program that executes processing on a computer.
2. The computer is further caused to execute a process of calculating a subtraction value obtained by subtracting the number of second respondents among the first respondents who gave the same answer as the second answer candidate indicated in the candidate for causal relationship from the number of first respondents who gave the same answer as the first answer candidate indicated in the candidate for causal relationship, based on the survey result data; In the process of searching for a combination of candidate causal relationships, searching for a solution to a first combinatorial optimization problem that finds a combination of the candidate causal relationships that minimizes the number of the candidate causal relationships to be combined, among combinations of the candidate causal relationships that satisfy the first constraint; searching for a solution to a second combinatorial optimization problem that finds a combination of the causal relationship candidates that minimizes the sum of the subtraction values of the causal relationship candidates to be combined, from combinations of the causal relationship candidates that satisfy the first constraint and a second constraint that the number of the causal relationship candidates to be combined is equal to the number of the causal relationship candidates included in the combination obtained as a solution to the first combinatorial optimization problem; 2. The questionnaire result analysis program according to claim 1.
3. The computer is further caused to execute a process of calculating a subtraction value obtained by subtracting the number of second respondents among the first respondents who gave the same answer as the second answer candidate indicated in the candidate for causal relationship from the number of first respondents who gave the same answer as the first answer candidate indicated in the candidate for causal relationship, based on the survey result data; In the process of searching for a combination of candidate causal relationships, searching for a solution to a second combinatorial optimization problem that finds a combination of the candidates for causal relationships that minimizes the sum of the subtraction values of the candidates for causal relationships to be combined, among combinations of the candidates for causal relationships that satisfy the first constraint; searching for a solution to the first combinatorial optimization problem that finds a combination of candidate causal relationships that minimizes the number of candidates for causal relationships to be combined, from among combinations of candidate causal relationships that satisfy the first constraint and a third constraint that the sum of the subtraction values indicated by the candidates for causal relationships to be combined is equal to the sum of the subtraction values indicated by the candidates for causal relationships included in the combination obtained as a solution to the second combinatorial optimization problem; 2. The questionnaire result analysis program according to claim 1.
4. Based on survey result data showing the answers of each of a plurality of respondents to a survey including one or more first questions regarding the attributes of the respondent and one or more second questions regarding the behavior of the respondent, a plurality of candidate causal relationships are generated, each candidate causal relationship including a pair of a first candidate answer to at least some of the one or more first questions and a second candidate answer to one of the one or more second questions; determining, for each of the plurality of respondents, whether or not the respondent has given the same answer as the pair of the first answer candidate and the second answer candidate in the candidate causal relationship based on the survey result data; calculating a subtraction value obtained by subtracting the number of second respondents who gave the same answer as the second answer candidate indicated in the candidate for the causal relationship from the number of first respondents who gave the same answer as the first answer candidate indicated in the candidate for the causal relationship based on the survey result data; a search is made for a combination of the candidate causal relationships that minimizes the sum of the subtraction values of the candidates for causal relationships to be combined, among the combinations of the candidate causal relationships that satisfy a first constraint condition that a predetermined percentage or more of the plurality of respondents give the same answer as a pair of the first answer candidate and the second answer candidate of any of the candidates for causal relationships to be combined; A survey result analysis program that executes processing on a computer.
5. In the process of generating a plurality of candidates for the causal relationship, selecting the first answer candidate for each of at least some of the one or more first questions, selecting one of the one or more second questions, selecting the second answer candidate for the selected second question based on an answer to the selected second question by a respondent who gave the same answer as the selected first answer candidate, and generating the causal relationship candidates including the selected first answer candidate and the selected second answer candidate; 5. The questionnaire result analysis program according to claim 1.
6. Based on survey result data showing the answers of each of a plurality of respondents to a survey including one or more first questions regarding the attributes of the respondent and one or more second questions regarding the behavior of the respondent, a plurality of candidate causal relationships are generated, each candidate causal relationship including a pair of a first candidate answer to at least some of the one or more first questions and a second candidate answer to one of the one or more second questions; determining, for each of the plurality of respondents, whether or not the respondent has given the same answer as the pair of the first answer candidate and the second answer candidate in the candidate causal relationship based on the survey result data; a search is made for a combination of the candidate causal relationships that minimizes the number of candidates for the causal relationships to be combined, among combinations of the candidate causal relationships that satisfy a first constraint that a predetermined percentage or more of the plurality of respondents give the same answer as a pair of the first answer candidate and the second answer candidate of any of the candidates for the causal relationships to be combined; A method for analyzing survey results in which processing is performed by a computer.
7. a processing unit that generates a plurality of causal relationship candidates, each of which includes a pair of a first answer candidate to at least some of the one or more first questions and a second answer candidate to one of the one or more second questions, based on survey result data indicating answers from each of a plurality of respondents to a survey including one or more first questions regarding attributes of the respondent and one or more second questions regarding behavior of the respondent, determines, for each of the causal relationship candidates, based on the survey result data, whether or not each of the plurality of respondents has given the same answer as the pair of the first answer candidate and the second answer candidate in the causal relationship candidates, and searches for a combination of the causal relationship candidates that minimizes the number of the causal relationship candidates to be combined, from among combinations of the causal relationship candidates that satisfy a first constraint condition that a predetermined percentage or more of the plurality of respondents have given the same answer as the pair of the first answer candidate and the second answer candidate in any of the causal relationship candidates to be combined; An information processing device having the above.
Citation Information
Patent Citations
Questionnaire data processing system
JP2006302107A
Circulation prediction system, method and program, and influence degree estimation system, method and program
JP2009238193A
Computer system for performing risk prediction in project and method and computer program thereof
JP2010108404A
Attribute processing apparatus and method
JP2011039760A
Satisfaction metric for customer tickets
US20170308903A1