Group testing device and computer program
Patent Information
- Application Number
- PCT/JP2025/012432
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012432_01102026_PF_FP_ABST
Abstract
Description
Collective testing apparatus and computer program
[0001] One aspect of the present invention relates to a collective testing apparatus and a computer program.
[0002] A technique for comparing a plurality of groups and statistically testing whether there is a bias in data constituting the groups is important. For a plurality of groups, when there is a bias beyond chance in the data constituting the groups, it can be inferred that some cause that generates the bias exists.
[0003] There are a plurality of methods for statistical testing. For example, when the data constituting a group is numerical data, t-test and Wilcoxon rank-sum test are generally used. For example, when the data constituting a group is categorical data belonging to specific categories, chi-square test and Fisher's exact test are generally used.
[0004] As described above, the applicable statistical testing method differs depending on the nature of the problem and the data constituting the groups, but when verifying whether there exists a cause that generates data bias in two groups, the procedure of any method is as follows.
[0005] First, the following null hypothesis and alternative hypothesis are established. Null hypothesis: There is no cause that generates a bias between the two groups (that is, the bias in the actually measured values is caused by chance). Alternative hypothesis: There is a cause that generates a bias between the two groups (that is, the bias in the actually measured values exceeds the range of chance).
[0006] No matter which statistical testing method is used, the hypothesis that the above-mentioned null hypothesis is correct is adopted, and the probability that the data matches the actually measured values is calculated using some model (more specifically, the probability that data more biased than the actually measured values will occur at least under the assumption). This probability is called the p-value (significance probability). Thereafter, the predetermined significance level is compared with the calculated p-value, and if the p-value is lower than the significance level, the null hypothesis is rejected. As the significance level used in this case, 5% and 1% are often used.
[0007] However, when performing the above-mentioned statistical tests on a large number of groups, there is a problem in that the null hypothesis will be rejected with a low probability for some groups, even if there is no cause for bias. For example, if we were to test whether bias exists in some data belonging to each municipality in Japan, with a significance level of 1%, since there are approximately 2200 municipalities in Japan, the null hypothesis would be rejected for about 22 municipalities (even if there is no cause for bias), and it would be inferred that there is a cause for bias specific to those municipalities.
[0008] Japanese Patent No. 6162021
[0009] "Explaining the Chi-squared Test Through Familiar Examples", [online], July 21, 2023, [retrieved March 18, 2025], Internet <URL: https: / / igakunote.com / カイ2乗検定を身近な例で解説する / >"What is Fisher's Exact Test? A Clear Explanation of the Difference From the Chi-squared Test", [online], August 5, 2022, [retrieved March 18, 2025], Internet <URL: https: / / best-biostatistics.com / contingency / fisher-exact.html>"What is the Chi-squared Test?", [online], July 5, 2022, [retrieved March 18, 2025], Internet <URL: https: / / www.business-research-lab.com / 220705-2 / >"A Simple Summary of the Chi-squared Distribution", [online], March 28, 2024, [retrieved March 18, 2025], Internet <URL: https: / / avilen.co.jp / personal / knowledge-article / chi-squared-distribution / >"Chi-squared Distribution Table", [online], [retrieved March 18, 2025], Internet <URL: https: / / www.saiensu.co.jp / book_support / 978-4-88384-140-0 / chi-square_distribution.pdf>"Why Do Women Leave Rural Areas Japan? -- A Case Study on Eliminating the "Gender Gap" Cited as One Contributing Factor", [online], November 26, 2024, [retrieved March 18, 2025], Internet <URL: https: / / japan.cnet.com / article / 35226536 / >"What is Wilcoxon Rank-sum Test? What is the Difference From the Mann-Whitney U Test?", [online], November 5, 2024, [retrieved March 18, 2025], Internet <URL: https: / / best-biostatistics.com / stat-test / w-test.html>
[0010] For example, Patent Document 1 describes a technique for detecting whether there is a URL specific to a particular type of malicious communication. In this technique, a Z-test is performed on the number of times a URL = a is communicated before and after a malicious communication for a particular type of malicious communication id = x and all other types of malicious communication id ≠ x. If the p-value calculated by the Z-test falls below the significance level, it is considered to be related to the type of malicious communication id = x, and URL = a is added to the blacklist. However, in this technique, when the number of URLs is very large, even if the test is performed at a significance level of 5% or 1%, some URLs that are not related to malicious communication will be added to the blacklist.
[0011] For example, Non-Patent Document 1 describes an example of conducting a chi-squared test on the pass rate of an anatomy exam. Specifically, the pass rate for the past 100 years of exams is compared with the pass rate for this year's exam, and a chi-squared test is performed under the null hypothesis that there was no difference between past and current pass rates. In this example, the null hypothesis was rejected in a chi-squared test at a significance level of 5%. This result suggests that there is a factor that lowers this year's pass rate compared to past pass rates (resulting in a difference), such as "the quality of this year's students was inferior to that of previous years." However, if the test is performed on 101 years of data at a significance level of 5%, the null hypothesis is rejected for 5 to 6 years, so this result does not necessarily mean that there was a change in the quality of this year's students.
[0012] In other words, the technologies described in Patent Document 1 and Non-Patent Document 1 above had the problem that they could not accurately determine whether or not there was a cause for bias in the data constituting a large number of groups.
[0013] This invention was made in view of the above circumstances and aims to provide a technology that can accurately determine whether or not there are factors that cause bias in the data constituting a large group.
[0014] To solve the above problems, one embodiment of the group testing device according to the present invention comprises a testing unit that performs a statistical test to determine whether or not there is a bias in multiple measured values constituting a group and calculates a significance probability, and a determination unit that determines whether or not the significance probabilities for each group calculated using multiple groups are evenly distributed in the range of 0 to 1, and if it is determined that the significance probabilities for each group are not evenly distributed in the range of 0 to 1, it determines that there is a bias in the measured values in the multiple groups.
[0015] According to one aspect of this invention, statistical tests are performed on each of the multiple groups to calculate the significance probability, and if the calculated significance probabilities for each group are not evenly distributed between 0 and 1, it is determined that there is a bias in the data across the multiple groups. This makes it possible to accurately determine, through statistical tests, whether or not there is a bias in the data constituting a group for a large number of groups, and to infer whether there is a cause for the bias in the data constituting the group.
[0016] Figure 1 is a schematic diagram showing one example configuration of the population testing device according to the embodiment. Figure 2 is a table showing the measured values of the male and female populations for the entire country and Kita Ward, Sakai City, in a specific example. Figure 3 is a table showing the expected frequency calculated for each element of the table shown in Figure 2. Figure 4 is a table showing the results of calculations to find the chi-squared value for each element of the table shown in Figure 2 based on the measured value and the expected frequency. Figure 5 is a flowchart showing an example of the processing procedure and processing content of the estimation process executed by the population testing device according to the embodiment. Figure 6 is a table sorted in ascending order of the p-values calculated in the specific example for all municipalities. Figure 7 is a graph showing the relationship between the p-value and rank (percentage) obtained by a chi-squared test of the male-female ratio for ages 0 to 4 in all municipalities in the specific example. Figure 8 is a graph showing the relationship between the p-value and rank (percentage) obtained by a chi-squared test of the male-female ratio for ages 20 to 24 in all municipalities in the specific example. Figure 9 is a table showing the number of calculated p-values and ideal numbers for each age group for the male-female ratio of 0 to 4 years old in all municipalities in the specific example. Figure 10 is a table showing the number of calculated p-values and ideal numbers for each age group for the male-female ratio of 20 to 24 years old in all municipalities in the specific example. Figure 11 is a flowchart showing an example of the processing procedure and content of the inference process using the second judgment method executed by the group testing device 1 according to the embodiment. Figure 12 is a graph showing the relationship between the p-value and rank (proportion) obtained by a chi-squared test of the male-female ratio of 20 to 24 years old in all municipalities in the specific example. Figure 13 is a flowchart showing an example of the processing procedure and content of the inference process using the third judgment method executed by the group testing device 1 according to the embodiment.
[0017] Embodiments of this invention will be described below with reference to the drawings.
[0018] Figure 1 is a schematic diagram showing one example of the configuration of a population testing device according to the embodiment. The population testing device 1 is, for example, a personal computer (PC) that uses statistical testing to determine whether or not there is a bias in the data constituting a group of multiple groups, and infers whether there is a cause for the bias in the data constituting the group. Here, a group is a collection of multiple data, and the data constituting a group will be referred to as measured values below. The population testing device 1 according to the embodiment can perform the inference process described later if the number of groups is two or more, but in order to obtain statistically meaningful results, it is desirable that the number of groups be a large number, for example, around 100.
[0019] In this embodiment, the group testing device 1 comprises a control unit 10, a storage unit 20, an input unit 30, an output unit 40, and a communication unit 50, and each component is connected via a bus.
[0020] The input unit 30 may include, for example, a user interface such as a mouse or keyboard for inputting changes to numerical values of various settings of the group testing device 1, a microphone, a touch panel, a camera, various sensors, etc.
[0021] The output unit 40 may include a display means such as a monitor that allows visual recognition of operation screens and the like necessary for operating the group testing device 1, and an audio output means such as a speaker that allows auditory recognition. The output unit 40 may also be configured to output image information and audio information to externally connected display means and audio output means.
[0022] The communication unit 50 can communicate with other external components via a network such as a LAN or the Internet, and performs the transmission and reception of various types of data.
[0023] The storage unit 20 includes, for example, a main storage unit and an auxiliary storage unit. The main storage unit may include, for example, ROM (read-only memory) and RAM (random-access memory). ROM is a non-volatile memory used exclusively for reading data, and can store data and various setting values used by the control unit 10 in performing various processes. RAM can be used as a so-called work area to temporarily store data when the control unit 10 performs various processes. In this embodiment, the main storage unit is, for example, RAM and is used as memory.
[0024] The auxiliary storage unit is a non-temporary computer-readable storage medium for a computer centered around the control unit 10. Examples of auxiliary storage units include EEPROM® (electric erasable programmable read-only memory), HDD (hard disk drive), or SSD (solid state drive). In addition to middleware such as the OS (operating system), the auxiliary storage unit can store computer programs used by the control unit 10 in performing various processes described later, as well as data or various setting values generated by processing in the control unit 10.
[0025] In one embodiment, the storage unit 20 includes a group data storage unit 21. The group data storage unit 21 stores multiple measured values that constitute multiple groups that are the subject of the inference process. The group data storage unit 21 stores each of the multiple measured values in association with information that can identify the group to which the measured value belongs (e.g., an identification number). Alternatively, the group data storage unit 21 stores, for example, a database for each group, which is composed of multiple measured values that constitute any group.
[0026] The control unit 10 is a processor that includes at least one CPU (Central Process Unit), MPU (micro processing unit), GPU (Graphics Processing Unit), FPGA (field-programmable gate array), etc., and can realize various functions of the group testing device 1 using computer programs such as system software, application software, or firmware stored in the storage unit 20.
[0027] In this embodiment, the control unit 10 comprises a testing unit 11 and a judgment unit 12. The testing unit 11 calculates a p-value (significance probability), which is a value indicating whether or not there is a bias in the observed values that constitute the group, for each of the multiple groups that are the subject of the inference process. Specifically, the testing unit 11 takes the null hypothesis that there is no bias in the observed values that constitute any group (the alternative hypothesis is that there is a bias in the observed values that constitute the group) and performs a statistical test according to the observed values that constitute the group, and calculates a p-value. For example, if the observed values that constitute the group are numerical data, the testing unit 11 performs a t-test or a Wilcoxon rank-sum test, etc. For example, if the observed values that constitute the group are categorical data belonging to a specific category, the testing unit 11 performs a chi-squared test or a Fisher exact test, etc.
[0028] The judgment unit 12 determines whether or not there is a bias in the measured values across multiple groups, based on the p-values for each group calculated by the testing unit 11. Specifically, the judgment unit 12 determines, using a predetermined judgment method, whether or not the p-values for each group calculated by the testing unit 11 are evenly distributed between 0 and 1. If it determines that they are not evenly distributed, it determines that there is a bias in the measured values across multiple groups. There are three judgment methods used by the judgment unit 12 to determine whether or not the p-values for each of the multiple groups are evenly distributed between 0 and 1: the first judgment method, the second judgment method, and the third judgment method. The specific processing of the first judgment method, the second judgment method, and the third judgment method will be described later.
[0029] Next, the statistical testing and inference processes performed by the population testing device 1 according to the embodiment will be explained based on specific examples.
[0030] In the specific example shown below, the population testing device 1 will examine whether there is a gender bias in each Japanese municipality, focusing on a certain age group. That is, in the specific example, the population to be estimated is each Japanese municipality, and the measured values that make up the population are the male population (hereinafter simply referred to as "men") and the female population (hereinafter simply referred to as "women") for each municipality.
[0031] For example, according to the population by age group in the Basic Resident Register as of January 1, 2024, there were 2,047,019 boys and 1,949,555 girls aged 0-4 in Japan, resulting in an overall male-to-female ratio of approximately 1.05. In contrast, Chiyoda Ward in Tokyo had 1,392 boys and 1,370 girls, with a male-to-female ratio of approximately 1.02, while Hiraya Village in Shimoina District, Nagano Prefecture, had 2 boys and 6 girls, with a male-to-female ratio of approximately 0.33. The purpose of the group testing device 1 in this specific example is to determine whether the male-to-female ratios in each municipality are biased compared to the national average, and to infer whether there are any causes for the bias in the male-to-female ratio in each municipality.
[0032] First, we will explain the statistical tests performed by the population testing device 1 according to this embodiment, using a specific example.
[0033] When focusing on a particular municipality, in order to determine whether the gender ratio in that municipality is skewed compared to the gender ratio of Japan as a whole, Fisher's exact test or chi-squared test is performed on the observed values of that municipality and the gender ratio of men and women in Japan as a whole, under the null hypothesis that the gender ratio in that municipality is not skewed compared to Japan as a whole. For the sake of simplicity, in the following specific example, we will perform the chi-squared test for all municipalities. However, if the number of data points (i.e., the value for at least one of the genders) is 5 or less, it is desirable to perform Fisher's exact test to improve accuracy (see Non-Patent Literature 2), and Fisher's exact test may be performed only when the number of data points for either the gender or gender of the municipality being tested is small, or Fisher's exact test may be performed for all municipalities.
[0034] First, as an example of a statistical test, we will explain below an example of a chi-squared test conducted to determine whether the male-female ratio for children aged 0 to 4 years is skewed in Kita Ward, Sakai City, Osaka Prefecture (see Non-Patent Document 3).
[0035] Figure 2 is a table showing the measured male and female populations for the entire country and Kita Ward, Sakai City, in a specific example. Note that each element in the "National" row of Figure 2 represents the total population data for Japan (male population, female population, and combined population) minus the measured values for Kita Ward, Sakai City. First, based on the measured values shown in Figure 2, we calculate the expected frequency of each element in the table, assuming that the male-female ratio (i.e., the measured values for each gender) in Kita Ward, Sakai City is not biased compared to the national average.
[0036] Figure 3 is a table showing the expected frequencies for each element of the table shown in Figure 2. Expected frequencies are the frequencies of each element that are expected to occur by working backward from the ratio of the sum of the row elements and the sum of the column elements in the table. For example, the expected frequency for the element "National - Male" can be calculated as: National total 3,996,574 × (Male total 2,050,008 ÷ Total 4,002,677) = 2,046,882.
[0037] Figure 4 is a table showing the results of calculations performed to determine the chi-squared value for each element of the table shown in Figure 2, based on the observed value and the expected frequency. That is, the table shown in Figure 4 shows the result for each element as follows: (observed value - expected frequency) 2 This is the value obtained by dividing by the expected frequency. As shown in Figure 4, the sum of each element in the table shown in Figure 4 is the chi-squared value. When converting the calculated chi-squared value to a p-value, degrees of freedom are needed. The degrees of freedom are (number of rows in the table - 1) × (number of columns in the table - 1), so in this case it is 1.
[0038] Here, the p-value for degrees of freedom n is the probability distribution function of the chi-squared distribution, so it can be calculated using the following formula (see Non-Patent Document 4). However, Γ(n / 2) is the gamma function and can be calculated using the following formula.
[0039] In practical terms, the test is performed by comparing the calculated chi-square value with a chi-square distribution table. A well-known chi-square distribution table, such as the one shown in Non-Patent Document 5, may be used. Looking at the entry for a degree of freedom of 1 and α = 0.01 in the chi-square distribution table, the value is 6.63. Since the chi-square value of 12.27504 shown in FIG. 4 is larger than 6.63, the null hypothesis is rejected at a significance level of 1%. That is, this suggests that the sex ratio of 0- to 4-year-olds in Sakai City's Kita Ward is biased to an extent that would only occur with a probability of 1% or less compared to that of Japan as a whole.
[0040] However, when a test at the 1% significance level is performed for each municipality, it is natural that the null hypothesis is rejected in 1% of the many municipalities, and this alone does not necessarily mean that there is a cause for a biased sex ratio in Sakai City's Kita Ward. Accordingly, the multiple testing apparatus 1 according to the embodiment performs the same chi-square test for all municipalities, and determines whether there is a bias for each municipality based on the calculated p-value of each municipality.
[0041] Next, the estimation processing executed by the multiple testing apparatus 1 according to the embodiment will be described based on a specific example.
[0042] FIG. 5 is a flowchart illustrating an example of the processing procedure and processing content of the estimation processing executed by the multiple testing apparatus according to the embodiment. The flowchart of FIG. 5 illustrates the processing procedure when the determination unit 12 uses a first determination method to make a determination. The control unit 10 of the multiple testing apparatus 1 starts the following processing in response to, for example, an operation from an operator instructing execution of the estimation processing.
[0043] First, the test unit 11 selects an arbitrary population (step ST1). The test unit 11 performs a statistical test according to actual measured values on comparison target data and the selected population to determine whether there is a bias in the actual measured values constituting the selected population (step ST2). The comparison target data is data composed of ideal values obtained when it is assumed that there is no bias in the actual measured values constituting the population. For example, in the specific example described above, the actual population data of males and females in each municipality are actual measured values, whereas the comparison target data is the population data of males and females for Japan as a whole.
[0044] The testing unit 11 calculates a p-value based on a test statistic calculated by testing (a chi-square value in the specific example) (step ST3). The calculated p-value is a probability representing information about whether there is a bias in the actually measured values constituting the selected population.
[0045] The testing unit 11 determines whether p-values have been calculated for all populations (step ST4). If the testing unit 11 has not calculated p-values for all populations (step ST4: NO), that is, if there still exists a population for which a p-value has not been calculated, the process returns to step ST1, and steps ST1 to ST4 are repeated until p-values have been calculated for all populations.
[0046] First, in the specific example, when the population testing apparatus 1 executes the processes of steps S1 to S4 described above, that is, for each municipality, under the null hypothesis that the sex ratio of the municipality is not biased compared to that of the whole of Japan, when a chi-square test is performed on the population data of men and women for the whole of Japan and the actually measured values of the municipality to calculate a p-value, the result is as follows. Note that the procedure of the chi-square test performed by the testing unit 11 in the specific example may be the same as that described above.
[0047] Fig. 6 is a table obtained by sorting data of all municipalities in ascending order of the p-values calculated in the specific example. As described above, when the above-described test is performed on each of all municipalities to calculate a p-value, if there is no bias in the sex ratio of people aged 0 to 4 years for each municipality, according to the definition of a p-value, the p-value for each municipality should follow a uniform distribution from 0 to 1.
[0048] Figure 7 is a graph showing the relationship between the p-value and rank (proportion) obtained by a chi-squared test of the male-female ratio for children aged 0 to 4 years in all municipalities in a specific example. In the graph shown in Figure 7, the vertical axis represents the p-value, and the horizontal axis represents the rank when each municipality is sorted in ascending order of p-value. Here, rank represents the position of each municipality within all municipalities as a percentage from 0 to 100% (0.0 to 1.0). In the graph shown in Figure 7, the solid line (p-value) plots the position of each municipality on the graph based on the calculated p-value and rank. In the graph shown in Figure 7, the dotted line (auxiliary line) plots the ideal p-value for each rank and is a straight line passing through (0,0) and (1,1).
[0049] As shown in Figure 7, the p-values obtained by a chi-square test on the male-female ratio of children aged 0 to 4 in each municipality are distributed almost uniformly between (0,0) and (1,1). Based on this, it can be inferred that there are no factors that cause a bias in the male-female ratio of children aged 0 to 4 in each municipality.
[0050] In contrast, if the group testing device 1 performs the same steps S1 to S4 on the male-female ratio of 20-24 year olds in each municipality, the results will be as follows.
[0051] Figure 8 is a graph showing the relationship between the p-value (p-value) and rank (proportion) obtained by a chi-squared test of the male-female ratio of 20-24 year olds in all municipalities in a specific example. Similar to Figure 7, the vertical axis of the graph in Figure 8 shows the p-value, and the horizontal axis shows the rank of each municipality when sorted in ascending order of p-value.
[0052] As shown in Figure 8, the p-values obtained by chi-square tests on the male-female ratio of 20-24 year olds in each municipality are skewed downwards and to the right relative to the guideline. This indicates that there are many gender ratio biases that occur with a low probability, suggesting that each municipality has its own cause for this gender ratio bias in this age group.
[0053] These test results are consistent with the claim that "the gender ratio becomes skewed due to employment and further education." For example, Non-Patent Document 6 shows that "in the case of Toyooka City, 52.2% of men who left the city for further education or other reasons later returned to the city, while only 26.7% of women returned."
[0054] Furthermore, if the distribution of p-values is skewed to the upper left of (0,0) to (1,1), it can be concluded that the male-female ratio is unnaturally uniform. It is also unlikely that the male-female ratios of all municipalities would be perfectly identical, and in such cases, the graph will be skewed to the upper left of the p-values. This suggests that there are factors at work that are causing the male-female ratios to be equal across municipalities.
[0055] In this way, by determining whether the p-values for each group are evenly distributed between 0 and 1, it is possible to determine whether there is a bias in the measured values across multiple groups. Furthermore, focusing on the area around (0.6, 0.4) indicated by the arrows in Figure 8, it can be seen that a gender ratio bias, which would occur in only 40% of cases according to an ideal p-value distribution, is occurring in as much as 60% of cases. The group testing device 1 according to one embodiment has the effect of detecting bias even in areas where the significance level does not fall below 5% or 1%.
[0056] The group testing apparatus 1 according to the embodiment uses one of the following determination methods—the first, second, and third determination methods—to determine whether the p-values for each group calculated in steps ST1 to ST4 are evenly distributed between 0 and 1, and to determine whether there is a bias in the measured values across multiple groups.
[0057] Returning to Figure 5, we will explain the inference process using the first judgment method. The first judgment method involves setting multiple classes by dividing the range from 0 to 1 into two or more natural numbers, assigning the significance probability for each group to each class and counting them, and comparing the number of significance probabilities counted for each class with the ideal number of significance probabilities for each class when the significance probabilities are evenly distributed between 0 and 1 to determine whether the significance probabilities for each group are evenly distributed between 0 and 1.
[0058] If the testing unit 11 determines that it has calculated p-values for all groups (step ST4: YES), the judgment unit 12 sets classes by dividing the p-values (ranging from 0 to 1) into N classes, and counts the calculated p-values for each group by assigning them to each class (step ST11). Here, N is a natural number greater than or equal to 2.
[0059] The judgment unit 12 generates ideal numbers for each class when the p-values are evenly distributed in the range of 0 to 1 (step ST12). The ideal numbers for each class are evenly distributed as the total number of people in the group ÷ N.
[0060] The decision unit 12 performs a chi-squared test on the number of p-values for each class counted in step ST11 and the ideal number for each class generated in step ST12, based on the null hypothesis that the distribution of p-values for each class in each class is no different from the ideal distribution (step ST13).
[0061] Subsequently, the judgment unit 12 compares the p-value based on the chi-squared value calculated by the test with the significance level (step ST14). The significance level used in step ST14 is typically a general significance level of 5% or 1%.
[0062] If the calculated p-value is below the significance level (step ST14: YES), the judgment unit 12 rejects the null hypothesis and determines that the distribution of p-values for each group in each class cannot be said to be without differences from the ideal distribution, that is, the p-values for each group are not evenly distributed between 0 and 1, and thus determines that there is a bias in the observed values in multiple groups (step ST15). Based on this result, it can be inferred that there is a cause for the bias in the observed values in each group.
[0063] On the other hand, if the calculated p-value is above the significance level (step ST14: NO), the judgment unit 12 determines that the null hypothesis is not rejected, and that the distribution of p-values for each group in each class is no different from the ideal distribution, that is, that the p-values for each group are evenly distributed between 0 and 1, and that there is no bias in the measured values among multiple groups (step ST16). Based on this result, it can be inferred that there is no cause for bias in the measured values among the groups.
[0064] Figure 9 is a table showing the number of calculated p-values and the ideal number for each class in the male-female ratio of 0 to 4 years old in all municipalities in the specific example. For example, if we divide the p-values for the male-female ratio of 0 to 4 years old in the specific example into 10 classes (i.e., N=10), and calculate the number of p-values (actual) and the ideal number of p-values (ideal) for each municipality in each class, we get the results shown in Figure 9.
[0065] When a chi-squared test is performed on the actual and ideal values for each class shown in Figure 9, the p-value is 0.54. In other words, there is a 54% chance that such a bias in the number of p-values for each class is due to chance, and it can be concluded that the p-values for each municipality are evenly distributed between 0 and 1. Therefore, it is difficult to say that there is a cause for the bias in the male-female ratio for ages 0 to 4 among municipalities.
[0066] Figure 10 is a table showing the number of calculated p-values and the ideal number for each age group of 20-24 year olds in all municipalities in the specific example. Similar to Figure 9, the p-values for the 20-24 year old age group in the specific example are divided into 10 classes, and the number of p-values for each municipality in each class and the ideal number of p-values in each class are calculated, as shown in Figure 10.
[0067] When a chi-squared test is performed on the measured and ideal values for each class shown in Figure 10, the p-value is 5.63 × 10⁻¹⁰. ―106 This results in an extremely small value. It is unlikely that such a bias in the number of p-values for each class would occur by chance, and since it can be determined that the p-values for each municipality are not evenly distributed between 0 and 1, it is suggested that there is a cause for the bias in the gender ratio between municipalities.
[0068] Next, we will explain the inference process using the second judgment method. The second judgment method is a method for determining whether the significance probabilities for each group are evenly distributed between 0 and 1 by comparing each significance probability for each group with the ideal significance probability at the rank of that significance probability when all significance probabilities are sorted in ascending order.
[0069] Figure 11 is a flowchart showing an example of the processing procedure and content of the inference process using the second judgment method executed by the group testing device 1 according to the embodiment. The flowchart of the inference process shown in Figure 11 differs from the flowchart shown in Figure 5 in that the processing from step ST4 onwards is different.
[0070] In step ST4 of the flowchart shown in Figure 5, if the testing unit 11 determines that it has calculated p-values for all groups, the judgment unit 12 calculates the absolute difference between each calculated p-value for each group and the rank percentage, and calculates the maximum value of the absolute difference (step ST21).
[0071] Figure 12 is a graph showing the relationship between the p-value and rank (proportion) obtained by a chi-squared test on the male-female ratio of 20-24 year olds in all municipalities in a specific example. As shown in Figure 12, the absolute value of the difference between the p-value and the rank proportion is equal to the distance between the p-value (solid line) for each rank and the auxiliary line measured parallel to the vertical axis. Here, since the auxiliary line on the graph represents the ideal p-value for each rank, calculating the maximum absolute value of the difference between the p-value and the rank proportion means determining the degree to which the p-value for each group deviates most from the ideal p-value.
[0072] The determination unit 12 compares the maximum absolute value of the calculated difference with a predetermined threshold (step ST22). The threshold can be set to an appropriate value (for example, 0.1) depending on the degree of rigor required for the determination.
[0073] If the judgment unit 12 determines that the maximum absolute value of the calculated difference is greater than a predetermined threshold (step ST22: YES), it determines that the p-values for each group deviate significantly from the ideal p-value, that is, that the p-values for each group are not evenly distributed between 0 and 1, and determines that there is a bias in the measured values across multiple groups (step ST23). Based on this result, it can be inferred that there is a cause for the bias in the measured values in each group.
[0074] On the other hand, if the judgment unit 12 determines that the maximum value of the calculated absolute value is not greater than a predetermined threshold (step ST22: NO), it determines that the discrepancy between the ideal p-value and the p-value for each group is small, that is, that the p-values for each group are evenly distributed between 0 and 1, and thus determines that there is no bias in the measured values among multiple groups (step ST24). Based on this result, it can be inferred that there is no cause for bias in the measured values among the groups.
[0075] For example, when calculating the absolute difference between the p-value for each municipality and the proportion of the rank of that municipality's p-value (the ideal p-value) for the male-female ratio of children aged 0 to 4 in a specific example, the maximum value is 0.051. If the threshold is set to 0.1, this value is smaller than the threshold, indicating that the p-values for each municipality do not deviate significantly from the ideal p-value, and that the p-values for each municipality are evenly distributed between 0 and 1. Therefore, it is difficult to say that there are any causes that would create a bias in the male-female ratio of children aged 0 to 4 among municipalities.
[0076] In contrast, when calculating the absolute difference between the p-value for each municipality and the proportion of the rank of that municipality's p-value (the ideal p-value) for the male-female ratio of 20 to 24-year-olds in a specific example, the maximum value is 0.278. Similarly, if the threshold is set to 0.1, this value is greater than the threshold, indicating that the p-values for each municipality deviate significantly from the ideal p-value. This suggests that the p-values for each municipality are not evenly distributed between 0 and 1, indicating that there are factors causing the bias in the male-female ratio among municipalities.
[0077] Next, we will explain the inference process using the third judgment method. The third judgment method involves creating a sequence of numbers equal to the total number of numbers in the group and evenly distributed between 0 and 1, and then comparing the significance probability for each group with the created sequence to determine whether or not the significance probability for each group is evenly distributed between 0 and 1.
[0078] Figure 13 is a flowchart showing an example of the processing procedure and content of the inference process using the third judgment method executed by the group testing device 1 according to the embodiment. The flowchart of the inference process shown in Figure 13 differs from the flowchart shown in Figure 5 in that the processing from step ST4 onwards is different.
[0079] In step ST4 of the flowchart shown in Figure 5, if the testing unit 11 determines that it has calculated p-values for all groups, the judgment unit 12 creates a sequence of numbers equal to the total number of p-values (total number of groups), which are evenly distributed between 0 and 1 (step ST31).
[0080] The judgment unit 12, under the null hypothesis that there is no bias between the p-values for each group and the sequence that is evenly distributed between 0 and 1, performs a Wilcoxon rank-sum test on the p-values for each group calculated in steps ST1 to ST4 and the sequence created in step ST31 (step ST32) (see Non-Patent Literature 7).
[0081] Subsequently, the judgment unit 12 compares the p-value based on the test statistic calculated by the test with the significance level (step ST33). The significance level used in step ST33 is typically a general significance level of 5% or 1%.
[0082] If the calculated p-value is below the significance level (step ST33: YES), the judgment unit 12 rejects the null hypothesis and determines that there is a bias between the p-values for each group and the sequence that is evenly distributed between 0 and 1. In other words, it determines that the p-values for each group are not evenly distributed between 0 and 1, and thus determines that there is a bias in the measured values for multiple groups (step ST34). Based on this result, it can be inferred that there is a cause for the bias in the measured values for each group.
[0083] On the other hand, if the calculated p-value is above the significance level (step ST33: NO), the judgment unit 12 determines that the null hypothesis is not rejected, that there is no bias between the p-values for each group and the sequence that is evenly distributed between 0 and 1, that is, that the p-values for each group are evenly distributed between 0 and 1, and that there is no bias in the measured values across multiple groups (step ST35). Based on this result, it can be inferred that there is no cause for bias in the measured values across groups.
[0084] For example, in this specific case, there are 2267 municipalities nationwide, so the sequence created in step ST31 will be 0 / 2266, 1 / 2266, 2 / 2266, ..., 2265 / 2266, 1.
[0085] When we perform a Wilcoxon rank-sum test on the male-female ratio for children aged 0 to 4 in the specific example, using the p-values for each municipality and the created sequence, the p-value is 0.00164, indicating that such a bias occurs with a probability of only 0.1%. This is because the Wilcoxon rank-sum test compares the ranks of the ideal value and the actual value, so the probability of a bias occurring where the actual p-value is always above the ideal value, as shown in the graph in Figure 7, is calculated to be low.
[0086] In contrast, when Wilcoxon rank-sum test is performed on the p-values for each municipality and the above sequence created for the male-female ratio of 20-24 year olds in the specific example, the p-value is 2.40 × 10⁻¹⁰. ―86 This results in an extremely small value. It is unlikely that such a bias in the p-values for each municipality would occur by chance, and since it can be determined that the p-values for each municipality are not evenly distributed between 0 and 1, it is suggested that there is a cause for the bias in the gender ratio among the municipalities.
[0087] Furthermore, the Wilcoxon rank-sum test performed in the inference process using the third decision method is essentially the same as the Mann-Whitney U test. Therefore, in the inference process using the third decision method, the population testing device 1 may perform the Mann-Whitney U test instead of the Wilcoxon rank-sum test. Specifically, in step ST32 of the flowchart shown in Figure 13, the decision unit 12 performs the Mann-Whitney U test on the p-values for all populations calculated in steps ST1 to ST4 and the sequence created in step ST31.
[0088] As described above, the group testing device 1 according to one embodiment performs a statistical test on each of the multiple groups to determine whether or not there is a bias in the measured values that make up the group, calculates the significance probability, determines whether or not the significance probabilities for each group are evenly distributed in the range of 0 to 1, and if it is determined that they are not evenly distributed, it is determined that there is a bias in the measured values in the multiple groups.
[0089] If it is determined that there is a bias in the measured values across multiple groups, it can be inferred that there is some cause that is causing the bias in the measured values that make up the groups. Therefore, according to the group testing device 1 of one embodiment, it is possible to accurately determine whether or not there is a bias in the data that makes up a large number of groups through statistical testing, and it is possible to estimate whether there is a cause that is causing the bias in the data that makes up the groups.
[0090] Furthermore, the group testing device 1 according to one embodiment determines whether the significance probabilities for each group are evenly distributed between 0 and 1 by setting multiple classes by dividing the range from 0 to 1 into two or more natural numbers, assigning and counting the significance probabilities for each group to each class, and comparing the number of significance probabilities for each class counted with the ideal number of significance probabilities for each class when the significance probabilities are evenly distributed between 0 and 1; a second determination method determines whether the significance probabilities for each group are evenly distributed between 0 and 1 by comparing each significance probability for each group with the ideal significance probability at the rank of that significance probability when all significance probabilities are sorted in ascending order; and a third determination method determines whether the significance probabilities for each group are evenly distributed between 0 and 1 by creating a sequence of numbers equal to the total number of groups and evenly distributed between 0 and 1, and comparing the significance probabilities for each group with the created sequence.
[0091] It should be noted that the present invention is not limited to the embodiments described above, and can be modified in various ways during implementation without departing from its essence. Furthermore, each embodiment may be combined as appropriate, and in that case, the combined effects can be obtained. Moreover, the above embodiments include various inventions, and various inventions can be extracted by selecting combinations from the multiple constituent elements disclosed. For example, if the problem can be solved and effects obtained even if some constituent elements are deleted from all the constituent elements shown in the embodiment, then the configuration with these deleted constituent elements can be extracted as an invention.
[0092] 1...Group testing device 10...Control unit 11...Testing unit 12...Decision unit 20...Storage unit 21...Group data storage unit 30...Input unit 40...Output unit 50...Communication unit
Claims
1. A group testing device comprising: a testing unit that performs a statistical test to determine whether or not there is a bias in multiple measured values constituting a group and calculates a significance probability; and a determination unit that determines whether or not the significance probabilities for each of the multiple groups calculated using the multiple groups are evenly distributed in the range of 0 to 1, and if it is determined that the significance probabilities for each of the groups are not evenly distributed in the range of 0 to 1, it determines that there is a bias in the measured values in the multiple groups.
2. The determination unit determines whether the significance probabilities for each group are evenly distributed between 0 and 1 by any of the following determination methods: a first determination method which determines whether the significance probabilities for each group are evenly distributed between 0 and 1 by setting up a plurality of classes obtained by dividing the range from 0 to 1 into two or more natural numbers, assigning the significance probabilities for each group to each class and counting them, and comparing the number of significance probabilities for each class that has been counted with the ideal number of significance probabilities for each class when the significance probabilities are evenly distributed between 0 and 1; a second determination method which determines whether the significance probabilities for each group are evenly distributed between 0 and 1 by comparing each of the significance probabilities for each group with the ideal significance probability at the rank of the significance probability when all the significance probabilities are sorted in ascending order; and a third determination method which determines whether the significance probabilities for each group are evenly distributed between 0 and 1 by creating a sequence of numbers equal to the total number of numbers in the group and evenly distributed between 0 and 1, and comparing the significance probabilities for each group with the created sequence; and 3. The group testing apparatus according to claim 2, wherein the judgment unit performs a chi-squared test on the number of significance probabilities for each class counted in the first judgment method and on the ideal number of significance probabilities for each class; in the second judgment method, for each of the significance probabilities for each group, calculates the absolute value of the difference between the ideal significance probability at the rank of the significance probability when all the significance probabilities are sorted in ascending order and the significance probability; compares the maximum value of the absolute values of all the calculated differences with a predetermined threshold; and in the third judgment method, performs a rank-sum test on the significance probabilities for each group and the created sequence.
4. A computer program that causes a computer processor to perform the processing performed by each part of the apparatus described in any one of claims 1 to 3.