Information Processing System, Information Processing Method, and Computer Program
The information processing system addresses the challenge of generating high-quality statistical data with fine granularity by combining personal data attributes and areas, using inverse distance weighting to enhance data quality and reliability.
Patent Information
- Application Number
- JP2025042928
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2045-03-17
AI Technical Summary
Existing techniques struggle to generate high-quality statistical data with fine granularity by combining personal data attributes and areas due to reduced sample sizes, leading to potential deterioration in data quality.
An information processing system and method that includes acquiring and statistically processing first and second explanatory data to generate area and attribute-specific statistical data, using inverse distance weighting to calculate values for empty areas, and combining these data sets to enrich information.
The system generates high-quality statistical data with fine granularity by addressing the issue of reduced sample sizes, ensuring accurate and reliable analysis of human behavior and consciousness.
Smart Images

Figure 0007717996000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing system and an information processing method.
Background Art
[0002] Conventionally, a technique for generating customer data rich in information amount by combining multiple types of data related to customers is known. For example, a technique of matching an existing customer database with external survey data using gender, age, and area as keys and adding the information stored in the external survey data to the existing customer database is known (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present inventors consider converting personal data into statistical data for each combination of attributes and areas by statistically processing first data that describes the characteristics of two or more corresponding individuals for each attribute and area for a plurality of individuals. The present inventors consider generating data with a large amount of information by combining the statistical data for each attribute and area thus generated with second data that describes the characteristics of people for each attribute and area.
[0005] Data related to people often has information on attributes and areas attached. However, it is difficult to meaningfully combine the data of a plurality of individuals with only the information on attributes and areas. On the other hand, after statistically processing personal data, it is possible to generate meaningful combined data by combining a plurality of data using attributes and areas as keys.
[0006] Also, when statistically processing personal data, if the granularity of the attributes and areas, which are the units of statistical processing, is reduced, it is possible to generate statistical data that represents the characteristics of the set of individuals belonging to the attributes and areas with a fine granularity.
[0007] However, when the granularity of the attributes and areas is reduced, the number of samples for each attribute and area decreases. Due to this small number of samples, when replacing personal data with statistical data through statistical processing, the quality of the statistical data may deteriorate.
[0008] Therefore, according to one aspect of the present disclosure, it is desirable to be able to provide a technique capable of generating high-quality statistical data when statistically summarizing data describing personal characteristics at a fine granularity to generate statistical data for data combination.
Means for Solving the Problem
[0009] According to one aspect of the present disclosure, an information processing system is provided. The information processing system includes a first acquisition unit, a generation unit, a second acquisition unit, and a combination unit. The first acquisition unit is configured to acquire first explanatory data that describes the characteristics of a plurality of individuals belonging to a first set. The first explanatory data describes, for each individual, the first feature quantity of the corresponding individual in association with the attributes of the corresponding individual and information on the area to which the corresponding individual belongs.
[0010] The generation unit is configured to generate statistical data by statistically processing the first explanatory data. The statistical data includes, for a plurality of areas, area statistical data for each area regarding the set of individuals belonging to the corresponding area.
[0011] The area statistical data includes, for a plurality of attributes, first feature data for each attribute. The first feature data describes statistical values of the first feature quantities regarding one or more individuals calculated by statistical processing on the first feature quantities of one or more individuals having the corresponding attribute among the set of individuals within the corresponding area.
[0012] The second acquisition unit is configured to acquire second explanatory data that describes the characteristics of a plurality of people belonging to the second set. The second explanatory data includes second feature data for each person. The second feature data describes the second feature amount of the corresponding person in association with the attribute of the corresponding person and the information on the area to which the corresponding person belongs.
[0013] The combining unit is configured to combine the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data in which the combinations of the attribute and the area match.
[0014] According to one aspect of the present disclosure, when there is an empty area, which is an area where there is no individual having the corresponding attribute in the first set, among the plurality of areas for each attribute, the generation unit executes statistical processing on the first feature amounts of one or more individuals belonging to one or more peripheral areas located around the empty area among the plurality of areas, thereby calculating a statistical value of the corresponding attribute for the empty area and generating area statistical data for the empty area.
[0015] According to the information processing system configured as described above, in order to combine data, when generating feature data for each area and attribute by statistically summarizing data that describes the characteristics of individuals, high-quality statistical data can be generated. That is, it is possible to suppress the occurrence of a lack of statistical values regarding combinations of attributes and areas for which there are no corresponding persons due to a small number of samples.
[0016] According to one aspect of the present disclosure, for each combination of an area and an attribute with respect to a plurality of areas and a plurality of attributes, the generation unit may be configured to calculate a statistical value of the corresponding attribute for the corresponding area by executing statistical processing based on the first feature amounts of one or more individuals having the corresponding attribute in the corresponding area and the first feature amounts of one or more individuals belonging to one or more peripheral areas, which are one or more areas located around the corresponding area.
[0017] According to the information processing system configured as described above, in order to combine data, when generating feature data for each area and attribute by statistically summarizing data that describes the characteristics of individuals, high-quality statistical data can be generated. That is, due to the small number of samples, the possibility of large variations in statistical values can be suppressed by calculating statistical values taking into account the surrounding areas.
[0018] According to one aspect of the present disclosure, when the corresponding area of the first set is an empty area where there is no individual having the corresponding attribute, the generation unit may be configured to calculate a statistical value of the corresponding attribute for the empty area by performing statistical processing on the first feature amounts of one or more individuals belonging to one or more areas located around the empty area as one or more surrounding areas.
[0019] According to one aspect of the present disclosure, the generation unit may be configured to calculate a representative value of the first feature amount in one or more individuals having the corresponding attribute in the corresponding area for each combination of area and attribute with respect to a plurality of areas and a plurality of attributes.
[0020] According to one aspect of the present disclosure, with respect to an empty area, the generation unit may be configured to calculate a statistical value of the corresponding attribute by a weighted sum of the representative values in each of one or more areas located around the empty area.
[0021] According to one aspect of the present disclosure, the generation unit may be configured to calculate, for each combination of area and attribute, the ratio of individuals whose first feature amount satisfies a specific condition among one or more individuals having the corresponding attribute in the corresponding area.
[0022] According to one aspect of the present disclosure, the generation unit may be configured to calculate a statistical value by a weighted sum of the representative value in the corresponding area and the representative values in each of one or more areas located around the corresponding area for each combination of area and attribute.
[0023] According to one aspect of the present disclosure, for each combination of an area and an attribute with respect to a plurality of areas and a plurality of attributes, the generation unit may calculate the ratio of individuals whose first feature amount satisfies a specific condition among one or more individuals having the corresponding attribute in the corresponding area.
[0024] According to one aspect of the present disclosure, for each combination of an area and an attribute, the generation unit may be configured to calculate a statistical value by a weighted sum of the ratio in the corresponding area and the ratios in each of one or more areas located around the corresponding area.
[0025] According to one aspect of the present disclosure, the generation unit may be configured to calculate a statistical value using the inverse distance weighting (IDW) method. According to one aspect of the present disclosure, the generation unit may be configured to calculate a statistical value by a weighted sum based on the inverse distance weighting (IDW) method. According to the information processing system configured in this way, it is possible to calculate a highly reliable statistical value with a small number of samples.
[0026] According to one aspect of the present disclosure, the first feature data may describe, as the first feature amount, a feature amount related to at least one of the behavior and consciousness of the corresponding individual. The information processing system configured in this way can generate combined data useful for analyzing at least one of human behavior and consciousness.
[0027] According to one aspect of the present disclosure, the first feature data may describe, as the first feature amount, a feature amount related to the behavior of the corresponding individual. The behavior may include at least one of viewing behavior, purchasing behavior, answering behavior for a questionnaire, and online behavior. The information processing system configured in this way can generate combined data useful for analyzing human behavior.
[0028] According to one aspect of the present disclosure, when the action includes a response action to a question, the first feature amount may represent the response of the corresponding individual to the question. The statistical value may correspond to the ratio of the individuals among one or more individuals who gave a specific response to the question. The information processing system configured in this way can generate combined data useful for analyzing the behavior of people based on a questionnaire survey.
[0029] According to one aspect of the present disclosure, when the action includes a viewing action, the first feature amount may represent the presence or absence of viewing or the viewing amount of the object by the corresponding individual. The statistical value may correspond to the number or ratio of the individuals among one or more individuals who viewed the object, the number or ratio of the individuals among one or more individuals who viewed the object more than a reference amount, or a representative value of the viewing amount for one or more individuals. The information processing system configured in this way can generate combined data useful for analyzing the viewing action.
[0030] According to one aspect of the present disclosure, when the action includes a purchasing action, the first feature amount may represent the presence or absence of purchasing or the purchasing amount of the object by the corresponding individual. The statistical value may correspond to the number or ratio of the individuals among one or more individuals who purchased the object, the number or ratio of the individuals among one or more individuals who purchased the object more than a reference amount, or a representative value of the purchasing amount for one or more individuals. The information processing system configured in this way can generate combined data useful for analyzing the purchasing action.
[0031] According to one aspect of the present disclosure, the second explanatory data may be data for explaining the characteristics of a plurality of individuals belonging to the second set. In this case, the second explanatory data may include second feature data for each individual. The second feature data may describe the second feature amount of the corresponding individual in association with the attributes of the corresponding individual and information on the area to which the corresponding individual belongs. The information processing system configured in this way can expand the individual data with statistical data and enrich the information.
[0032] According to one aspect of the present disclosure, when combining first feature data and second feature data, the combining part may be configured to associate a weight-back value for each combination of area and attribute regarding a plurality of areas and a plurality of attributes with the combined data. The weight-back value is based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute in the corresponding area of the second explanatory data.
[0033] There may be statistical bias in the combined data due to the difference between the composition of the population and the composition of the sample. According to the information processing system configured in this way, the combined data can be generated so that analysis considering the difference between the composition of the population and the composition of the sample is possible. This information processing system can generate combined data useful for analyzing the population.
[0034] According to one aspect of the present disclosure, the weight-back value may correspond to a value obtained by dividing the population of the attribute corresponding to the combination in the area corresponding to the combination by the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data. The information processing system may hold the weight-back value together with the population information of the population. The population information may include information that can identify the total population in a plurality of areas. The information processing system configured in this way is useful for data analysis considering the population.
[0035] According to one aspect of the present disclosure, when the first explanatory data describes, as a first feature amount, a feature amount related to the behavior of a corresponding individual, the generation unit may be configured to calculate, for each combination of area and attribute, an estimated number of persons who take an action in which the first feature amount satisfies a specific condition in the corresponding area.
[0036] According to one aspect of the present disclosure, the estimated number can be calculated based on the ratio of individuals among one or more individuals having the corresponding attribute in the corresponding area whose first feature amount satisfies a specific condition and the population of the corresponding attribute in the corresponding area. The generation unit may generate first feature data describing the estimated number as a statistical value.
[0037] According to the information processing system configured as described above, a user can obtain information regarding the number of persons who satisfy a specific condition in the population based on the estimation from the sample. According to one aspect of the present disclosure, when the action includes an answering action to a question, the estimated number can be the estimated number of respondents who give a specific answer to the question.
[0038] According to one aspect of the present disclosure, the information processing system may include an output unit. When the statistical value corresponds to the ratio of individuals among one or more individuals whose first feature quantity satisfies a specific condition, the output unit, based on the statistical value and the population for one or more combinations, on the condition that one or more combinations are specified, with respect to the combination of an area and an attribute, may be configured to output at least one of the ratio of individuals whose first feature quantity satisfies the specific condition and the total number of individuals whose first feature quantity satisfies the specific condition in a group corresponding to one or more combinations of the population.
[0039] According to one aspect of the present disclosure, the one or more combinations may be specified by specifying conditions regarding at least one of an area, an attribute, a first feature quantity, and a second feature quantity.
[0040] According to one aspect of the present disclosure, the information processing system can provide meaningful information regarding the population desired by the user about persons whose first feature quantity satisfies a specific condition, based on, for example, the above-mentioned specification from the user.
[0041] According to one aspect of the present disclosure, an information processing method corresponding to the above-described information processing system may be provided. The information processing method can be executed by a computer.
[0042] According to one aspect of the present disclosure, the information processing method may include obtaining first explanatory data that describes the features of a plurality of individuals belonging to a first set. The first explanatory data may describe, for each individual, the first feature quantity of the corresponding individual in association with the information of the attribute of the corresponding individual and the area to which the corresponding individual belongs.
[0043] The information processing method may include generating statistical data by statistically processing first explanatory data. The statistical data may include, for each of a plurality of areas, area statistical data regarding a set of individuals belonging to the corresponding area.
[0044] The area statistical data may include, for a plurality of attributes, first feature data for each attribute. The first feature data may describe a statistical value of a first feature quantity regarding one or more individuals calculated by statistically processing the first feature quantities of one or more individuals having the corresponding attribute among the set of individuals within the corresponding area.
[0045] The information processing method may further include obtaining second explanatory data for explaining the characteristics of a plurality of people belonging to a second set. The second explanatory data may include second feature data for each person. The second feature data may describe the second feature quantity of the corresponding person in association with the attributes of the corresponding person and information on the area to which the corresponding person belongs.
[0046] The information processing method may include combining the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data in which the combination of the attribute and the area matches.
[0047] According to one aspect of the present disclosure, generating may include, for each attribute, when there is an empty area, which is an area in the first set where there are no individuals having the corresponding attribute, among the plurality of areas, performing statistical processing on the first feature quantities of one or more individuals belonging to one or more peripheral areas located around the empty area among the plurality of areas, thereby calculating a statistical value of the corresponding attribute for the empty area and generating area statistical data for the empty area.
[0048] According to one aspect of the present disclosure, generating may include performing statistical processing based on, for each combination of an area and an attribute with respect to a plurality of areas and a plurality of attributes, a first feature amount of one or more individuals having the corresponding attribute in the corresponding area and a first feature amount of one or more individuals belonging to one or more peripheral areas that are one or more areas located around the corresponding area, to calculate a statistical value of the corresponding attribute for the corresponding area.
[0049] These information processing methods have the same effects as the information processing system described above. According to one aspect of the present disclosure, a computer program for causing a computer to execute the above-described information processing method may be provided. The computer program may be recorded on a non-transitory computer-readable recording medium.
Brief Description of the Drawings
[0050]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Mode for Carrying Out the Invention
[0051] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. [First Embodiment] The information processing system 10 of the present embodiment shown in FIG. 1 combines a first living person table L1 that describes the characteristics of each individual regarding a plurality of individuals belonging to a first set, and a second living person table L2 that describes the characteristics of each person regarding a plurality of people belonging to a second set after processing the first living person table L1, so as to expand the second living person table L2.
[0052] The first set can be, for example, a set of individuals targeted for information collection by a first means, or a set of customers in a first company. The second set can be a set of individuals targeted for information collection by a second means, or a set of customers in a second company. The second means can be different from the first means. The second company can be different from the first company. The generation of a table rich in information by combination is useful for analyzing at least one of the behavior and awareness of living persons.
[0053] As shown in FIG. 1, the information processing system 10 includes a processor 11, a memory 13, a storage 15, a user interface 17, and a communication interface 19. The processor 11 is configured to execute processing according to a computer program stored in the storage 15.
[0054] The memory 13 is a main memory device and is used as a working memory when the processor 11 executes processing. The storage 15 is an auxiliary storage device such as a hard disk drive and a solid state drive. The storage 15 is configured to store various data used for the processing executed by the processor 11 in addition to computer programs.
[0055] The user interface 17 includes a display unit 17A and an operation unit 17B. The display unit 17A is controlled by the processor 11 and is configured to display information for an operator who operates the information processing system 10.
[0056] The display unit 17A includes one or more display devices such as a liquid crystal display and an organic EL display. The operation unit 17B is configured to input an operation signal from the operator to the processor 11. The operation unit 17B includes one or more input devices such as a keyboard and a pointing device.
[0057] The communication interface 19 is configured to be communicable with an external device connected to a wide area network. The information processing system 10 can acquire necessary data from an external server through the communication interface 19.
[0058] Based on an instruction from the operator input through the user interface 17, the processor 11 executes the table processing shown in FIG. 2. As a result, the processor 11 converts the feature amount for each individual described in the first resident table L1 into a statistical value for each segment (details will be described later), and generates a statistical table Ls1 which is a table obtained by statistically processing the information in the first resident table L1.
[0059] An exemplary first resident table L1 is shown in the upper part of FIG. 3. As can be understood from FIG. 3, the first resident table L1 includes a record for each individual (hereinafter referred to as "personal ticket data") regarding a plurality of individuals belonging to the first set.
[0060] The individual data describes the area and basic attributes of the corresponding individual in association with the identification code (i.e., ID) of the corresponding individual. Each individual data further describes the feature quantities X1, ..., X2 of the corresponding individual regarding other items 1, ..., M in association with information describing the area and basic attributes of the corresponding individual. M where M is a natural number greater than or equal to 1.
[0061] The area of the corresponding individual is the area to which the corresponding individual belongs, specifically the area where the corresponding individual lives. When the first set is a sample set of residents in Japan, each of the multiple areas may be a region defined by dividing Japan into municipalities. The individual data may include information on the prefecture and municipality where the corresponding individual lives as information describing the area of the corresponding individual.
[0062] The basic attributes include gender and age group. The age group is defined by dividing the age into predetermined intervals. For example, the age group may be defined by dividing the age into 10-year intervals, such as "20s" and "30s." The predetermined interval may be 1-year intervals or 5-year intervals, and the age group with 1-year intervals is the same as the age.
[0063] The features X1,…,X for items 1,…,M shown in the top row of Figure 3 M Each of the variables is a variable that indicates a value of 1 when the corresponding individual has the corresponding characteristic, and a value of 0 when the corresponding individual does not have the corresponding characteristic.
[0064] As a first example, the first consumer table L1 contains features X1, ..., X obtained from a questionnaire survey. M In this case, the feature value X of each item m (m=1, 2, ..., M) m is a variable that indicates a value of 1 when the corresponding individual answers the mth (m=1, 2, ..., M) question included in the questionnaire in a positive manner, and a value of 0 when the corresponding individual answers in the negative manner. In other words, the feature quantity X m can be a variable representing an individual's answer to the m-th question. The unindexed feature X described below is the feature Xm is a simplified expression.
[0065] As a second example, the first lifestyle table L1 can describe the feature quantities X1, …, X obtained from a survey of purchasing behavior. M In this case, the feature quantity X can be a variable that indicates a value of 1 when the purchase quantity of the m-th (m = 1, 2, …, M) product by the corresponding individual exceeds a reference quantity during a past predetermined period, and indicates a value of 0 when the purchase quantity is less than or equal to the reference quantity. The purchase quantity can be, but is not limited to, the number of purchases or the purchase amount. The reference quantity can be defined as a value of zero or more. When the reference quantity is zero, the feature quantity X represents the presence or absence of the purchase of the m-th product during a past predetermined period.
[0066] As a third example, the first lifestyle table L1 can describe the feature quantities X1, …, X obtained from a survey of viewing behavior. M In this case, the feature quantity X can be a variable that indicates a value of 1 when the viewing quantity of the m-th (m = 1, 2, …, M) program by the corresponding individual exceeds a reference quantity, and indicates a value of 0 when the viewing quantity is less than or equal to the reference quantity. The viewing quantity can be, but is not limited to, the viewing time. The reference quantity can be predetermined as a value of zero or more. The program can be, for example, a television broadcast program. When the reference quantity is zero, the feature quantity X represents the presence or absence of viewing of program m.
[0067] However, the first lifestyle table L1 is not limited to the first to third examples described above. The first lifestyle table L1 can be a table that describes feature quantities X1, …, X related to the behavior of lifestyle users not limited to, for example, response behavior to questions, purchasing behavior, and viewing behavior. The behavior referred to here includes online behavior and offline behavior. The first lifestyle table L1 may include feature quantities X1, …, X related to the awareness of lifestyle users. M Here, the behavior includes online behavior and offline behavior. The first lifestyle table L1 may include feature quantities X1, …, X related to the awareness of lifestyle users. M may be included.
[0068] The feature quantity X does not have to be a variable represented by two values of "0" or "1". For example, the feature quantity X may be a variable indicating the purchase quantity of the m-th product by the corresponding individual in a past predetermined period. The feature quantity X may be a variable indicating the viewing quantity of the m-th program by the corresponding individual.
[0069] The lower part of FIG. 3 shows the configuration of the statistical table Ls1 generated by the table processing (see FIG. 2) for the first-living-person table L1 illustrated in the upper part of FIG. 3. The statistical table Ls1 has records (hereinafter referred to as "individual ticket statistical data") for each combination of area and basic attributes. A group of individual ticket statistical data is generated by statistical processing on the first-living-person table L1.
[0070] Specifically, each individual ticket statistical data is generated by statistical processing on a group of records of individuals belonging to the corresponding area and basic attributes. The combination of area and basic attributes here is a combination of area, age group, and gender.
[0071] Each individual ticket statistical data is for the feature quantities X1,..., X of the set of individuals having the corresponding basic attributes (gender and age group) in the corresponding area among the first set corresponding to the first-living-person table L1. M It is generated by statistical processing.
[0072] Hereinafter, the set of individuals classified by the combination of area and basic attributes is expressed as a "segment". The individual ticket statistical data for each combination of area and basic attributes is the individual ticket statistical data for each segment. The "group of individual ticket statistical data" for each area corresponds to the statistical data regarding the set of individuals belonging to the corresponding area, that is, the area statistical data.
[0073] The individual ticket statistical data describes the feature quantity Ys of each item m (m = 1, 2, …, M), together with the information explaining the area and basic attributes of the corresponding segment. The area of the corresponding segment is the area to which the corresponding segment belongs, specifically, the living area of the corresponding segment. The segment feature quantity Ys of item m is the statistical value of the feature quantity X of item m for the corresponding segment.
[0074] When an execution instruction for table processing is input through the user interface 17, the processor 11 executes the table processing shown in FIG. 2. The execution instruction designates the first resident table L1 to be processed.
[0075] When starting the table processing, the processor 11 acquires the first resident table L1 designated by the execution instruction (S110). The processor 11 can acquire the designated first resident table L1 by reading it from the storage 15.
[0076] In the subsequent S120, the processor 11 selects one from among a plurality of areas defining the segment and one from among a plurality of basic attributes defining the segment, thereby selecting the processing target segment, that is, the combination of the processing target area and basic attributes (S120). Here, the area and basic attributes selected as the processing target are expressed as area z and basic attribute r.
[0077] In the subsequent S130, the processor 11 calculates the feature quantity Y m for each item m (m = 1, 2, …, M) regarding the processing target segment by referring to a group of individual ticket data corresponding to the processing target segment.
[0078] The feature quantity Y calculated in S130 mis the statistical value, specifically the average, of the feature quantity X of one or more individuals belonging to the processing target segment among a plurality of individuals belonging to the first set. The calculation of the average corresponds to an example of statistical processing. The average of the feature quantity X corresponds to an example of the representative value of the feature quantity X.
[0079] That is, the feature quantity Y m is the average of the feature quantity X of the corresponding item m of one or more individuals having the basic attribute r, which is a set of individuals within the area z. The individuals within the area z are individuals for whom the area z is the living area. The feature quantity Y without an index used below is a simplified expression of the feature quantity Y. m is a simplified expression of.
[0080] As shown in FIG. 4, when the feature quantity X is represented by two values, the feature quantity Y as the average of the feature quantity X in a certain segment corresponds to the ratio of individuals for whom the feature quantity X is the value 1 in the same segment. That is, the feature quantity Y corresponds to the ratio of individuals for whom the feature regarding the item m satisfies a specific condition in one segment.
[0081] When the feature quantity X is the feature quantity in the first example described above, the feature quantity Y corresponds to the ratio of individuals who answered "yes" in response to the m-th question asking "yes" or "no".
[0082] When the feature quantity X is the feature quantity in the second example described above, the feature quantity Y corresponds to the ratio of individuals who purchased more than the reference quantity of the m-th product in the past predetermined period. When the reference quantity is zero, the feature quantity Y corresponds to the ratio of individuals who purchased the m-th product in the past predetermined period. When the feature quantity X is represented by the purchase quantity, the feature quantity Y is the average of the purchase quantity.
[0083] When the feature quantity X is the feature quantity in the third example described above, the feature quantity Y corresponds to the ratio of individuals who watched more than the reference quantity of the m-th program. When the reference quantity is zero, the feature quantity Y corresponds to the ratio of individuals who watched the program. When the feature quantity X is represented by the viewing quantity, the feature quantity Y is the average of the viewing quantity.
[0084] In subsequent S140, the processor 11 determines the feature amount Y for each calculated item (i.e., the average of the feature amounts X) as the segment feature amount Ys of the corresponding item. That is, for m = 1, 2, …, M, the processor 11 determines the feature amount Y of item m as the segment feature amount Ys of item m.
[0085] In subsequent S150, the processor 11 determines whether all combinations of the area and the basic attribute have been selected as the processing target. If it is determined that not all combinations have been selected (No in S150), the processor 11 changes the processing target segment, that is, the processing target area z and the basic attribute r, in S120, and executes the processing after S130. In this way, for each combination of the area and the basic attribute, the processor 11 determines the segment feature amount Ys of each item m (m = 1, 2, …, M) to be described in the individual ticket statistical data.
[0086] If it is determined in S150 that all combinations have been selected (Yes in S150), the processor 11 generates and outputs a statistical table Ls1 including the individual ticket statistical data for each segment, that is, for each combination of the area and the basic attribute, in S160.
[0087] The individual ticket statistical data for each segment in the generated statistical table Ls1 describes the segment feature amount Ys of each item m (m = 1, 2, …, M) in association with the information explaining the corresponding area and basic attribute.
[0088] The output destination of the statistical table Ls1 is, for example, the storage 15. That is, the generated statistical table Ls1 is stored in the storage 15. After executing the processing of S160, the processor 11 ends the table processing shown in FIG. 2. In this way, the processor 11 converts the first resident table L1 into the statistical table Ls1.
[0089] Further, the processor 11 executes the association related processing shown in FIG. 5 according to an execution instruction from an operator input through the user interface 17, thereby associating the statistical table Ls1 with the second life subject table L2 specified by the operator. As a result, the processor 11 generates an association table Le2.
[0090] The association table Le2 is a table obtained by associating the statistical table Ls1 with the second life subject table L2. The association table Le2 corresponds to a table obtained by expanding the second life subject table L2 using the statistical table Ls1. The association here includes associating a part of the information included in the statistical table Ls1 with the second life subject table L2.
[0091] In the execution instruction for the association related processing, the second life subject table L2 and the statistical table Ls1 to be processed are specified. The specified second life subject table L2 can be a table having individual slip data for each individual regarding a plurality of individuals belonging to a second set, similar to the first life subject table L1.
[0092] Alternatively, the second life subject table L2 can be a table (see FIG. 6A) having records for explaining the characteristics of corresponding segments in association with the area and basic attribute information of the corresponding segments, not for each individual but for each segment.
[0093] Each of the segments here is a set of individuals having the same area and basic attributes among a plurality of individuals belonging to the second set. The term "person" as used in this specification should be understood as a term including, in addition to individuals, sets of individuals classified by segments.
[0094] FIG. 6A shows an example of the second life subject table L2 and the statistical table Ls1 to be associated. The statistical table Ls1 shown in FIG. 6A has records (individual slip statistical data) for explaining the characteristics of corresponding segments regarding questionnaire response behavior of the corresponding segments in association with the area and basic attribute information of the corresponding segments for each segment regarding the first set.
[0095] This statistical table Ls1 describes the segment feature amount Ys corresponding to the feature amount X for each question obtained by the questionnaire survey. The statistical table Ls1 describes, as a plurality of items, for M = M1 questions, the segment feature amount Ys for each question, for each segment, in association with the information on the corresponding area and basic attributes.
[0096] The second-lifer table L2 shown in FIG. 6A includes records that explain the characteristics related to the purchasing behavior of the corresponding segment, in association with the information on the area and basic attributes of the corresponding segment, for each segment as a person belonging to the second set.
[0097] Each record of the second-lifer table L2 describes, as a plurality of items, for M = M2 products, the feature amount for each product (in other words, the segment feature amount Ys), in association with the information on the corresponding area and basic attributes. The feature amount of each record is the feature amount of the corresponding segment. The feature amount for each product may indicate the average purchase amount of the corresponding product.
[0098] When starting the join-related process shown in FIG. 5, the processor 11 acquires the second-lifer table L2 specified in the execution instruction (S210). In subsequent S220, the processor 11 acquires the statistical table Ls1 specified in the execution instruction (S220).
[0099] Specifically, the processor 11 can acquire the second-lifer table L2 and the statistical table Ls1 by reading the second-lifer table L2 and the statistical table Ls1 from the storage 15 (S210, S220).
[0100] In subsequent S230, the processor 11 performs data fusion to join the specified statistical table Ls1 to the specified second-lifer table L2. That is, the processor 11 joins the statistical table Ls1 to the second-lifer table L2 so as to associate the records of the statistical table Ls1 with the same combination of area and basic attributes to the records of the specified second-lifer table L2.
[0101] As a result, the processor 11 generates a combined table Le2 by combining the statistical table Ls1 with the second resident table L2. As described above, the combined table Le2 corresponds to a table obtained by expanding each record of the second resident table L2 using the statistical table Ls1.
[0102] FIG. 6B shows an example of the combined table Le2 generated by combining the second resident table L2 and the statistical table Ls1 shown in FIG. 6A. The combined table Le2 shown in FIG. 6B is configured such that segment feature amounts Ys for each question indicated by the individual ticket statistical data of the corresponding area and basic attributes are added as extended data to each record of the second resident table L2.
[0103] That is, the combined table Le2 includes records that describe segment feature amounts Ys for each question and feature amounts for each product, in association with the information on the corresponding area and basic attributes, for each combination of area and basic attributes.
[0104] When generating the combined table Le2 in S230 (see FIG. 5), the processor 11 then outputs the combined table Le2 in S240 and ends the combination-related process. The output destination of the combined table Le2 in S240 can be the storage 15. That is, the processor 11 can store the combined table Le2 generated in S230 in the storage 15.
[0105] In the storage 15 to which the combined table Le2 is output, population information for each segment, that is, for each combination of area and basic attributes, is pre-stored as one table (hereinafter referred to as "population table") Lp for analysis considering the population of the combined table Le2.
[0106] The population table Lp shown in FIG. 7 includes records that describe the population of the corresponding age group and gender in the corresponding area for each combination of area, age group, and gender.
[0107] The above-mentioned first set and second set are, for example, sample sets when considering the set of individuals living in Japan as the population. The population table Lp describes the population of the population including the first set and the second set. The population table Lp can be created, for example, based on the population information provided by the Geospatial Information Authority of Japan in Japan.
[0108] When generating the combined table Le2, the processor 11 may attach the population information of the corresponding segment to each record of the combined table Le2. That is, each record of the combined table Le2 may be attached with population information that describes the population of residents of the corresponding age group and gender in the corresponding area (see FIG. 8).
[0109] By analyzing this combined table Le2, the operator of the information processing system 10 can convert the segment feature amount Ys for each segment into a feature amount considering the population distribution of the population, and analyze the features of each segment.
[0110] For example, as shown in FIG. 8, consider the case where the combined table Le2 describes the affirmative response rate for each question obtained by a questionnaire survey as the segment feature amount Ys. When the first lifestyle table L1 describes the feature amount X of the above-mentioned first example regarding the question, the segment feature amount Ys corresponding to the feature amount X is the ratio of the number of people who answered affirmatively to the corresponding question in the corresponding segment in the first set, that is, the affirmative response rate in the corresponding segment.
[0111] By multiplying the segment feature amount Ys indicating this affirmative response rate by the population P of the corresponding segment, it is possible to estimate the number of affirmative respondents, which is the number of individuals with the corresponding basic attributes living in the corresponding area in the population, who answered affirmatively to the question.
[0112] Figure 8 illustrates that when the positive response rate in the i-th (i = 1, 2, …) segment is Ys[i] and the population of the i-th segment is P[i], the number of positive respondents (estimated value) in the area and basic attributes corresponding to the segment can be calculated as Ys[i]·P[i]. Ys[i] represents the segment feature amount Ys of the i-th segment.
[0113] In this way, by associating population information with the combined table Le2, it is possible to estimate the number of positive respondents based on the population standard, which is convenient. However, the segment feature amount Ys is not limited to the positive response rate.
[0114] For example, the segment feature amount Ys corresponding to item m can represent the ratio of individuals whose features corresponding to item m satisfy specific conditions. In this case, by multiplying the segment feature amount Ys and the population P of the corresponding segment, in the population, among a plurality of individuals living in the corresponding area and having the corresponding basic attributes, the estimated number of individuals whose features corresponding to item m satisfy specific conditions can be calculated.
[0115] For example, the segment feature amount Ys corresponding to item m can represent the average purchase quantity of the product corresponding to item m. In this case, by multiplying the segment feature amount Ys and the population P of the corresponding segment, in the population, the estimated value of the total purchase quantity of the corresponding product by a plurality of individuals living in the corresponding area and having the corresponding basic attributes can be calculated.
[0116] In addition, Figure 8 shows that if the number of positive respondents Ys[i]·P[i] for each segment is used, the positive response rate Ys * of the target group G, which is a group of segments of interest among a plurality of segments, can be calculated based on the population standard.
[0117] For example, when the target group G is a group of the first segment, the second segment, the third segment, and the fourth segment, the sum Σ of the number of positive respondents Ys[i]·P[i] of the i-th (i = 1, 2, 3, 4) segmenti∈G Ys[i]·P[i] = (Ys[1]·P[1] + Ys[2]·P[2] + Ys[3]·P[3] + Ys[4]·P[4]) is divided by the total population Σ of the target group G i∈G by P[i] = (P[1] + P[2] + P[3] + P[4]), so as to obtain the positive response rate Ys of the target group G * = Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] can be calculated based on the population standard
[0118] Here, the variable i is the index of the segment as described above, and i ∈ G means the segment belonging to the target group G. Σ i∈G F[i] means calculating the sum of F[i] of each segment corresponding to the target group G. F[i] is, for example, Ys[i]·P[i] or P[i].
[0119] The information processing system 10 has such a function of calculating and outputting the number of positive respondents Σ i∈G Ys[i]·P[i] and the positive response rate Ys * = Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] according to the instruction from the operator. The positive response rate Ys * corresponds to the positive response rate of the target group G with guaranteed representativeness
[0120] For example, the processor 11 can realize this function by executing the response analysis process shown in FIG. 9 according to the instruction from the operator. When starting the response analysis process shown in FIG. 9, the processor 11 obtains from the operator information specifying the combined table Le2 to be processed through the user interface 17 (S310).
[0121] In the subsequent S320, the processor 11 obtains from the operator information specifying the target group G through the user interface 17. The operation of specifying the target group G corresponds to the operation of specifying one or more segments, and corresponds to the operation of specifying one or more combinations regarding the area and basic attributes
[0122] The target group G may be all combinations of area and basic attributes. The target group G may be a group specified only by area regardless of basic attributes. The target group G may be a group specified only by basic attributes regardless of area.
[0123] In subsequent S330, the processor 11 calculates, for each item m (m = 1, 2,..., M), the number of positive respondents Σ i∈G Ys[i]·P[i] in the target group G with reference to the combined table Le2. In subsequent S340, the processor 11 calculates the total population Σ i∈G P[i] of the target group G.
[0124] In subsequent S350, for each item m (m = 1, 2,..., M), the processor 11 calculates the positive response rate Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] of the target group G with guaranteed representativeness.
[0125] The positive response rate Ys * of item m corresponds to an estimated value of the proportion of individuals whose feature quantity X of item m satisfies a specific condition (X = 1) among the set of individuals belonging to one or more specified segments in the population. The number of positive respondents of item m corresponds to an estimated value of the total number of individuals whose feature quantity X of item m satisfies a specific condition (X = 1) among the set of individuals belonging to one or more specified segments in the population.
[0126] In subsequent S360, for each item m (m = 1, 2,..., M), the processor 11 outputs the positive response rate Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the number of positive respondents Σ i∈G Ys[i]·P[i].
[0127] The processor 11 calculates the positive response rate Ys * =Σi∈G Ys[i]·P[i] / Σ i∈G P[i] and the total number of affirmative respondents Σ i∈G Ys[i]·P[i] can be output to the operator through the display unit 17A.
[0128] The processor 11 calculates the affirmative response rate Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the total number of affirmative respondents Σ i∈G By storing the data file describing Ys[i]·P[i] in the storage 15, the corresponding information may be output. Then, the processor 11 ends the response analysis process.
[0129] The above-described response analysis process is a process of analyzing the combined table Le2 that describes the affirmative response rate as the segment feature amount Ys. However, this process can be generalized to the analysis of the combined table Le2 that describes parameters other than the affirmative response rate as the segment feature amount Ys.
[0130] When the segment feature amount Ys[i] of the i-th segment indicates the ratio of individuals whose characteristics regarding item m satisfy specific conditions in the same segment, the value Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the value Σ i∈G Ys[i]·P[i] respectively indicate the estimated ratio and number of individuals whose characteristics regarding item m satisfy specific conditions in the target group G of the population.
[0131] When the segment feature amount Ys[i] of the i-th segment represents the average purchase quantity of the m-th product in the same segment, the value Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the value Σ i∈G Ys[i]·P[i] respectively indicate the estimated average and total of the purchase quantity of the m-th product in the target group G of the population. The purchase quantity of the m-th product may be read as the viewing volume of the m-th program.
[0132] The target group G may be designated as a group with a purchase quantity equal to or greater than a specified quantity, or may be designated as a group with a viewing quantity equal to or greater than a specified quantity. In the combined table Le2, when a plurality of types of feature quantities such as the positive response rate, average purchase quantity, and average viewing quantity are described as the segment feature quantity Ys, a group that satisfies a plurality of conditions, such as a group in which the average purchase quantity satisfies a specific condition and the average viewing quantity satisfies a specific condition, may be designated as the target group G.
[0133] That is, among the segment feature quantities Ys of a plurality of items, a group in which one or more segment feature quantities Ys satisfy a specific condition may be designated as the target group G. In this case, the number Σ of the target group G in the population i∈G P[i] is the value Ys * = Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the value Σ i∈G Ys[i]·P[i] may be output together.
[0134] According to the information processing system 10 of the present embodiment described above, after statistically processing personal data, a plurality of data are combined using attributes and areas as keys. Therefore, it is possible to generate combined data in which a plurality of data are meaningfully combined, and combined data useful for analyzing consumers can be generated.
[0135] Data related to people often has information on attributes and areas attached. However, it is difficult to meaningfully combine the data of a plurality of individuals only with the information on attributes and areas. In contrast, according to the present embodiment, since personal data is statistically processed and combined with other data, a plurality of data can be meaningfully combined using a small amount of information such as attributes and areas as keys, and data with a large amount of information can be generated.
[0136] Furthermore, according to the information processing system 10 of the present embodiment, using the population table Tp, the feature amount of the sample set can be converted into the feature amount based on the population standard. Therefore, the information processing system 10 of the present embodiment is very useful for data analysis considering the population.
[0137] By the way, the above-described information processing system 10 may be modified as follows. That is, in S160, the processor 11 may generate the statistical table Ls12 shown in FIG. 10 instead of the above-described statistical table Ls1.
[0138] The statistical table Ls12 of the modification example includes individual ticket statistical data for each segment, similar to the above-described statistical table Ls1. The individual ticket statistical data for the i-th segment includes information explaining the population P[i] of the i-th segment in addition to the same information as the statistical table Ls1.
[0139] The individual ticket statistical data for the i-th segment further includes, for each item, the estimated number of individuals Ys[i]·P[i] of individuals whose feature amount X of the corresponding item m satisfies a specific condition (X = 1), appended to the segment feature amount Ys[i] indicating the ratio of individuals whose feature amount X of the corresponding item m satisfies the specific condition (X = 1).
[0140] That is, each individual ticket statistical data in the statistical table Ls12 of the modification example includes, as shown in FIG. 10, information on the population P[i] of the corresponding segment, the segment feature amount Ys[i] of each item in the corresponding segment, and information on the estimated number of individuals Ys[i]·P[i] where X = 1. The estimated number can be, for example, the estimated number of persons who take a specific action regarding answering behavior, purchasing behavior, viewing behavior, etc. The estimated number can be, for example, the estimated value of the number of affirmative respondents described above.
[0141] In S160, the processor 11 can calculate the estimated number of individuals Ys[i]·P[i] of each item m based on the population P[i] and the segment feature amount Ys[i] of each item m in the corresponding segment for each segment, and generate the statistical table Ls12 appended with these.
[0142] The same process may be executed in the process of S240 instead of S160. That is, in S240, the processor 11 can attach the population P[i] of the corresponding segment and the estimated population Ys[i]·P[i] of each item m to each record of the join table Le2. The information processing system 10 according to this modified example of generating the statistical table Ls12 and the join table Le2 is very useful for data analysis considering the population.
[0143] [Second Embodiment] Subsequently, the information processing system 10 of the second embodiment will be described. The information processing system 10 of the second embodiment is configured in the same manner as the first embodiment except that the table processing executed by the processor 11 is different. The hardware configuration of the information processing system 10 in the second embodiment is the same as that in the first embodiment. Therefore, hereinafter, the same components as those in the first embodiment included in the information processing system 10 of the second embodiment are denoted by the same reference numerals as those in the first embodiment, and the description thereof will be omitted.
[0144] According to this embodiment, the processor 11 executes the table processing shown in FIG. 11 instead of the table processing shown in FIG. 2. When starting the table processing shown in FIG. 11, the processor 11 acquires the first inhabitant table L1 specified by an execution instruction in the same manner as the process of S110 (S410). In subsequent S420, the processor 11 selects a processing target segment in the same manner as the process of S120 (S420).
[0145] That is, the processor 11 selects one area to be processed from a plurality of areas and one basic attribute to be processed from a plurality of basic attributes, thereby selecting a combination of the area to be processed and the basic attribute (S420).
[0146] In subsequent S430, for one or more peripheral areas that are one or more areas located around the area z to be processed, for each peripheral area, the processor 11 calculates a feature quantity Y, which is the average of the feature quantities X of one or more individuals belonging to the basic attribute r to be processed in the corresponding peripheral area.
[0147] Here, for each of the one or more peripheral areas located around the area z to be processed, a symbol z k (k = 1, 2, …, K) is attached. The one or more peripheral areas can be one or more areas adjacent to the area z to be processed.
[0148] When K areas are adjacent to the area z to be processed, the area z to be processed is adjacent to the peripheral areas z1, …, z K . Hereinafter, the feature quantity Y of the item m regarding one or more individuals having the basic attribute r in the peripheral area z k is referred to as the feature quantity Y m (z k , r) or Y(z k , r).
[0149] In S430, for each peripheral area z k (k = 1, 2, …, K), the feature quantity Y(z k , r) for each item is calculated. The calculation method of the feature quantity Y m (z k , r) of the m-th item is as described in the first embodiment.
[0150] That is, the feature quantity Y k of the peripheral area z regarding the basic attribute r m (z k , r) is calculated as the average of the feature quantities X of the corresponding item m of one or more individuals having the basic attribute r among the individuals within the area z k in the first set. The individuals within the area z k are individuals for whom the area z k is the living area.
[0151] In subsequent S440, the processor 11 determines the inverse of the distance d(z, z k (k = 1, 2, …, K)) between the area z to be processed and each surrounding area z k and calculates the weight W(z, z k ) = 1 / d(z, z k ) corresponding thereto. The distance d(z, z k ) between the area z to be processed and the surrounding area z k can be the distance between the representative point of the area z to be processed and the representative point of the surrounding area z k as shown in FIG. 12. Examples of representative points include the geometric center, geometric centroid, and population centroid.
[0152] In subsequent S450, for each item m (m = 1, 2, …, M), the processor 11 uses the inverse distance weighted (IDW) method to calculate the estimated feature amount Ya(z, r) of the segment to be processed in the area z by the weighted average of the feature amounts Y(z k , r) of one or more surrounding areas z k . The estimated feature amount Ya(z, r) is the weighted sum (i.e., weighted average) of the feature amounts Y(z k , r) using the weight W(z, z k ). The weighted average corresponds to an example of statistical processing, and the weighted sum corresponds to an example of a statistical value. Specifically, the estimated feature amount Ya(z, r) is calculated by the following formula.
[0153] Ya(z, r) = Σ k {W(z, z k ) · Y(z k , r)} / Σ k W(z, z k ) In subsequent S460, the processor 11 determines whether or not a sample of the segment to be processed exists in the area z to be processed. This determination corresponds to determining whether or not the area z to be processed is an empty area where no sample (individual) having the attribute r to be processed exists. When no sample of the segment to be processed exists in the area z to be processed, the area z to be processed is an empty area where no sample (individual) having the attribute r to be processed exists.
[0154] By determining whether the individual ticket data of one or more individuals belonging to the segment to be processed exists in the first-life table L1, the processor 11 can determine whether the sample of the segment to be processed exists in the processing target area z.
[0155] Among the multiple individuals belonging to the first set, the processor 11 searches within the first-life table L1 for the individuals within the processing target area z and having the basic attribute r, and can determine the presence or absence of the sample based on the presence or absence of the corresponding individuals.
[0156] When the processor 11 determines that the sample of the segment to be processed does not exist (No in S460), in the statistical table Ls1, the segment feature amount Ys of each item m (m = 1, 2,..., M) to be described in the individual ticket statistical data of the segment to be processed is determined as the estimated feature amount Ya(z, r) of each item m (m = 1, 2,..., M) calculated in S450 (S470), and the process proceeds to the process of S500.
[0157] On the other hand, when the processor 11 determines that the sample of the segment to be processed exists in the processing target area z (Yes in S460), for each item m (m = 1, 2,..., M), the observed feature amount Y m (z, r) is calculated (S480). Y(z, r) hereinafter is a simplified expression of Y m (z, r).
[0158] The observed feature amount Y(z, r) of the processing target area z is the feature amount Y of the item m regarding one or more individuals having the basic attribute r of the processing target in the processing target area z. That is, the observed feature amount Y(z, r) is the average of the feature amounts X of one or more individuals belonging to the basic attribute r of the processing target in the processing target area z.
[0159] In subsequent S490, the processor 11 determines, within the statistical table Ls1, the segment feature amount Ys of each item m (m = 1, 2, …, M) to be described in the individual ticket statistical data of the segment to be processed as the average {Ya(z, r) + Y(z, r)} / 2 of the estimated feature amount Ya(z, r) of the corresponding item m (m = 1, 2, …, M) calculated in S450 and the observed feature amount Y(z, r) of the corresponding m (m = 1, 2, …, M) calculated in S480.
[0160] In subsequent S500, the processor 11 determines, in the same manner as the process in S150, whether all combinations of the area and the basic attributes have been selected as the processing target. If it is determined that not all combinations have been selected (No in S500), the processor 11 changes the segment to be processed, that is, the area z and the basic attribute r of the processing target, in S420, and executes the processes after S430. In this way, the processor 11 determines the segment feature amount Ys of each item m (m = 1, 2, …, M) to be described in the individual ticket statistical data for each combination of the area and the basic attributes.
[0161] If it is determined in S500 that all combinations have been selected (Yes in S500), the processor 11 generates and outputs, in S510, a statistical table Ls1 having individual ticket statistical data for each segment, that is, for each combination of the area and the basic attributes, in the same manner as the process in S160. An example of the statistical table Ls1 is as shown in FIG. 3. In S510, a statistical table Ls12 shown in FIG. 10 may be generated and output.
[0162] The configurations of the statistical tables Ls1 and Ls12 generated in S510 are the same as those in the first embodiment except that the segment feature amount Ys of each item m (m = 1, 2, …, M) described in each individual ticket statistical data is the value determined in S470 and S490.
[0163] After executing the process in S510, the processor 11 ends the table processing shown in FIG. 11. In this way, the processor 11 converts the first resident table L1 into the statistical table Ls1.
[0164] According to this embodiment, the segment feature amount Ys for each segment is calculated and determined using the IDW method. Therefore, even when the number of samples in the first resident table L1 is small, the feature amount Y(z k of the peripheral area z k , r) can be used to accurately calculate the segment feature amount Ys of the processing target area z without omission, and it is possible to generate a highly reliable statistical table Ls1.
[0165] [Third Embodiment] Subsequently, the information processing system 10 of the third embodiment will be described. The information processing system 10 of the third embodiment is configured by adding new functions to the information processing system 10 of the first embodiment or the second embodiment. Therefore, hereinafter, the same components as those in the above embodiments included in the information processing system 10 of the third embodiment are denoted by the same reference numerals as those in the above embodiments, and the description thereof is omitted.
[0166] In the information processing system 10 of this embodiment, the processor 11 executes the association-related process shown in FIG. 5 as the first association-related process. The processor 11 is further configured to execute the second association-related process shown in FIG. 13 to combine the third resident table L3 with the association table Le2 generated by the first association-related process to expand the third resident table L3. Hereinafter, the table generated by expanding the third resident table L3 with the association table Le2 is referred to as an expanded table Le3.
[0167] As another example, the expanded table Le3 may be generated using the statistical table Ls1 instead of the association table Le2. That is, the processor 11 may combine the third resident table L3 with the statistical table Ls1 to generate the expanded table Le3. In this case, in the second association-related process described below, the statistical table Ls1 is used instead of the association table Le2. As a further alternative example, the statistical table Ls12 shown in FIG. 10 may be used instead of the statistical table Ls1 for generating the expanded table Le3.
[0168] Processor 11 executes the second association-related process shown in FIG. 13 in accordance with an execution instruction from an operator input through user interface 17. When starting the second association-related process, processor 11 acquires the third resident table L3 specified in the execution instruction (S610). In subsequent S620, processor 11 acquires the association table Le2 specified in the execution instruction. Thereafter, processor 11 executes the process of S630.
[0169] Processor 11 can acquire the third resident table L3 and the association table Le2 by reading the third resident table L3 and the association table Le2 from storage 15 (S610, S620).
[0170] As shown in FIG. 14, the third resident table L3 is a table that describes the characteristics of a plurality of individuals belonging to a third set and includes a record for each individual (hereinafter referred to as an "individual record").
[0171] Similar to the individual ticket data in the first resident table L1, the individual record describes the area and basic attributes (i.e., gender and age group) of the corresponding individual in association with the identification code (i.e., ID) of the corresponding individual.
[0172] Each individual record further describes other characteristics of the corresponding individual in association with the information on the area and basic attributes of the corresponding individual. The other characteristics include attributes of the individual other than the basic attributes (hereinafter referred to as "additional attributes") as shown in FIG. 14. The additional attribute shown in FIG. 14 is occupation. The additional attribute is a single feature quantity. Although not shown, the individual record may further include information on the feature quantity X for each item related to the behavior and / or awareness of the individual, similar to the individual ticket data in the first resident table L1.
[0173] The association table Le2 combined with the third resident table L3 configured in this way includes records for each segment (i.e., a combination of area and basic attributes), similar to the first embodiment.
[0174] Each record in the association table Le2 is associated with the area corresponding to the segment and information on the basic attributes, and includes information for explaining the characteristics of the set of individuals corresponding to the segment by the segment feature quantity Ys for each item, similar to the individual ticket statistical data of the first embodiment. However, the record for each segment in the association table Le2 may be a record showing the segment feature quantity Ys for each item calculated using the IDW method, similar to the second embodiment.
[0175] In S630, the processor 11 performs data fusion for associating the specified association table Le2 with the specified third party table L3. That is, the processor 11 associates the records of the specified third party table L3 with the records of the association table Le2 whose combinations of area and basic attributes match.
[0176] As a result, the processor 11 generates an extended table Le3 obtained by associating the association table Le2 with the third party table L3. The extended table Le3 corresponds to a table in which each record of the third party table L3 is data-expanded using the association table Le2.
[0177] In S640, the processor 11 outputs the extended table Le3 generated in S630 and ends the second association related process. The output destination of the extended table Le3 in S640 may be the storage 15. That is, the processor 11 can store the extended table Le3 generated in S630 in the storage 15.
[0178] FIG. 15 shows an example of the extended table Le3 generated by associating the third party table L3 and the association table Le2 shown in FIG. 14. The extended table Le3 shown in FIG. 15 is configured such that each record of the third party table L3 describes the segment feature quantity Ys for each item indicated by the corresponding records of the area and basic attributes included in the association table Le2. Hereinafter, each record included in the extended table Le3 is also referred to as an extended record.
[0179] The extended record further describes a weight-back (WB) value V that takes into account the population P of the corresponding segment and the number H of individuals belonging to the same segment in the third-party table L3. The weight-back value V of a certain segment is the value V = P / H obtained by dividing the population P of the segment by the number H of individuals belonging to the same segment in the third-party table L3. The extended table Le3 describes the weight-back value V for each segment.
[0180] This weight-back value V is used to perform an operation considering the segment composition between the population and the sample set. The segment composition here refers to the composition ratio of the segment, that is, the ratio of individuals belonging to each segment. The population here is the set of individuals living in Japan.
[0181] The sample set corresponds to the third set and is the set of individuals having records in the third-party table L3. The sample set is a subset of the set of individuals living in Japan (the population). The above-mentioned number H is the sample size of the corresponding segment in the sample set (the third set). Hereinafter, the number H will be expressed as the sample size H.
[0182] The processor 11 can determine the population P of each segment by referring to the population table Lp shown in FIG. 7. The processor 11 can determine the sample size H by counting the number of records of individuals belonging to the corresponding segment in the third-party table L3 or the extended table Le3 for each segment.
[0183] The weight-back value V is a value for each segment. As can be understood from FIG. 15, the same weight-back value V is described in the extended records of individuals belonging to the same segment.
[0184] By using this weight-back value V, the affirmative response rate Yu of each question described in the extended record can be adjusted to a value Yu considering the segment composition between the sample set and the population. *can be corrected. In the following, this value Yu * is referred to as the corrected response rate Yu * and expressed as such.
[0185] The positive response rate Yu is the segment feature quantity Ys for the corresponding question. This segment feature quantity Ys is described in the combined table Le2 combined with the third party table L3. In the combined table Le2, the positive response rate Yu for each question described as the segment feature quantity Ys is as described in the first embodiment or the second embodiment.
[0186] The positive response rate Yu for each question described in each extended record can be converted into the corrected response rate Yu * in a group of a plurality of segments of interest (hereinafter referred to as the "segment group of interest") according to the formula Yu * = Σ(Yu·V) / ΣV.
[0187] The positive response rate Yu for each question described in each extended record is the positive response rate for each question in one segment to which the corresponding third party belongs in the sample set (i.e., the third set). According to the above formula Yu * = Σ(Yu·V) / ΣV, such a positive response rate Yu for each question can be converted into the corrected response rate Yu * as the positive response rate for each question in the set of third parties corresponding to the segment group of interest in the population.
[0188] The numerator in the formula Yu * = Σ(Yu·V) / ΣV is the sum of (Yu·V) for a group of samples belonging to the segment group of interest in the sample set. Therefore, the numerator Σ(Yu·V) represents the number of positive respondents within the segment group of interest in the population. The denominator ΣV is the sum of the weight-back values V for a group of samples belonging to the segment group of interest. Therefore, the denominator ΣV represents the population of the segment group of interest in the population. From this explanation, the formula Yu *= The corrected response rate Yu for each question calculated by Σ(Yu·V) / ΣV * It can be understood that it is the affirmative response rate for each question in the set of consumers corresponding to the target segment group in the population.
[0189] The target segment group may be a group of all segments. In this case, the denominator ΣV corresponds to the total population of the population, and the corrected response rate Yu * is the corrected response rate Yu for each question in the population * and represents it.
[0190] The corrected response rate Yu * corresponds to the value obtained by correcting the affirmative response rate Yu in the sample set to the affirmative response rate of the population in consideration of the difference in segment composition between the sample set and the population in this way.
[0191] The operator of the information processing system 10 can calculate the corrected response rate Yu * based on the affirmative response rate Yu and the weight back value V described in each extended record in the extended table Le3.
[0192] According to the information processing system 10 configured in this way, personal data is expanded using statistical data to generate an extended table Le3 with rich information. Using this extended table Le3, the operator of the information processing system 10 can analyze the characteristics of consumers in detail, and at this time, it is possible to perform a meaningful feature analysis based on the population standard using the weight back value.
[0193] According to this embodiment, the processor 11 may execute the same processing as the processing shown in FIG. 9 regarding the calculation of the corrected response rate Yu * That is, the processor 11 may obtain information for designating a target segment group from the operator through the user interface 17 (S320). The processor 11 can calculate and output the corrected response rate Yu * for each question in the designated target segment group according to the above formula (S350, S360).
[0194] [Other Embodiments] The present disclosure is not limited to the above-described embodiments and can take various forms. For example, in the above-described embodiments, an example has been described in which a lifestyle table (first lifestyle table L1) obtained by a questionnaire survey and a lifestyle table (second lifestyle table L2) obtained by a purchase survey are handled. However, the technology of the present disclosure can be applied to various types of data related to human behavior and awareness.
[0195] In the above-described embodiments, as the feature quantity Y, the ratio of individuals in the corresponding segment who satisfy the specific condition for the feature quantity X and the average of the feature quantity X in the corresponding segment are adopted. However, the feature quantity Y may be the number of samples (individuals) in the corresponding segment who satisfy the specific condition for the feature quantity X.
[0196] For example, the feature quantity Y may be the number of individuals who answered affirmatively to the m-th question, the number of individuals who purchased the m-th product, the number of individuals who purchased the m-th product more than the reference amount, the number of individuals who watched the m-th program, or the number of individuals who watched the m-th program more than the reference amount in the corresponding segment. In this case, by attaching information on the number of samples in the corresponding segment to the feature quantity Y or the corresponding segment feature quantity Ys, it is possible to change the number of people indicated by the feature quantity Y to a ratio, and statistical tables Ls1 and Ls12 can be constructed.
[0197] The functions of one component in the above-described embodiments may be provided dispersedly in a plurality of components. The functions of a plurality of components may be integrated into one component. A part of the configuration of the above-described embodiments may be omitted. At least a part of the configuration of the above-described embodiments may be added to or replaced with the configuration of other above-described embodiments. All aspects included in the technical idea specified from the language described in the claims are embodiments of the present disclosure.
[0198] [Technical Idea Disclosed in this Specification] It can be understood that the following technical idea is disclosed in this specification. [Item 1] First explanatory data for explaining the characteristics of a plurality of individuals belonging to the first set, wherein for each individual, the first feature amount of the corresponding individual is described in association with the attribute of the corresponding individual and the information of the area to which the corresponding individual belongs, and a first acquisition unit configured to acquire the first explanatory data A generation unit that generates statistical data by statistically processing the first explanatory data, wherein the statistical data includes, for a plurality of areas, area statistical data for each area regarding the set of individuals belonging to the corresponding area, and the area statistical data includes, for a plurality of attributes, first feature data for each attribute, and the first feature data describes a statistical value of the first feature amount regarding the one or more individuals calculated by statistical processing on the one or more individuals having the corresponding attribute among the set of individuals within the corresponding area, and a generation unit configured as such Second explanatory data for explaining the characteristics of a plurality of people belonging to the second set, comprising second feature data for each person, and a second acquisition unit configured to acquire the second explanatory data, wherein the second feature data describes the second feature amount of the corresponding person in association with the attribute of the corresponding person and the information of the area to which the corresponding person belongs A combining unit configured to combine the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data in which the combination of the attribute and the area matches Comprising When there is an empty area, which is an area in which there are no individuals having the corresponding attribute in the first set, among the plurality of areas, the generation unit calculates the statistical value of the corresponding attribute for the empty area by performing statistical processing on the first feature amounts of one or more individuals belonging to one or more peripheral areas located around the empty area among the plurality of areas, and is configured to generate the area statistical data for the empty area An information processing system [Item 2] First explanatory data for explaining the characteristics of a plurality of individuals belonging to a first set, the first acquisition unit being configured to acquire, for each individual, the first explanatory data that describes the first feature amount of the corresponding individual in association with the attribute of the corresponding individual and the information of the area to which the corresponding individual belongs. A generation unit that generates statistical data by statistically processing the first explanatory data, the statistical data including, for each of a plurality of areas, area statistical data regarding the set of individuals belonging to the corresponding area, the area statistical data including, for each of a plurality of attributes, first feature data, and the first feature data being configured to describe a statistical value of the first feature amount regarding the one or more individuals calculated by statistical processing on the one or more individuals having the corresponding attribute among the set of individuals in the corresponding area. Second explanatory data for explaining the characteristics of a plurality of people belonging to a second set, the second acquisition unit being configured to acquire, for each person, the second explanatory data that includes second feature data for the corresponding person and describes the second feature amount of the corresponding person in association with the attribute of the corresponding person and the information of the area to which the corresponding person belongs. A combining unit configured to combine the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data in which the combination of the attribute and the area matches. Comprising The generation unit is configured to calculate the statistical value of the corresponding attribute for the corresponding area by performing statistical processing based on the first feature amount of one or more individuals having the corresponding attribute in the corresponding area and the first feature amount of one or more individuals belonging to one or more peripheral areas that are one or more areas located around the corresponding area, for each combination of area and attribute regarding the plurality of areas and the plurality of attributes. An information processing system. [Item 3] The information processing system according to Item 2, wherein When, with respect to the first set, the corresponding area is an empty area where there is no individual having the attribute corresponding to the corresponding area, the generation unit calculates the statistical value of the corresponding attribute for the empty area by performing statistical processing on the first feature amounts of one or more individuals belonging to one or more areas located around the empty area as the one or more peripheral areas. Information processing system. [Item 4] The information processing system according to Item 1, wherein the generation unit is configured to calculate a representative value of the first feature amount in the one or more individuals having the corresponding attribute in the corresponding area for each combination of the area and the attribute with respect to the plurality of areas and the plurality of attributes, and is configured to calculate the statistical value of the corresponding attribute by a weighted sum of the representative values in each of the one or more areas located around the empty area with respect to the empty area. [Item 5] The information processing system according to Item 1, wherein the generation unit is configured to calculate, for each combination of the area and the attribute with respect to the plurality of areas and the plurality of attributes, the ratio of individuals in the one or more individuals having the corresponding attribute in the corresponding area whose first feature amount satisfies a specific condition, and is configured to calculate the statistical value of the corresponding attribute by a weighted sum of the ratios in each of the one or more areas located around the empty area with respect to the empty area. [Item 6] The information processing system according to Item 2, wherein the generation unit, with respect to the plurality of areas and the plurality of attributes, for each combination of the area and the attribute, calculates a representative value of the first feature amount in the one or more individuals having the corresponding attribute in the corresponding area, An information processing system configured to calculate the statistical value by a weighted sum of the representative value in the corresponding area and the representative values in each of one or more areas located around the corresponding area, for each combination of the area and the attribute. [Item 7] The information processing system according to Item 2, wherein the generation unit, with respect to the plurality of areas and the plurality of attributes, for each combination of the area and the attribute, calculates a ratio of individuals among the one or more individuals having the corresponding attribute in the corresponding area, who satisfy a specific condition with respect to the first feature amount, and is configured to calculate the statistical value by a weighted sum of the ratio in the corresponding area and the ratios in each of one or more areas located around the corresponding area, for each combination of the area and the attribute. [Item 8] The information processing system according to any one of Items 1 to 7, wherein the generation unit calculates the statistical value using an inverse distance weighted (IDW) method. [Item 9] The information processing system according to any one of Items 4 to 7, wherein the generation unit is configured to calculate the statistical value by the weighted sum based on the inverse distance weighted (IDW) method. [Item 10] The information processing system according to any one of Items 1 to 3, wherein the first feature data describes, as the first feature amount, a feature amount related to at least one of the behavior and consciousness of the corresponding individual. [Item 11] The information processing system according to any one of Items 1 to 10, wherein the first feature data describes, as the first feature amount, a feature amount related to the behavior of the corresponding individual, The information processing system, wherein the action includes at least one of a viewing action, a purchasing action, an answering action for a question, and an online action. [Item 12] The information processing system according to Item 10, wherein the action includes an answering action for a question, the first feature amount represents an answer of the corresponding individual to the question, and the statistical value corresponds to a ratio of individuals among the one or more individuals who gave a specific answer to the question. [Item 13] The information processing system according to Item 10, wherein the action includes a viewing action, the first feature amount represents whether or not the corresponding individual has viewed the target or the viewing amount, and the statistical value corresponds to the number or ratio of individuals among the one or more individuals who have viewed the target, the number or ratio of individuals among the one or more individuals who have viewed the target equal to or more than a reference, or a representative value of the viewing amount for the one or more individuals. [Item 14] The information processing system according to Item 10, wherein the action includes a purchasing action, the first feature amount represents whether or not the corresponding individual has purchased the target or the purchase amount, and the statistical value corresponds to the number or ratio of individuals among the one or more individuals who have purchased the target, the number or ratio of individuals among the one or more individuals who have purchased the target equal to or more than a reference, or a representative value of the purchase amount for the one or more individuals. [Item 15] The information processing system according to any one of Items 1 to 14, wherein the second explanatory data is data for explaining characteristics of a plurality of individuals belonging to the second set, and each individual is provided with the second feature data, the second feature data describes the second feature amount of the corresponding individual in association with the attribute of the corresponding individual and information on the area to which the corresponding individual belongs, When combining the first feature data and the second feature data, the combining unit is configured to associate a weight-back value for each combination of area and attribute with respect to the plurality of areas and the plurality of attributes to the combined data. The weight-back value is an information processing system based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute in the corresponding area of the second explanatory data. [Item 16] The information processing system according to Item 15, The weight-back value is an information processing system corresponding to a value obtained by dividing the population of the attribute corresponding to the combination in the area corresponding to the combination by the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data. [Item 17] The information processing system according to any one of Items 1 to 16, The first explanatory data describes, as the first feature amount, a feature amount related to the behavior of the corresponding individual. For each combination of the area and the attribute, the generation unit calculates an estimated number of persons who take an action in which the first feature amount satisfies the specific condition in the corresponding area based on the ratio of the individuals among the one or more individuals having the corresponding attribute in the corresponding area whose first feature amount satisfies the specific condition and the population of the corresponding attribute in the corresponding area, and generates the first feature data describing the estimated number as the statistical value. [Item 18] The information processing system according to Item 17, The action is a response action to a questionnaire. The estimated number is an estimated number of respondents who give a specific answer to the questionnaire. [Item 19] The information processing system according to any one of Items 1 to 3, Item 5, and Item 7, The statistical value corresponds to the ratio of the individuals among the one or more individuals whose first feature amount satisfies a specific condition. The information processing system further Based on the statistical value and population for one or more combinations regarding the combination of the area and the attribute, on the condition that one or more combinations are specified, at least one of the ratio of individuals in the group corresponding to the one or more combinations of the population where the first feature quantity satisfies the specific condition and the total number of individuals where the first feature quantity satisfies the specific condition is output. An output unit configured to output An information processing system comprising [Item 20] The information processing system according to Item 19, An information processing system in which the one or more combinations are specified by specifying a condition regarding at least one of the area, the attribute, the first feature quantity, and the second feature quantity. [Item 21] An information processing method executed by a computer, Obtaining first explanatory data that describes the characteristics of a plurality of individuals belonging to a first set, and for each individual, associates the first feature quantity of the corresponding individual with information on the attribute of the corresponding individual and the area to which the corresponding individual belongs. First explanatory data to be described, Generating statistical data by statistically processing the first explanatory data, where the statistical data includes area statistical data for each area regarding a plurality of areas, and the area statistical data includes first feature data for each attribute regarding a plurality of attributes. The first feature data describes the statistical value of the first feature quantity regarding the one or more individuals calculated by statistical processing of the first feature quantity of one or more individuals having the corresponding attribute among the set of individuals in the corresponding area. Generating, Obtaining second explanatory data that describes the characteristics of a plurality of people belonging to a second set, and for each person, includes second feature data for the corresponding person, and the second feature data associates the second feature quantity of the corresponding person with information on the attribute of the corresponding person and the area to which the corresponding person belongs. Second explanatory data to be described, Combining the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data where the combinations of the attribute and the area match; including When generating, for each attribute, in the first set, there is an empty area among the plurality of areas that is an area where there is no individual having the corresponding attribute, and when there is such an empty area among the plurality of areas, by performing statistical processing on the first feature amounts of one or more individuals belonging to one or more peripheral areas located around the empty area among the plurality of areas, calculating the statistical value of the corresponding attribute for the empty area, and generating the area statistical data for the empty area; An information processing method. [Item 22] An information processing method executed by a computer, obtaining first explanatory data that describes the characteristics of a plurality of individuals belonging to a first set, where for each individual, the first feature amount of the corresponding individual is described in association with the information on the attribute of the corresponding individual and the area to which the corresponding individual belongs; generating statistical data by performing statistical processing on the first explanatory data, where the statistical data includes area statistical data for each area regarding a plurality of areas, the area statistical data includes first feature data for each of a plurality of attributes, and the first feature data describes the statistical value of the first feature amount regarding the one or more individuals calculated by performing statistical processing on the first feature amounts of one or more individuals having the corresponding attribute among the set of individuals in the corresponding area; obtaining second explanatory data that describes the characteristics of a plurality of people belonging to a second set, where the second feature data for each person is provided, and the second feature data describes the second feature amount of the corresponding person in association with the information on the attribute of the corresponding person and the area to which the corresponding person belongs; Combining the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data in which the combinations of the attribute and the area match. including The generating includes calculating the statistical value of the corresponding attribute for the corresponding area by performing statistical processing based on the first feature amount of one or more individuals having the corresponding attribute in the corresponding area and the first feature amount of one or more individuals belonging to one or more peripheral areas which are one or more areas located around the corresponding area, for each combination of area and attribute with respect to the plurality of areas and the plurality of attributes. An information processing method. [Item 23] A computer program for causing a computer to execute the information processing method according to Item 21 or Item 22.
Explanation of Signs
[0199] 10… Information processing system, 11… Processor, 13… Memory, 15… Storage, 17… User interface, 19… Communication interface, L1, L2, L3… Lifestyle table, Le2… Combined table, Le3… Extended table, Lp… Population table, Ls1, Ls12… Statistical table.
Claims
1. First explanatory data for explaining characteristics of a plurality of individuals belonging to a first set, the first acquisition unit being configured to acquire, for each individual, the first explanatory data that describes the first feature quantity of the corresponding individual in association with the attribute of the corresponding individual and information on the area to which the corresponding individual belongs; A generation unit that generates statistical data by statistically processing the first explanatory data, the statistical data including, for a plurality of areas, area statistical data for each area regarding the set of individuals belonging to the corresponding area, the area statistical data including, for a plurality of attributes, first feature data for each attribute, the first feature data being configured to describe a statistical value of the first feature quantity regarding the one or more individuals calculated by statistically processing the first feature quantities of one or more individuals having the corresponding attribute among the set of individuals within the corresponding area; Second explanatory data for explaining characteristics of a plurality of people belonging to a second set, the second explanatory data including second feature data for each person, the second acquisition unit being configured to acquire the second explanatory data that describes the second feature quantity of the corresponding person in association with the attribute of the corresponding person and information on the area to which the corresponding person belongs; A combining unit configured to combine the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data for which the combination of the attribute and the area matches; Comprising; When there is an empty area, which is an area in which there are no individuals having the corresponding attribute in the first set, among the plurality of areas, the generation unit executes statistical processing on the first feature quantities of one or more individuals belonging to one or more peripheral areas located around the empty area among the plurality of areas, thereby calculating the statistical value of the corresponding attribute for the empty area and generating the area statistical data for the empty area. An information processing system.
2. First explanatory data for explaining characteristics of a plurality of individuals belonging to a first set, the first acquisition unit being configured to acquire, for each individual, the first explanatory data that describes the first feature quantity of the corresponding individual in association with the attribute of the corresponding individual and information on the area to which the corresponding individual belongs; A generation unit that generates statistical data by statistically processing the first explanatory data, wherein the statistical data includes, for each of a plurality of areas, area statistical data regarding a set of individuals belonging to the corresponding area, the area statistical data includes, for each of a plurality of attributes, first feature data, and the first feature data is configured to describe a statistical value of the first feature quantity regarding the one or more individuals calculated by statistically processing the first feature quantity of the one or more individuals having the corresponding attribute among the set of individuals in the corresponding area. Second explanatory data for explaining the characteristics of a plurality of people belonging to a second set, including second feature data for each person, and a second acquisition unit configured to acquire the second explanatory data in which the second feature data describes the second feature quantity of the corresponding person in association with the attribute of the corresponding person and information on the area to which the corresponding person belongs. A combining unit configured to combine the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data in which the combination of the attribute and the area matches. Comprising The generation unit is configured to calculate the statistical value of the corresponding attribute for the corresponding area by performing statistical processing based on the first feature quantity of one or more individuals having the corresponding attribute in the corresponding area and the first feature quantity of one or more individuals belonging to one or more peripheral areas, which are one or more areas located around the corresponding area, for each combination of area and attribute with respect to the plurality of areas and the plurality of attributes. An information processing system.
3. The information processing system according to claim 2, The generation unit is configured to calculate the statistical value of the corresponding attribute for the empty area by performing statistical processing on the first feature quantity of one or more individuals belonging to one or more areas located around the empty area as the one or more peripheral areas when the corresponding area is an empty area in which there are no individuals having the corresponding attribute with respect to the first set. An information processing system.
4. The information processing system according to claim 1, The generation unit is configured to calculate a representative value of the first feature amount in the one or more individuals having the corresponding attribute in the corresponding area for each combination of an area and an attribute with respect to the plurality of areas and the plurality of attributes, and for the empty area, calculate the statistical value of the corresponding attribute by a weighted sum of the representative values in each of the one or more areas located around the empty area. An information processing system configured as such.
5. The information processing system according to claim 1, wherein the generation unit is configured to calculate, for each combination of an area and an attribute with respect to the plurality of areas and the plurality of attributes, a ratio of individuals in the one or more individuals having the corresponding attribute in the corresponding area, among whom the first feature amount satisfies a specific condition, and for the empty area, calculate the statistical value of the corresponding attribute by a weighted sum of the ratios in each of the one or more areas located around the empty area. An information processing system configured as such.
6. The information processing system according to claim 2, wherein the generation unit, with respect to the plurality of areas and the plurality of attributes, for each combination of the area and the attribute, calculates a representative value of the first feature amount in the one or more individuals having the corresponding attribute in the corresponding area, and for each combination of the area and the attribute, calculates the statistical value by a weighted sum of the representative value in the corresponding area and the representative values in each of the one or more areas located around the corresponding area. An information processing system configured as such.
7. The information processing system according to claim 2, wherein the generation unit, with respect to the plurality of areas and the plurality of attributes, for each combination of the area and the attribute, calculates a ratio of individuals in the one or more individuals having the corresponding attribute in the corresponding area, among whom the first feature amount satisfies a specific condition, and for each combination of the area and the attribute, calculates the statistical value by a weighted sum of the ratio in the corresponding area and the ratios in each of the one or more areas located around the corresponding area. An information processing system configured as such.
8. The information processing system according to any one of claims 1 to 7, wherein the generation unit calculates the statistical value using an inverse distance weighted (IDW) method. An information processing system configured as such.
9. An information processing system according to any one of Claims 4 to 7, wherein the generation unit is configured to calculate the statistical value by the weighted sum based on the inverse distance weighting (IDW) method.
10. An information processing system according to any one of Claims 1 to 3, wherein the first feature data describes, as the first feature amount, a feature amount related to at least one of the behavior and consciousness of the corresponding individual.
11. An information processing system according to any one of Claims 1 to 7, wherein the first feature data describes, as the first feature amount, a feature amount related to the behavior of the corresponding individual, and the behavior includes at least one of a viewing behavior, a purchasing behavior, an answering behavior to a questionnaire, and an online behavior.
12. The information processing system according to Claim 10, wherein the behavior includes an answering behavior to a questionnaire, the first feature amount represents the answer of the corresponding individual to the questionnaire, and the statistical value corresponds to the ratio of the individuals among the one or more individuals who gave a specific answer to the questionnaire.
13. The information processing system according to Claim 10, wherein the behavior includes a viewing behavior, the first feature amount represents the presence or absence of viewing or the viewing amount of the target by the corresponding individual, and the statistical value corresponds to the number or ratio of the individuals among the one or more individuals who viewed the target, the number or ratio of the individuals among the one or more individuals who viewed the target more than a reference, or a representative value of the viewing amount related to the one or more individuals.
14. The information processing system according to Claim 10, wherein the behavior includes a purchasing behavior, the first feature amount represents the presence or absence of purchase or the purchase amount of the target by the corresponding individual, and the statistical value corresponds to the number or ratio of the individuals among the one or more individuals who purchased the target, the number or ratio of the individuals among the one or more individuals who purchased the target more than a reference, or a representative value of the purchase amount related to the one or more individuals.
15. An information processing system according to any one of Claims 1 to 7, wherein the second explanatory data is data for explaining the characteristics of a plurality of individuals belonging to the second set, and each individual is provided with the second feature data. The second characteristic data describes the second characteristic amount of the corresponding individual in association with the attribute of the corresponding individual and information on the area to which the corresponding individual belongs. When combining the first characteristic data and the second characteristic data, the combining unit is configured to associate a weight-back value for each combination of an area and an attribute with respect to the plurality of areas and the plurality of attributes with the combined data. The weight-back value is an information processing system based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute in the corresponding area of the second explanatory data.
16. The information processing system according to claim 15, The weight-back value is an information processing system corresponding to a value obtained by dividing the population of the attribute corresponding to the combination in the area corresponding to the combination by the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data.
17. The information processing system according to any one of claims 1 to 7, The first explanatory data describes, as the first characteristic amount, a characteristic amount related to the behavior of the corresponding individual. For each combination of the area and the attribute, the generation unit calculates an estimated number of persons who take an action in which the first characteristic amount satisfies the specific condition in the corresponding area, based on the ratio of individuals among the one or more individuals having the corresponding attribute in the corresponding area whose first characteristic amount satisfies the specific condition and the population of the corresponding attribute in the corresponding area, and generates the first characteristic data describing the estimated number as the statistical value.
18. The information processing system according to claim 17, The action is a response action to a questionnaire. The estimated number is an estimated number of respondents who give a specific response to the questionnaire.
19. The information processing system according to any one of claims 1 to 3, claim 5, and claim 7, The statistical value corresponds to the ratio of individuals among the one or more individuals whose first characteristic amount satisfies a specific condition. The information processing system further Output unit configured to output at least one of the ratio of individuals in a group corresponding to one or more combinations in the population and the total number of individuals whose first feature quantity satisfies the specific condition, based on the statistical value and the population for the one or more combinations, on the condition that one or more combinations are specified for the combination of the area and the attribute. An information processing system comprising the same. **Claim 20** The information processing system according to claim 19, An information processing system in which the one or more combinations are specified by specifying a condition related to at least one of the area, the attribute, the first feature quantity, and the second feature quantity. **Claim 21** An information processing method executed by a computer, comprising: Obtaining first explanatory data for explaining the characteristics of a plurality of individuals belonging to a first set, the first explanatory data describing, for each individual, the first feature quantity of the corresponding individual in association with the information on the attribute of the corresponding individual and the area to which the corresponding individual belongs; Generating statistical data by statistically processing the first explanatory data, the statistical data including area statistical data for each area regarding a plurality of areas, the area statistical data including first feature data for each of a plurality of attributes, the first feature data describing the statistical value of the first feature quantity regarding the one or more individuals calculated by statistically processing the first feature quantity of the one or more individuals having the corresponding attribute among the set of individuals in the corresponding area; Obtaining second explanatory data for explaining the characteristics of a plurality of people belonging to a second set, the second explanatory data including second feature data for each person, the second feature data describing the second feature quantity of the corresponding person in association with the information on the attribute of the corresponding person and the area to which the corresponding person belongs; Combining the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data in which the combination of the attribute and the area matches; including The generating includes, for each of the attributes, when there is an empty area, which is an area in the first set where there is no individual having the corresponding attribute, among the plurality of areas, calculating a statistical value of the corresponding attribute for the empty area by performing statistical processing on the first feature amounts of one or more individuals belonging to one or more peripheral areas located around the empty area among the plurality of areas, and generating the area statistical data for the empty area. Information processing method.
22. An information processing method executed by a computer, comprising: obtaining first explanatory data for explaining characteristics of a plurality of individuals belonging to a first set, the first explanatory data describing, for each individual, a first feature amount of the corresponding individual in association with the attribute of the corresponding individual and information on the area to which the corresponding individual belongs; generating statistical data by performing statistical processing on the first explanatory data, the statistical data including, for a plurality of areas, area statistical data for each area regarding a set of individuals belonging to the corresponding area, the area statistical data including, for a plurality of attributes, first feature data for each attribute, the first feature data describing a statistical value of the first feature amount regarding the one or more individuals calculated by performing statistical processing on the first feature amounts of one or more individuals having the corresponding attribute among the set of individuals in the corresponding area; obtaining second explanatory data for explaining characteristics of a plurality of people belonging to a second set, the second explanatory data including second feature data for each person, the second feature data describing a second feature amount of the corresponding person in association with the attribute of the corresponding person and information on the area to which the corresponding person belongs; combining the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data where the combination of the attribute and the area matches; and including The generating includes calculating the statistical value of the corresponding attribute for the corresponding area by performing statistical processing based on, for each combination of an area and an attribute with respect to the plurality of areas and the plurality of attributes, a first feature amount of one or more individuals having the corresponding attribute in the corresponding area and the first feature amount of one or more individuals belonging to one or more peripheral areas that are one or more areas located around the corresponding area. An information processing method.
23. A computer program for causing a computer to execute the information processing method according to claim 21 or claim 22.
Citation Information
Patent Citations
Analysis method for marketing information, information processor and medium
JP2002140490A
General-purpose data fusion system and general-purpose data fusion method
JP5638675B1
Clean horse with cable
KR102813629B1
Information-processing system
WO2016021726A1
Power demand prediction device, power demand prediction method, and program therefor
WO2019207622A1