Information Processing System, Information Processing Method, and Computer Program
The information processing system addresses the challenge of combining data with differing population and sample compositions by using weight-back values to generate meaningful insights into human behavior and consciousness.
Patent Information
- Application Number
- JP2025042927
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2045-03-17
AI Technical Summary
Existing techniques struggle to meaningfully combine data related to individuals with attributes and areas due to statistical bias arising from differences between population and sample compositions, lacking methods to account for these differences in data analysis.
An information processing system that acquires and combines first and second explanatory data, generating statistical data with weight-back values to account for population and sample composition differences, allowing for meaningful data analysis by associating feature data with area and attribute combinations.
Enables data analysis that considers population and sample composition differences, providing accurate and meaningful insights into human behavior and consciousness by generating combined data with associated weight-back values.
Smart Images

Figure 0007717995000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing system and an information processing method.
Background Art
[0002] Conventionally, a technique for generating customer data with a large amount of information by combining multiple types of data related to customers is known. For example, a technique is known in which existing customer databases and external survey data are matched based on gender, age, and area as keys, and the information stored in the external survey data is added to the existing customer databases (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present inventors consider converting personal data into statistical data for each combination of attributes and areas by statistically processing first data that describes the characteristics of two or more corresponding individuals for each attribute and area regarding a plurality of individuals. The present inventors consider generating data with a large amount of information by combining the thus-generated statistical data for each attribute and area with second data that describes the characteristics of people for each attribute and area.
[0005] Data related to people often has information on attributes and areas attached. However, it is difficult to meaningfully combine the data of a plurality of individuals with only the information on attributes and areas. On the other hand, after statistically processing the personal data, it is possible to generate meaningful combined data by combining a plurality of data with attributes and areas as keys.
[0006] However, the combined data may contain statistical bias due to differences between the composition of the population and the composition of the sample. Techniques for performing an analysis that takes into account such statistical bias for combined data are not known.
[0007] Therefore, according to one aspect of the present disclosure, it is desirable to be able to provide a novel technique for realizing an analysis of combined data that takes into account differences between the composition of the population and the composition of the sample.
Means for Solving the Problems
[0008] According to one aspect of the present disclosure, an information processing system is provided. The information processing system includes a first acquisition unit, a generation unit, a second acquisition unit, and a combination unit.
[0009] The first acquisition unit is configured to acquire first explanatory data that describes the characteristics of a plurality of individuals belonging to a first set. The first explanatory data describes, for each individual, the first feature quantity of the corresponding individual in association with the attributes of the corresponding individual and information on the area to which the corresponding individual belongs.
[0010] The generation unit is configured to generate statistical data by statistically processing the first explanatory data. The statistical data includes, for a plurality of areas, area statistical data for each area regarding the set of individuals belonging to the corresponding area.
[0011] The area statistical data includes, for a plurality of attributes, first feature data for each attribute. The first feature data describes a statistical value of the first feature quantity regarding one or more individuals calculated by statistically processing the first feature quantities of one or more individuals having the corresponding attribute among the set of individuals within the corresponding area.
[0012] The second acquisition unit is configured to acquire second explanatory data that describes the characteristics of a plurality of people belonging to a second set. The second explanatory data includes second feature data for each person. The second feature data describes the second feature quantity of the corresponding person in association with the attributes of the corresponding person and information on the area to which the corresponding person belongs.
[0013] The combining unit is configured to combine statistical data based on first explanatory data and second explanatory data so as to associate first feature data and second feature data in which combinations of attributes and areas match.
[0014] According to one aspect of the present disclosure, when combining first feature data and second feature data, the combining unit may be configured to associate a weight-back value for each combination of area and attribute for a plurality of areas and a plurality of attributes with the combined data.
[0015] The weight-back value is based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data.
[0016] According to the information processing system configured in this way, it is possible to generate combined data so that analysis considering the difference between the population composition and the sample composition is possible.
[0017] According to one aspect of the present disclosure, the weight-back value may correspond to a value obtained by dividing the population of the attribute corresponding to the combination in the area corresponding to the combination by the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data. The information processing system may hold the weight-back value together with the population information of the population. The population information may include information capable of identifying the total population in a plurality of areas. The information processing system configured in this way is useful for data analysis considering the population.
[0018] According to one aspect of the present disclosure, the statistical value may correspond to a representative value of a first feature quantity regarding one or more individuals, or a ratio of individuals among one or more individuals whose first feature quantity satisfies a specific condition.
[0019] According to one aspect of the present disclosure, the first feature data can describe, as a first feature quantity, a feature quantity related to at least one of the behavior and consciousness of the corresponding individual. The information processing system configured in this way is useful for analyzing at least one of human behavior and consciousness.
[0020] According to one aspect of the present disclosure, the first feature data can describe, as a first feature quantity, a feature quantity related to the behavior of the corresponding individual. The behavior may include at least one of viewing behavior, purchasing behavior, answering behavior for questions, and online behavior. The information processing system configured in this way is useful for analyzing human behavior.
[0021] According to one aspect of the present disclosure, when the behavior includes answering behavior for questions, the first feature quantity may represent the answer of the corresponding individual to the question. The statistical value may correspond to the ratio of the individuals among one or more individuals who gave a specific answer to the question. The information processing system configured in this way is useful for analyzing human behavior based on a questionnaire survey.
[0022] According to one aspect of the present disclosure, when the behavior includes viewing behavior, the first feature quantity may represent the presence or absence of viewing or the viewing amount of the target by the corresponding individual. The statistical value may correspond to the number or ratio of individuals among one or more individuals who viewed the target, the number or ratio of individuals among one or more individuals who viewed the target above a reference, or a representative value of the viewing amount for one or more individuals. The information processing system configured in this way is useful for analyzing viewing behavior.
[0023] According to one aspect of the present disclosure, when the behavior includes purchasing behavior, the first feature quantity may represent the presence or absence of purchasing or the purchasing amount of the target by the corresponding individual. The statistical value may correspond to the number or ratio of individuals among one or more individuals who purchased the target, the number or ratio of individuals among one or more individuals who purchased the target above a reference, or a representative value of the purchasing amount for one or more individuals. The information processing system configured in this way is useful for analyzing purchasing behavior.
[0024] According to one aspect of the present disclosure, when the first feature data describes, as a first feature amount, a feature amount related to the behavior of a corresponding individual, the generation unit may be configured to calculate, for each combination of an area and an attribute, an estimated number of persons who take an action in the corresponding area where the first feature amount satisfies a specific condition.
[0025] According to one aspect of the present disclosure, the estimated number can be calculated based on the ratio of individuals among one or more individuals having the corresponding attribute in the corresponding area whose first feature amount satisfies a specific condition and the population of the corresponding attribute in the corresponding area. The generation unit may be configured to generate the first feature data describing the estimated number.
[0026] According to the information processing system configured as described above, the user can obtain information regarding the number of persons who satisfy a specific condition in the population based on the estimation from the sample. According to one aspect of the present disclosure, when the behavior includes a response behavior to a questionnaire, the estimated number may be the estimated number of respondents who give a specific response to the questionnaire.
[0027] According to one aspect of the present disclosure, the information processing system may include an output unit. When the statistical value corresponds to the ratio of individuals among one or more individuals whose first feature amount satisfies a specific condition, the output unit, on the condition that one or more combinations are specified, regarding the combination of the area and the attribute, based on the statistical value and the population for the one or more combinations, outputs at least one of the ratio of individuals whose first feature amount satisfies a specific condition and the total number of individuals whose first feature amount satisfies a specific condition in a group corresponding to one or more combinations of the population.
[0028] According to one aspect of the present disclosure, the information processing system can provide, as information desired by the user, based on, for example, the above-mentioned specification of the combination from the user, the ratio and / or the number of individuals who are individuals corresponding to the specified combination in the population and whose first feature amount satisfies a specific condition.
[0029] According to one aspect of the present disclosure, one or more combinations can be specified by specifying conditions related to at least one of an area, an attribute, a first feature amount, and a second feature amount. In this case, the information processing system can provide meaningful information about the population desired by the user, for example, based on the above specification from the user.
[0030] According to one aspect of the present disclosure, an information processing method corresponding to the above-described information processing system may be provided. The information processing method can be executed by a computer.
[0031] According to one aspect of the present disclosure, the information processing method may include obtaining first explanatory data that describes the characteristics of a plurality of individuals belonging to a first set. The first explanatory data may describe, for each individual, the first feature amount of the corresponding individual in association with the attribute of the corresponding individual and information on the area to which the corresponding individual belongs.
[0032] The information processing method may include generating statistical data by statistically processing the first explanatory data. The statistical data may include, for a plurality of areas, area statistical data for each area regarding the set of individuals belonging to the corresponding area.
[0033] The area statistical data may include, for a plurality of attributes, first feature data for each attribute. The first feature data may describe statistical values of the first feature amounts regarding one or more individuals calculated by statistically processing the first feature amounts of one or more individuals having the corresponding attribute among the set of individuals within the corresponding area.
[0034] The information processing method may include obtaining second explanatory data that describes the characteristics of a plurality of people belonging to a second set. The second explanatory data may include second feature data for each person. The second feature data may describe the second feature amount of the corresponding person in association with the attribute of the corresponding person and information on the area to which the corresponding person belongs.
[0035] The information processing method may include combining statistical data based on first explanatory data and second explanatory data so as to associate first feature data and second feature data in which combinations of attributes and areas match each other.
[0036] According to one aspect of the present disclosure, combining may include associating, with combined data, weight-back values for each combination of an area and an attribute for a plurality of areas and a plurality of attributes when combining the first feature data and the second feature data. The weight-back value is based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data.
[0037] This information processing method has the same effect as the above-described information processing system. According to one aspect of the present disclosure, there may be provided a computer program for causing a computer to execute the above-described information processing method. The computer program may be recorded on a non-transitory computer-readable recording medium.
Brief Description of Drawings
[0038]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
[0039] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. [First Embodiment] The information processing system 10 of the present embodiment shown in FIG. 1 is configured to expand the second life table L2 by combining a first life table L1 that describes the characteristics of each individual regarding a plurality of individuals belonging to a first set and a second life table L2 that describes the characteristics of each person regarding a plurality of people belonging to a second set after processing the first life table L1.
[0040] The first set can be, for example, a set of individuals who are the target of information collection by a first means, or a set of customers in a first company. The second set can be a set of individuals who are the target of information collection by a second means, or a set of customers in a second company. The second means can be a means different from the first means. The second company can be a company different from the first company. The generation of a table with abundant information by combination is useful for analyzing at least one of the behavior and consciousness of the consumers.
[0041] As shown in FIG. 1, the information processing system 10 includes a processor 11, a memory 13, a storage 15, a user interface 17, and a communication interface 19. The processor 11 is configured to execute processing according to a computer program stored in the storage 15.
[0042] The memory 13 is a main memory device and is used as a working memory when the processor 11 executes processing. The storage 15 is an auxiliary storage device such as a hard disk drive and a solid state drive. The storage 15 is configured to store various data used for the processing executed by the processor 11 in addition to the computer program.
[0043] The user interface 17 includes a display unit 17A and an operation unit 17B. The display unit 17A is controlled by the processor 11 and is configured to display information for an operator who operates the information processing system 10.
[0044] The display unit 17A includes one or more display devices such as a liquid crystal display and an organic EL display. The operation unit 17B is configured to input an operation signal from the operator to the processor 11. The operation unit 17B includes one or more input devices such as a keyboard and a pointing device.
[0045] The communication interface 19 is configured to be communicable with an external device connected to a wide area network. The information processing system 10 can acquire necessary data from an external server through the communication interface 19.
[0046] Based on an instruction from the operator input through the user interface 17, the processor 11 executes the table processing shown in FIG. 2. As a result, the processor 11 converts the feature amount for each individual described in the first life table L1 into a statistical value for each segment (details will be described later), and generates a statistical table Ls1 which is a table obtained by statistically processing the information in the first life table L1.
[0047] In the upper part of FIG. 3, an exemplary first life-person table L1 is shown. As can be understood from FIG. 3, the first life-person table L1 includes individual records (hereinafter referred to as "personal ticket data") for a plurality of individuals belonging to the first set.
[0048] The personal ticket data describes the area and basic attributes of the corresponding individual in association with the identification code (i.e., ID) of the corresponding individual. Each personal ticket data further includes, in association with information explaining the area and basic attributes of the corresponding individual, feature quantities X1,..., X M regarding other items 1,..., M of the corresponding individual. M is a natural number of 1 or more.
[0049] The area of the corresponding individual is the area to which the corresponding individual belongs, specifically the living area of the corresponding individual. When the first set is a sample set of domestic residents in Japan, each of the plurality of areas can be a region defined by partitioning Japan at the granularity of municipalities. The personal ticket data can include information on the prefecture and municipality of the region where the corresponding individual lives as information explaining the area of the corresponding individual.
[0050] The basic attributes include gender and age group. The age group is defined by dividing age at predetermined intervals. For example, the age group can be defined by dividing age at 10-year intervals, such as "in their 20s" and "in their 30s". The predetermined interval may be at 1-year intervals or 5-year intervals, and the age group at 1-year intervals is the same as the age.
[0051] The feature quantities X1,..., X M regarding items 1,..., M shown in the upper part of FIG. 3 are variables that indicate a value of 1 when the corresponding individual has the corresponding feature and a value of 0 when the corresponding individual does not have the corresponding feature.
[0052] As a first example, the first life-person table L1 can describe the feature quantities X1,..., X M obtained by a questionnaire survey. In this case, the feature quantity X mcan be a variable that indicates a value of 1 when the corresponding individual gives an affirmative answer to the $m^{th}$ ($m = 1, 2, \ldots, M$) question included in the questionnaire, and indicates a value of 0 when giving a negative answer. That is, the feature quantity $X$ m can be a variable representing an individual's answer to the $m^{th}$ question. The feature quantity $X$ without an index described below is a simplified expression of the feature quantity $X$ m .
[0053] As a second example, the first lifestyle table L1 can describe the feature quantities $X1, \ldots, X$ M . In this case, the feature quantity $X$ can be a variable that indicates a value of 1 when the purchase quantity of the $m^{th}$ ($m = 1, 2, \ldots, M$) product by the corresponding individual exceeds the reference quantity during a past predetermined period, and indicates a value of 0 when the purchase quantity is less than or equal to the reference quantity. The purchase quantity can be, but is not limited to, the number of purchases or the purchase amount. The reference quantity can be determined as a value of zero or more. When the reference quantity is zero, the feature quantity $X$ represents the presence or absence of the purchase of the $m^{th}$ product during a past predetermined period.
[0054] As a third example, the first lifestyle table L1 can describe the feature quantities $X1, \ldots, X$ M . In this case, the feature quantity $X$ can be a variable that indicates a value of 1 when the viewing quantity of the $m^{th}$ ($m = 1, 2, \ldots, M$) program by the corresponding individual exceeds the reference quantity, and indicates a value of 0 when the viewing quantity is less than or equal to the reference quantity. The viewing quantity can be, but is not limited to, the viewing time. The reference quantity can be predetermined as a value of zero or more. The program can be, for example, a television broadcast program. When the reference quantity is zero, the feature quantity $X$ represents the presence or absence of viewing program $m$.
[0055] However, the first lifestyle table L1 is not limited to the first to third examples described above. The first lifestyle table L1 can be a table that describes the feature quantities $X1, \ldots, X$ M relating to the behavior of lifestyle individuals not limited to, for example, answer behavior, purchase behavior, and viewing behavior to questions. The behavior referred to here includes online behavior and offline behavior. The first lifestyle table L1 can include the feature quantities $X1, \ldots, X$ M relating to the awareness of lifestyle individuals.It may be included.
[0056] The feature quantity X does not have to be a variable represented by two values of "0" or "1". For example, the feature quantity X may be a variable indicating the purchase quantity of the m-th product by the corresponding individual in a past predetermined period. The feature quantity X may be a variable indicating the viewing amount of the m-th program by the corresponding individual.
[0057] The lower part of FIG. 3 shows the configuration of the statistical table Ls1 generated by the table processing (see FIG. 2) for the first life table L1 illustrated in the upper part of FIG. 3. The statistical table Ls1 has records (hereinafter referred to as "individual ticket statistical data") for each combination of area and basic attributes. A group of individual ticket statistical data is generated by statistical processing on the first life table L1.
[0058] Specifically, each individual ticket statistical data is generated by statistical processing on a group of records of individuals belonging to the corresponding area and basic attributes. The combination of area and basic attributes mentioned here is a combination of area, age group, and gender.
[0059] Each individual ticket statistical data is for the feature quantities X1,..., X of the set of individuals having the corresponding basic attributes (gender and age group) in the corresponding area among the first set corresponding to the first life table L1. M It is generated by statistical processing.
[0060] Hereinafter, the set of individuals classified by the combination of area and basic attributes is expressed as a "segment". The individual ticket statistical data for each combination of area and basic attributes is the individual ticket statistical data for each segment. The "group of individual ticket statistical data" for each area corresponds to the statistical data regarding the set of individuals belonging to the corresponding area, that is, the area statistical data.
[0061] The individual ticket statistical data describes the segment feature amounts Ys of each item m (m = 1, 2, …, M), together with the information explaining the area and basic attributes of the corresponding segment. The area of the corresponding segment is the area to which the corresponding segment belongs, specifically, the living area of the corresponding segment. The segment feature amount Ys of item m is the statistical value of the feature amount X of item m for the corresponding segment.
[0062] When an execution instruction for table processing is input through the user interface 17, the processor 11 executes the table processing shown in FIG. 2. The execution instruction designates the first inhabitant table L1 to be processed.
[0063] When starting the table processing, the processor 11 acquires the first inhabitant table L1 designated by the execution instruction (S110). The processor 11 can acquire the designated first inhabitant table L1 by reading it from the storage 15.
[0064] In the subsequent S120, the processor 11 selects one from among a plurality of areas defining the segment and selects one from among a plurality of basic attributes defining the segment, thereby selecting the processing target segment, that is, the combination of the processing target area and basic attributes (S120). Here, the area and basic attributes selected as the processing target are expressed as area z and basic attribute r.
[0065] In the subsequent S130, the processor 11 calculates the feature amount Y m of each item m (m = 1, 2, …, M) regarding the processing target segment by referring to a group of individual ticket data corresponding to the processing target segment.
[0066] The feature amount Y calculated in S130 mis the statistical value, specifically the average, of the feature quantity X of one or more individuals belonging to the processing target segment among a plurality of individuals belonging to the first set. The calculation of the average corresponds to an example of statistical processing. The average of the feature quantity X corresponds to an example of a representative value of the feature quantity X.
[0067] That is, the feature quantity Y m is the average of the feature quantity X of one or more individuals having the basic attribute r in the set of individuals within the area z. The individuals within the area z are individuals for whom the area z is the living area. The feature quantity Y without an index used below is a simplified expression of the feature quantity Y. m is a simplified expression of
[0068] As shown in FIG. 4, when the feature quantity X is represented by a binary value, the feature quantity Y as the average of the feature quantity X in a certain segment corresponds to the ratio of individuals for whom the feature quantity X is the value 1 in the same segment. That is, the feature quantity Y corresponds to the ratio of individuals for whom the feature regarding item m satisfies a specific condition in one segment.
[0069] When the feature quantity X is the feature quantity of the first example described above, the feature quantity Y corresponds to the ratio of individuals who answered "yes" in response to the m-th question asking "yes" or "no".
[0070] When the feature quantity X is the feature quantity of the second example described above, the feature quantity Y corresponds to the ratio of individuals who purchased more of the m-th product than the reference quantity in the past predetermined period. When the reference quantity is zero, the feature quantity Y corresponds to the ratio of individuals who purchased the m-th product in the past predetermined period. When the feature quantity X is represented by the purchase quantity, the feature quantity Y is the average of the purchase quantity.
[0071] When the feature quantity X is the feature quantity of the third example described above, the feature quantity Y corresponds to the ratio of individuals who watched more of the m-th program than the reference quantity. When the reference quantity is zero, the feature quantity Y corresponds to the ratio of individuals who watched the program. When the feature quantity X is represented by the viewing quantity, the feature quantity Y is the average of the viewing quantity.
[0072] In subsequent S140, the processor 11 determines the feature amount Y for each calculated item (i.e., the average of the feature amounts X) as the segment feature amount Ys of the corresponding item. That is, for m = 1, 2, …, M, the processor 11 determines the feature amount Y of item m as the segment feature amount Ys of item m.
[0073] In subsequent S150, the processor 11 determines whether all combinations of the area and the basic attribute have been selected as the processing target. If it is determined that not all combinations have been selected (No in S150), the processor 11 changes the processing target segment, i.e., the area z and the basic attribute r to be processed, in S120 and executes the processing after S130. In this way, for each combination of the area and the basic attribute, the processor 11 determines the segment feature amount Ys of each item m (m = 1, 2, …, M) to be described in the individual ticket statistical data.
[0074] If it is determined in S150 that all combinations have been selected (Yes in S150), the processor 11 generates and outputs a statistical table Ls1 having individual ticket statistical data for each segment, i.e., for each combination of the area and the basic attribute, in S160.
[0075] The individual ticket statistical data for each segment in the generated statistical table Ls1 describes the segment feature amount Ys of each item m (m = 1, 2, …, M) in association with the information explaining the corresponding area and basic attribute.
[0076] The output destination of the statistical table Ls1 is, for example, the storage 15. That is, the generated statistical table Ls1 is stored in the storage 15. After executing the processing of S160, the processor 11 ends the table processing shown in FIG. 2. In this way, the processor 11 converts the first resident table L1 into the statistical table Ls1.
[0077] Further, the processor 11 executes the association-related process shown in FIG. 5 according to an execution instruction from an operator input through the user interface 17, thereby associating the statistical table Ls1 with the second life table L2 specified by the operator. As a result, the processor 11 generates an association table Le2.
[0078] The association table Le2 is a table obtained by associating the statistical table Ls1 with the second life table L2. The association table Le2 corresponds to a table obtained by expanding the second life table L2 using the statistical table Ls1. The association here includes associating a part of the information included in the statistical table Ls1 with the second life table L2.
[0079] In the execution instruction for the association-related process, the second life table L2 and the statistical table Ls1 to be processed are specified. The specified second life table L2 can be a table having individual slip data for each individual regarding a plurality of individuals belonging to the second set, similar to the first life table L1.
[0080] Alternatively, the second life table L2 can be a table (see FIG. 6A) having records that describe the characteristics of the corresponding segment in association with the area and basic attribute information of the corresponding segment, not for each individual but for each segment.
[0081] Each of the segments referred to here is a set of individuals having the same area and basic attributes among a plurality of individuals belonging to the second set. The term "person" as used in this specification should be understood as a term including, in addition to individuals, a set of individuals classified into segments.
[0082] FIG. 6A shows an example of the second life table L2 and the statistical table Ls1 to be associated. The statistical table Ls1 shown in FIG. 6A has records (individual slip statistical data) that describe the characteristics of the corresponding segment regarding the questionnaire response behavior of the corresponding segment in association with the area and basic attribute information of the corresponding segment for each segment regarding the first set.
[0083] This statistical table Ls1 describes the segment feature amounts Ys corresponding to the feature amounts X for each question obtained by the questionnaire survey. The statistical table Ls1 describes, as a plurality of items, for M = M1 questions, the segment feature amounts Ys for each question, for each segment, in association with the information on the corresponding area and basic attributes.
[0084] The second-lifer table L2 shown in FIG. 6A includes records that explain the characteristics related to the purchasing behavior of the corresponding segments, in association with the information on the area and basic attributes of the corresponding segments, for each segment as a person belonging to the second set.
[0085] Each record of the second-lifer table L2 describes, as a plurality of items, for M = M2 products, the feature amounts (in other words, the segment feature amounts Ys) for each product, in association with the information on the corresponding area and basic attributes. The feature amounts of each record are the feature amounts of the corresponding segment. The feature amounts for each product may indicate the average purchase amount of the corresponding product.
[0086] When starting the join-related process shown in FIG. 5, the processor 11 acquires the second-lifer table L2 specified in the execution instruction (S210). In subsequent S220, the processor 11 acquires the statistical table Ls1 specified in the execution instruction (S220).
[0087] Specifically, the processor 11 can acquire the second-lifer table L2 and the statistical table Ls1 by reading the second-lifer table L2 and the statistical table Ls1 from the storage 15 (S210, S220).
[0088] In subsequent S230, the processor 11 performs data fusion to join the specified statistical table Ls1 to the specified second-lifer table L2. That is, the processor 11 joins the statistical table Ls1 to the second-lifer table L2 so as to associate the records of the statistical table Ls1 whose combinations of area and basic attributes match with the records of the specified second-lifer table L2.
[0089] As a result, the processor 11 generates a combined table Le2 by combining the statistical table Ls1 with the second resident table L2. As described above, the combined table Le2 corresponds to a table obtained by expanding each record of the second resident table L2 using the statistical table Ls1.
[0090] FIG. 6B shows an example of the combined table Le2 generated by combining the second resident table L2 and the statistical table Ls1 shown in FIG. 6A. The combined table Le2 shown in FIG. 6B is configured such that segment feature amounts Ys for each question indicated by the individual vote statistical data of the corresponding area and basic attributes are added as expansion data to each record of the second resident table L2.
[0091] That is, the combined table Le2 includes, for each combination of an area and basic attributes, a record that describes segment feature amounts Ys for each question and feature amounts for each product, in association with the information on the corresponding area and basic attributes.
[0092] When generating the combined table Le2 in S230 (see FIG. 5), the processor 11 then outputs the combined table Le2 in S240 and ends the combination-related process. The output destination of the combined table Le2 in S240 may be the storage 15. That is, the processor 11 can store the combined table Le2 generated in S230 in the storage 15.
[0093] In the storage 15 to which the combined table Le2 is output, population information for each segment, that is, for each combination of an area and basic attributes, is preliminarily stored as one table (hereinafter referred to as the "population table") Lp for analysis in consideration of the population of the combined table Le2.
[0094] The population table Lp shown in FIG. 7 includes, for each combination of an area, an age group, and a gender, a record that describes the population of the corresponding age group and gender in the corresponding area.
[0095] The above-described first set and second set are, for example, sample sets when considering the set of individuals living in Japan as the population. The population table Lp describes the population of the population including the first set and the second set. The population table Lp can be created, for example, based on population information provided by the Geospatial Information Authority of Japan in Japan.
[0096] When generating the join table Le2, the processor 11 may attach the population information of the corresponding segment to each record of the join table Le2. That is, each record of the join table Le2 may be attached with population information explaining the population of residents of the corresponding age group and gender in the corresponding area (see FIG. 8).
[0097] By analyzing this join table Le2, the operator of the information processing system 10 can convert the segment feature amount Ys for each segment into a feature amount considering the population distribution of the population, and analyze the features of each segment.
[0098] For example, as shown in FIG. 8, consider the case where the join table Le2 describes the affirmative response rate for each question obtained by a questionnaire survey as the segment feature amount Ys. When the first-lifer table L1 describes the feature amount X of the above-described first example regarding the question, the segment feature amount Ys corresponding to the feature amount X is the ratio of the number of people who answered affirmatively to the corresponding question in the corresponding segment in the first set, that is, the affirmative response rate in the corresponding segment.
[0099] By multiplying the segment feature amount Ys indicating this affirmative response rate by the population P of the corresponding segment, it is possible to estimate the number of affirmative respondents, which is the number of individuals with the corresponding basic attributes living in the corresponding area in the population who answered affirmatively to the question.
[0100] Figure 8 illustrates that when the positive response rate in the \(i\)th (\(i = 1, 2, \ldots\)) segment is \(Ys[i]\) and the population of the \(i\)th segment is \(P[i]\), the number of positive respondents (estimated value) in the area and basic attributes corresponding to the segment can be calculated as \(Ys[i] \cdot P[i]\). \(Ys[i]\) represents the segment feature quantity \(Ys\) of the \(i\)th segment.
[0101] In this way, by associating population information with the combined table Le2, the number of positive respondents based on the population standard can be estimated, which is convenient. However, the segment feature quantity \(Ys\) is not limited to the positive response rate.
[0102] For example, the segment feature quantity \(Ys\) corresponding to item \(m\) can represent the proportion of individuals whose features corresponding to item \(m\) satisfy specific conditions. In this case, by multiplying the segment feature quantity \(Ys\) and the population \(P\) of the corresponding segment, in the population, among multiple individuals living in the corresponding area and having the corresponding basic attributes, the estimated number of individuals whose features corresponding to item \(m\) satisfy specific conditions can be calculated.
[0103] For example, the segment feature quantity \(Ys\) corresponding to item \(m\) can represent the average purchase quantity of the product corresponding to item \(m\). In this case, by multiplying the segment feature quantity \(Ys\) and the population \(P\) of the corresponding segment, in the population, the estimated value of the total purchase quantity of the corresponding product by multiple individuals living in the corresponding area and having the corresponding basic attributes can be calculated.
[0104] In addition, Figure 8 shows that by using the number of positive respondents \(Ys[i] \cdot P[i]\) for each segment, the positive response rate \(Ys\) of the target group \(G\), which is a group of segments of interest among multiple segments, * can be calculated based on the population standard.
[0105] For example, when the target group \(G\) is a group of the first segment, the second segment, the third segment, and the fourth segment, the sum \(\sum\) of the number of positive respondents \(Ys[i] \cdot P[i]\) in the \(i\)th (\(i = 1, 2, 3, 4\)) segmenti∈G Ys[i]·P[i] = (Ys[1]·P[1] + Ys[2]·P[2] + Ys[3]·P[3] + Ys[4]·P[4]) is divided by the total population Σ of the target group G i∈G by P[i] = (P[1] + P[2] + P[3] + P[4]) to obtain the affirmative response rate Ys of the target group G * = Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] can be calculated based on the population standard.
[0106] The variable i here is the index of the segment as described above, and i ∈ G means the segment belonging to the target group G. Σ i∈G F[i] means calculating the sum of F[i] for each segment corresponding to the target group G. F[i] is, for example, Ys[i]·P[i] or P[i].
[0107] The information processing system 10 has such a function of calculating and outputting the number of affirmative respondents Σ i∈G Ys[i]·P[i] and the affirmative response rate Ys * = Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] according to the instruction from the operator. The affirmative response rate Ys * corresponds to the affirmative response rate of the target group G with guaranteed representativeness.
[0108] The processor 11 can realize this function by, for example, executing the response analysis process shown in FIG. 9 according to the instruction from the operator. When starting the response analysis process shown in FIG. 9, the processor 11 obtains from the operator information specifying the combined table Le2 to be processed through the user interface 17 (S310).
[0109] In the subsequent S320, the processor 11 obtains from the operator information specifying the target group G through the user interface 17. The operation of specifying the target group G corresponds to the operation of specifying one or more segments and corresponds to the operation of specifying one or more combinations regarding the area and basic attributes.
[0110] The target group G may be all combinations of areas and basic attributes. The target group G may be a group specified only by area regardless of the basic attributes. The target group G may be a group specified only by basic attributes regardless of the area.
[0111] In subsequent S330, the processor 11 calculates, for each item m (m = 1, 2,..., M), the number of affirmative respondents Σ i∈G Ys[i]·P[i] in the target group G with reference to the combined table Le2. In subsequent S340, the processor 11 calculates the total population Σ i∈G P[i] of the target group G.
[0112] In subsequent S350, for each item m (m = 1, 2,..., M), the processor 11 calculates the affirmative response rate Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] of the target group G with guaranteed representativeness.
[0113] The affirmative response rate Ys * of item m corresponds to an estimated value of the proportion of individuals whose feature quantity X of item m satisfies a specific condition (X = 1) among the set of individuals belonging to one or more specified segments in the population. The number of affirmative respondents of item m corresponds to an estimated value of the total number of individuals whose feature quantity X of item m satisfies a specific condition (X = 1) among the set of individuals belonging to one or more specified segments in the population.
[0114] In subsequent S360, for each item m (m = 1, 2,..., M), the processor 11 outputs the affirmative response rate Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the number of affirmative respondents Σ i∈G Ys[i]·P[i].
[0115] The processor 11 calculates the affirmative response rate Ys * =Σi∈G Ys[i]·P[i] / Σ i∈G P[i] and the number of affirmative respondents Σ i∈G Ys[i]·P[i] can be output to the operator through the display unit 17A.
[0116] The processor 11 calculates the affirmative response rate Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the number of affirmative respondents Σ i∈G The data file describing Ys[i]·P[i] may be stored in the storage 15 to output the corresponding information. Then, the processor 11 ends the response analysis process.
[0117] The above-described response analysis process is a process of analyzing the combined table Le2 that describes the affirmative response rate as the segment feature amount Ys. However, this process can be generalized to the analysis of the combined table Le2 that describes parameters other than the affirmative response rate as the segment feature amount Ys.
[0118] When the segment feature amount Ys[i] of the i-th segment indicates the ratio of individuals whose features regarding item m satisfy specific conditions in the same segment, the value Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the value Σ i∈G Ys[i]·P[i] respectively indicate the estimated ratio and number of individuals whose features regarding item m satisfy specific conditions in the target group G of the population.
[0119] When the segment feature amount Ys[i] of the i-th segment represents the average purchase quantity of the m-th product in the same segment, the value Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the value Σ i∈G Ys[i]·P[i] respectively indicate the estimated average and total of the purchase quantity of the m-th product in the target group G of the population. The purchase quantity of the m-th product may be read as the viewing quantity of the m-th program.
[0120] The target group G may be designated as a group with a purchase quantity equal to or greater than a specified quantity, or may be designated as a group with a viewing quantity equal to or greater than a specified quantity. In the combined table Le2, when a plurality of types of feature quantities such as a positive response rate, an average purchase quantity, and an average viewing quantity are described as the segment feature quantity Ys, a group that satisfies a plurality of conditions, such as a group in which the average purchase quantity satisfies a specific condition and the average viewing quantity satisfies a specific condition, may be designated as the target group G.
[0121] That is, a group in which one or more of the segment feature quantities Ys among the segment feature quantities Ys of a plurality of items satisfy a specific condition may be designated as the target group G. In this case, the number Σ of the target group G in the population i∈G P[i] is the value Ys * =Σ i∈G Ys[i]·P[i] / Σ i∈G P[i] and the value Σ i∈G Ys[i]·P[i] may be output together.
[0122] According to the information processing system 10 of the present embodiment described above, after statistically processing personal data, a plurality of data are combined using attributes and areas as keys. Therefore, it is possible to generate combined data in which a plurality of data are meaningfully combined, and combined data useful for analyzing consumers can be generated.
[0123] Data related to people often has information on attributes and areas attached. However, it is difficult to meaningfully combine the data of a plurality of individuals with only the information on attributes and areas. In contrast, according to the present embodiment, since personal data is statistically processed and combined with other data, a plurality of data can be meaningfully combined using a small amount of information such as attributes and areas as keys, and data with a large amount of information can be generated.
[0124] Furthermore, according to the information processing system 10 of the present embodiment, using the population table Tp, the feature amount of the sample set can be converted into the feature amount based on the population standard. Therefore, the information processing system 10 of the present embodiment is very useful for data analysis considering the population.
[0125] Incidentally, the above-described information processing system 10 may be modified as follows. That is, in S160, the processor 11 may generate the statistical table Ls12 shown in FIG. 10 instead of the above-described statistical table Ls1.
[0126] Similar to the above-described statistical table Ls1, the statistical table Ls12 of the modification example includes the individual ticket statistical data for each segment. The individual ticket statistical data of the i-th segment includes information for explaining the population P[i] of the i-th segment in addition to the same information as the statistical table Ls1.
[0127] The individual ticket statistical data of the i-th segment further includes, for each item, the estimated number of individuals Ys[i]·P[i] of individuals whose feature amount X of the corresponding item m satisfies the specific condition (X = 1), appended with the segment feature amount Ys[i] indicating the ratio of individuals whose feature amount X of the corresponding item m satisfies the specific condition (X = 1).
[0128] That is, as shown in FIG. 10, each individual ticket statistical data in the statistical table Ls12 of the modification example includes information on the population P[i] of the corresponding segment, the segment feature amount Ys[i] of each item in the corresponding segment, and the estimated number of individuals Ys[i]·P[i] where X = 1. The estimated number can be, for example, the estimated number of persons taking a specific action regarding response behavior, purchase behavior, viewing behavior, etc. The estimated number can be, for example, the estimated value of the number of affirmative respondents described above.
[0129] In S160, the processor 11 can calculate the estimated number of individuals Ys[i]·P[i] of each item m based on the population P[i] and the segment feature amount Ys[i] of each item m in the corresponding segment for each segment, and generate the statistical table Ls12 appended with these.
[0130] The same process may be executed in the process of S240 instead of S160. That is, in S240, the processor 11 can attach the population P[i] of the corresponding segment and the estimated population Ys[i]·P[i] of each item m to each record of the join table Le2. The information processing system 10 of this modified example that generates the statistical table Ls12 and the join table Le2 is very useful for data analysis considering the population.
[0131] [Second Embodiment] Next, the information processing system 10 of the second embodiment will be described. The information processing system 10 of the second embodiment is configured in the same manner as the first embodiment except that the table processing executed by the processor 11 is different. The hardware configuration of the information processing system 10 in the second embodiment is the same as that in the first embodiment. Therefore, hereinafter, the same components as those in the first embodiment included in the information processing system 10 of the second embodiment are denoted by the same reference numerals as those in the first embodiment, and the description thereof is omitted.
[0132] According to this embodiment, the processor 11 executes the table processing shown in FIG. 11 instead of the table processing shown in FIG. 2. When starting the table processing shown in FIG. 11, the processor 11 acquires the first inhabitant table L1 designated by an execution instruction in the same manner as the process of S110 (S410). In subsequent S420, the processor 11 selects a processing target segment in the same manner as the process of S120 (S420).
[0133] That is, the processor 11 selects a processing target area from among a plurality of areas and a processing target basic attribute from among a plurality of basic attributes, thereby selecting a combination of the processing target area and the basic attribute (S420).
[0134] In subsequent S430, for each of one or more peripheral areas that are one or more areas located around the area z to be processed, the processor 11 calculates, for each peripheral area, a feature quantity Y that is the average of the feature quantities X of one or more individuals belonging to the basic attribute r to be processed in the corresponding peripheral area.
[0135] Here, for each of one or more peripheral areas located around the area z to be processed, a symbol z k (k = 1, 2, …, K) is attached. The one or more peripheral areas can be one or more areas adjacent to the area z to be processed.
[0136] When K areas are adjacent to the area z to be processed, the area z to be processed is adjacent to the peripheral areas z1, …, z K . Hereinafter, the feature quantity Y of the item m regarding one or more individuals having the basic attribute r in the peripheral area z k is referred to as the feature quantity Y m (z k , r) or Y(z k , r).
[0137] In S430, for each peripheral area z k (k = 1, 2, …, K), the feature quantity Y(z k , r) for each item is calculated. The calculation method of the feature quantity Y m (z k , r) of the m-th item is as described in the first embodiment.
[0138] That is, the feature quantity Y k of the peripheral area z regarding the basic attribute r m (z k , r) is calculated as the average of the feature quantities X of the corresponding item m of one or more individuals having the basic attribute r, among the plurality of individuals belonging to the first set and being the set of individuals within the area z k . The individuals within the area z k are individuals for whom the area z k is the living area.
[0139] In subsequent S440, the processor 11 calculates the weight W(z,z k (k = 1,2,…,K)) corresponding to the reciprocal of the distance d(z,z k ) between the processing target area z and each peripheral area z k ) = 1 / d(z,z k ). The distance d(z,z k ) between the processing target area z and the peripheral area z k can be the distance between the representative point of the processing target area z and the representative point of the peripheral area z k , as shown in FIG. 12. Examples of representative points include the geometric center, geometric centroid, and population centroid, etc.
[0140] In subsequent S450, for each item m (m = 1,2,…,M), the processor 11 uses the inverse distance weighted (IDW) method to calculate the estimated feature amount Ya(z,r) of the processing target segment in the area z by the weighted average of the feature amounts Y(z k ,r) of one or more peripheral areas z k . The estimated feature amount Ya(z,r) is the weighted sum (i.e., weighted average) of the feature amounts Y(z k ,r) using the weight W(z,z k ). The weighted average corresponds to an example of statistical processing, and the weighted sum corresponds to an example of a statistical value. Specifically, the estimated feature amount Ya(z,r) is calculated by the following formula.
[0141] Ya(z,r)=Σ k {W(z,z k )·Y(z k ,r)} / Σ k W(z,z k ) In subsequent S460, the processor 11 determines whether a sample of the processing target segment exists in the processing target area z. This determination corresponds to determining whether the processing target area z is an empty area where there is no sample (individual) having the attribute r of the processing target. When a sample of the processing target segment does not exist in the processing target area z, the processing target area z is an empty area where there is no sample (individual) having the attribute r of the processing target.
[0142] The processor 11 can determine whether a sample of the processing target segment exists in the processing target area z by determining whether the individual ticket data of one or more individuals belonging to the processing target segment exists in the first life table L1.
[0143] Among the multiple individuals belonging to the first set, the processor 11 searches for individuals having the basic attribute r within the first life table L1, which are the individuals within the processing target area z, and can determine the presence or absence of a sample based on the presence or absence of the corresponding individuals.
[0144] When the processor 11 determines that there is no sample of the processing target segment (No in S460), in the statistical table Ls1, the segment feature amount Ys of each item m (m = 1, 2,..., M) to be described in the individual ticket statistical data of the processing target segment is determined to be the estimated feature amount Ya(z, r) of each item m (m = 1, 2,..., M) calculated in S450 (S470), and the process proceeds to the process of S500.
[0145] On the other hand, when the processor 11 determines that a sample of the processing target segment exists in the processing target area z (Yes in S460), for each item m (m = 1, 2,..., M), the observed feature amount Y m (z, r) of the processing target area z is calculated (S480). Y(z, r) hereinafter is a simplified expression of Y m (z, r).
[0146] The observed feature amount Y(z, r) of the processing target area z is the feature amount Y of the item m regarding one or more individuals having the basic attribute r of the processing target in the processing target area z. That is, the observed feature amount Y(z, r) is the average of the feature amounts X of one or more individuals belonging to the basic attribute r of the processing target in the processing target area z.
[0147] In subsequent S490, the processor 11 determines, within the statistical table Ls1, the segment feature amount Ys of each item m (m = 1, 2, …, M) to be described in the individual ticket statistical data of the segment to be processed as the average {Ya(z,r) + Y(z,r)} / 2 of the estimated feature amount Ya(z,r) of the corresponding item m (m = 1, 2, …, M) calculated in S450 and the observed feature amount Y(z,r) of the corresponding m (m = 1, 2, …, M) calculated in S480.
[0148] In subsequent S500, the processor 11 determines, in the same manner as the processing in S150, whether all combinations of the area and the basic attributes have been selected as the processing target. If it is determined that not all combinations have been selected (No in S500), the processor 11 changes the segment to be processed, that is, the area z and the basic attribute r of the processing target, in S420 and executes the processing after S430. In this way, the processor 11 determines the segment feature amount Ys of each item m (m = 1, 2, …, M) to be described in the individual ticket statistical data for each combination of the area and the basic attributes.
[0149] In S500, if it is determined that all combinations have been selected (Yes in S500), the processor 11 generates and outputs, in S510, a statistical table Ls1 having individual ticket statistical data for each segment, that is, for each combination of the area and the basic attributes, in the same manner as the processing in S160. An example of the statistical table Ls1 is as shown in FIG. 3. In S510, the statistical table Ls12 shown in FIG. 10 may be generated and output.
[0150] The configurations of the statistical tables Ls1 and Ls12 generated in S510 are the same as those in the first embodiment except that the segment feature amount Ys of each item m (m = 1, 2, …, M) described in each individual ticket statistical data is the value determined in S470 and S490.
[0151] After executing the processing in S510, the processor 11 ends the table processing shown in FIG. 11. In this way, the processor 11 converts the first resident table L1 into the statistical table Ls1.
[0152] According to this embodiment, the segment feature amount Ys for each segment is calculated and determined using the IDW method. Therefore, even when the number of samples in the first inhabitant table L1 is small, the feature amount Y(z k of the peripheral area z k ,r) can be used to accurately calculate the segment feature amount Ys of the processing target area z without omission, and it is possible to generate a highly reliable statistical table Ls1.
[0153] [Third Embodiment] Subsequently, the information processing system 10 of the third embodiment will be described. The information processing system 10 of the third embodiment is configured by adding new functions to the information processing system 10 of the first embodiment or the second embodiment. Therefore, hereinafter, the same components as those in the above embodiments provided in the information processing system 10 of the third embodiment are denoted by the same reference numerals as those in the above embodiments, and the description thereof is omitted.
[0154] In the information processing system 10 of this embodiment, the processor 11 executes the association-related processing shown in FIG. 5 as the first association-related processing. The processor 11 is further configured to execute the second association-related processing shown in FIG. 13 to combine the third inhabitant table L3 with the combined table Le2 generated by the first association-related processing to expand the third inhabitant table L3. Hereinafter, the table generated by expanding the third inhabitant table L3 with the combined table Le2 is expressed as an expanded table Le3.
[0155] As another example, the expanded table Le3 may be generated using the statistical table Ls1 instead of the combined table Le2. That is, the processor 11 may combine the third inhabitant table L3 with the statistical table Ls1 to generate the expanded table Le3. In this case, in the second association-related processing described below, the statistical table Ls1 is used instead of the combined table Le2. As a further alternative example, the statistical table Ls12 shown in FIG. 10 may be used instead of the statistical table Ls1 for generating the expanded table Le3.
[0156] The processor 11 executes the second association-related process shown in FIG. 13 in accordance with an execution instruction from an operator input through the user interface 17. When starting the second association-related process, the processor 11 acquires the third party table L3 specified in the execution instruction (S610). In subsequent S620, the processor 11 acquires the association table Le2 specified in the execution instruction. Thereafter, the processor 11 executes the process of S630.
[0157] The processor 11 can acquire the third party table L3 and the association table Le2 by reading the third party table L3 and the association table Le2 from the storage 15 (S610, S620).
[0158] As shown in FIG. 14, the third party table L3 is a table that describes the characteristics of a plurality of individuals belonging to a third set, and includes records for each individual (hereinafter referred to as "individual records").
[0159] Similar to the individual ticket data in the first party table L1, the individual record describes the area and basic attributes (i.e., gender and age group) of the corresponding individual in association with the identification code (i.e., ID) of the corresponding individual.
[0160] Each individual record further describes other characteristics of the corresponding individual in association with information on the area and basic attributes of the corresponding individual. The other characteristics include attributes of the individual other than the basic attributes (hereinafter referred to as "additional attributes") as shown in FIG. 14. The additional attribute shown in FIG. 14 is occupation. The additional attribute is one feature quantity. Although not shown, the individual record may further include information on the feature quantity X for each item related to the behavior and / or consciousness of the individual, similar to the individual ticket data in the first party table L1.
[0161] The association table Le2 to be associated with the third party table L3 configured in this way includes records for each segment (i.e., a combination of area and basic attributes), similar to the first embodiment.
[0162] Each record in the association table Le2, similar to the individual ticket statistical data of the first embodiment, is associated with the area corresponding to the segment and information on the basic attributes, and includes information for explaining the characteristics of the set of individuals corresponding to the segment by the segment feature amounts Ys for each item. However, the record for each segment in the association table Le2 may be a record indicating the segment feature amounts Ys for each item calculated using the IDW method, similar to the second embodiment.
[0163] In S630, the processor 11 performs data fusion to associate the designated association table Le2 with the designated third-party living person table L3. That is, the processor 11 associates the records of the association table Le2 whose combination of area and basic attributes matches with each record of the designated third-party living person table L3, thereby associating the association table Le2 with the third-party living person table L3.
[0164] As a result, the processor 11 generates an extended table Le3 in which the association table Le2 is associated with the third-party living person table L3. The extended table Le3 corresponds to a table in which each record of the third-party living person table L3 is data-expanded using the association table Le2.
[0165] In S640, the processor 11 outputs the extended table Le3 generated in S630 and ends the second association-related process. The output destination of the extended table Le3 in S640 may be the storage 15. That is, the processor 11 can store the extended table Le3 generated in S630 in the storage 15.
[0166] FIG. 15 shows an example of the extended table Le3 generated by associating the third-party living person table L3 and the association table Le2 shown in FIG. 14. The extended table Le3 shown in FIG. 15 is configured such that for each record of the third-party living person table L3, the segment feature amounts Ys for each item indicated by the corresponding records of the area and basic attributes included in the association table Le2 are described. Hereinafter, each record included in the extended table Le3 is also referred to as an extended record.
[0167] The extended record further describes a weight-back (WB) value V that takes into account the population P of the corresponding segment and the number H of individuals belonging to the same segment in the third-party table L3. The weight-back value V of a certain segment is the value V = P / H obtained by dividing the population P of the segment by the number H of individuals belonging to the same segment in the third-party table L3. The extended table Le3 describes the weight-back value V for each segment.
[0168] This weight-back value V is used to perform an operation considering the segment composition between the population and the sample set. The segment composition here means the composition ratio of the segment, that is, the ratio of individuals belonging to each segment. The population here is the set of individuals living in Japan.
[0169] The sample set corresponds to the third set and is the set of individuals having records in the third-party table L3. The sample set is a subset of the set of individuals living in Japan (the population). The number H described above is the sample size of the corresponding segment in the sample set (the third set). Hereinafter, the number H will be expressed as the sample size H.
[0170] The processor 11 can determine the population P of each segment by referring to the population table Lp shown in FIG. 7. The processor 11 can determine the sample size H by counting the number of records of individuals belonging to the corresponding segment in the third-party table L3 or the extended table Le3 for each segment.
[0171] The weight-back value V is a value for each segment. As can be understood from FIG. 15, the same weight-back value V is described in the extended records of individuals belonging to the same segment.
[0172] Using this weight-back value V, the affirmative response rate Yu of each question described in the extended record can be adjusted to a value Yu considering the segment composition between the sample set and the population. *can be corrected. Below, this value Yu * is referred to as the corrected response rate Yu * and expressed as such.
[0173] The positive response rate Yu is the segment feature amount Ys for the corresponding question. This segment feature amount Ys is described in the combined table Le2 combined with the third party table L3. In the combined table Le2, the positive response rate Yu of each question described as the segment feature amount Ys is as described in the first embodiment or the second embodiment.
[0174] The positive response rate Yu of each question described in each extended record can be converted into the corrected response rate Yu in a group of a plurality of segments to be noted (hereinafter referred to as "noted segment group") according to the formula Yu * =Σ(Yu·V) / ΣV. *
[0175] The positive response rate Yu of each question described in each extended record is the positive response rate of each question in one segment to which the corresponding person belongs in the sample set (that is, the third set). According to the above formula Yu * =Σ(Yu·V) / ΣV, such a positive response rate Yu of each question can be converted into the positive response rate of each question in the set of persons corresponding to the noted segment group in the population as the corrected response rate Yu *
[0176] The numerator in the formula Yu * =Σ(Yu·V) / ΣV is the sum of (Yu·V) for a group of samples belonging to the noted segment group in the sample set. Therefore, the numerator Σ(Yu·V) represents the number of positive respondents within the noted segment group in the population. The denominator ΣV is the sum of the weight-back values V for a group of samples belonging to the noted segment group. Therefore, the denominator ΣV represents the population of the noted segment group in the population. From this explanation, the formula Yu * The corrected response rate Yu for each question calculated by =Σ(Yu·V) / ΣV * can be understood to be the affirmative response rate for each question in the set of consumers corresponding to the target segment group in the population.
[0177] The target segment group may be a group of all segments. In this case, the denominator ΣV corresponds to the total population of the population, and the corrected response rate Yu * represents the corrected response rate Yu for each question in the population. * represents.
[0178] The corrected response rate Yu * corresponds to the value obtained by correcting the affirmative response rate Yu in the sample set in consideration of the difference in segment composition between the sample set and the population to the affirmative response rate of the population.
[0179] The operator of the information processing system 10 can calculate the corrected response rate Yu * based on the affirmative response rate Yu and the weight-back value V described in each extended record in the extended table Le3.
[0180] According to the information processing system 10 configured as described above, personal data is extended using statistical data to generate an extended table Le3 with abundant information. Using this extended table Le3, the operator of the information processing system 10 can analyze the characteristics of consumers in detail, and at this time, it is possible to perform a meaningful feature analysis based on the population standard using the weight-back value.
[0181] According to the present embodiment, the processor 11 may execute processing similar to the processing shown in FIG. 9 regarding the calculation of the corrected response rate Yu * . That is, the processor 11 may acquire information for designating a target segment group from the operator through the user interface 17 (S320). The processor 11 can calculate and output the corrected response rate Yu * for each question in the designated target segment group according to the above formula (S350, S360).
[0182] [Other Embodiments] The present disclosure is not limited to the above-described embodiments and can take various forms. For example, in the above-described embodiments, an example of handling a lifestyle table (first lifestyle table L1) obtained through a questionnaire survey and a lifestyle table (second lifestyle table L2) obtained through a purchase survey has been described. However, the technology of the present disclosure can be applied to various types of data related to human behavior and awareness.
[0183] In the above-described embodiments, as the feature quantity Y, the ratio of individuals in the corresponding segment who satisfy the specific condition for the feature quantity X and the average of the feature quantity X in the corresponding segment were adopted. However, the feature quantity Y may be the number of samples (individuals) in the corresponding segment who satisfy the specific condition for the feature quantity X.
[0184] For example, the feature quantity Y may be the number of individuals who answered affirmatively to the m-th question, the number of individuals who purchased the m-th product, the number of individuals who purchased more than the reference amount of the m-th product, the number of individuals who watched the m-th program, or the number of individuals who watched more than the reference amount of the m-th program in the corresponding segment. In this case, by attaching information on the number of samples in the corresponding segment to the feature quantity Y or the corresponding segment feature quantity Ys, it is possible to change the number of people indicated by the feature quantity Y to a ratio, and statistical tables Ls1 and Ls12 can be constructed.
[0185] The functions of one component in the above-described embodiments may be provided distributively to a plurality of components. The functions of a plurality of components may be integrated into one component. A part of the configuration of the above-described embodiments may be omitted. At least a part of the configuration of the above-described embodiments may be added or replaced with the configuration of other above-described embodiments. All aspects included in the technical idea specified from the language described in the claims are embodiments of the present disclosure.
[0186] [Technical Ideas Disclosed in this Specification] It can be understood that the following technical ideas are disclosed in this specification. [Item 1] First explanatory data for explaining characteristics of a plurality of individuals belonging to a first set, the first acquisition unit being configured to acquire first explanatory data that describes, for each individual, a first feature amount of the corresponding individual in association with an attribute of the corresponding individual and information on an area to which the corresponding individual belongs; A generation unit configured to generate statistical data by statistically processing the first explanatory data, the statistical data including, for a plurality of areas, area statistical data for each area regarding a set of individuals belonging to the corresponding area, the area statistical data including, for a plurality of attributes, first feature data for each attribute, the first feature data describing a statistical value of the first feature amount regarding the one or more individuals calculated by statistical processing of the first feature amounts of the one or more individuals having the corresponding attribute among the set of individuals in the corresponding area; Second explanatory data for explaining characteristics of a plurality of people belonging to a second set, the second acquisition unit being configured to acquire second explanatory data including second feature data for each person, the second feature data describing a second feature amount of the corresponding person in association with an attribute of the corresponding person and information on an area to which the corresponding person belongs; A combining unit configured to combine the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data where the combination of the attribute and the area matches; Comprising; When combining the first feature data and the second feature data, the combining unit is configured to associate a weight-back value for each combination of area and attribute for the plurality of areas and the plurality of attributes with the combined data; The weight-back value is an information processing system based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data. [Item 2] The information processing system according to Item 1, wherein The weight back value corresponds to information in a system that divides the population of the attribute corresponding to the combination in the area corresponding to the combination by the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data. [Item 3] The information processing system according to Item 1 or Item 2, wherein the statistical value corresponds to a representative value of the first feature amount for the one or more individuals, or a ratio of individuals among the one or more individuals whose first feature amount satisfies a specific condition. [Item 4] The information processing system according to Item 1 or Item 2, wherein the first feature data describes, as the first feature amount, a feature amount related to at least one of the behavior and consciousness of the corresponding individual. [Item 5] The information processing system according to Item 1 or Item 2, wherein the first feature data describes, as the first feature amount, a feature amount related to the behavior of the corresponding individual, and the behavior includes at least one of viewing behavior, purchasing behavior, answering behavior for a questionnaire, and online behavior. [Item 6] The information processing system according to Item 4 or Item 5, wherein the behavior includes answering behavior for a questionnaire, the first feature amount represents an answer of the corresponding individual to the questionnaire, and the statistical value corresponds to a ratio of individuals among the one or more individuals who gave a specific answer to the questionnaire. [Item 7] The information processing system according to Item 4 or Item 5, wherein the behavior includes viewing behavior, the first feature amount represents whether or not the corresponding individual viewed the target or the viewing amount, The statistical value corresponds to an information processing system for the number or ratio of individuals among the one or more individuals who viewed the target, the number or ratio of individuals among the one or more individuals who viewed the target equal to or more than a reference, or a representative value of the viewing amount for the one or more individuals. [Item 8] The information processing system according to item 4 or item 5, wherein the action includes a purchasing action, the first feature amount represents the presence or absence or the purchase amount of the target by the corresponding individual, The statistical value corresponds to an information processing system for the number or ratio of individuals among the one or more individuals who purchased the target, the number or ratio of individuals among the one or more individuals who purchased the target equal to or more than a reference, or a representative value of the purchase amount for the one or more individuals. [Item 9] The information processing system according to any one of items 1 to 5, the first feature data describes, as the first feature amount, a feature amount related to the action of the corresponding individual, The generation unit calculates, for each combination of the area and the attribute, an estimated number of persons who take an action in which the first feature amount satisfies a specific condition in the corresponding area, based on the ratio of individuals among the one or more individuals having the corresponding attribute in the corresponding area whose first feature amount satisfies the specific condition and the population of the corresponding attribute in the corresponding area, and generates the first feature data describing the estimated number. [Item 10] The information processing system according to item 9, wherein the action is a response action to a question, the estimated number is an estimated number of respondents who give a specific response to the question. [Item 11] The information processing system according to any one of items 1 to 5, the statistical value corresponds to the ratio of individuals among the one or more individuals whose first feature amount satisfies a specific condition, the information processing system further Based on the statistical value and the population for the one or more combinations, on the condition that one or more combinations of the area and the attribute are specified, output at least one of the proportion of individuals in the group corresponding to the one or more combinations of the population where the first feature amount satisfies the specific condition and the total number of individuals where the first feature amount satisfies the specific condition. An output unit configured as such An information processing system comprising [Item 12] The information processing system according to Item 11, wherein The one or more combinations are specified by specifying a condition related to at least one of the area, the attribute, the first feature amount, and the second feature amount. An information processing system [Item 13] An information processing method executed by a computer, comprising: Obtaining first explanatory data that describes the characteristics of a plurality of individuals belonging to a first set, and for each individual, the first explanatory data that describes the first feature amount of the corresponding individual in association with the information of the attribute of the corresponding individual and the area to which the corresponding individual belongs Generating statistical data by statistically processing the first explanatory data, the statistical data including area statistical data for each area regarding a plurality of areas, the area statistical data including first feature data for each of a plurality of attributes, and the first feature data being the statistical value of the first feature amount regarding the one or more individuals calculated by statistically processing the first feature amounts of the one or more individuals having the corresponding attribute among the set of individuals in the corresponding area. Generating Obtaining second explanatory data that describes the characteristics of a plurality of people belonging to a second set, the second explanatory data including second feature data for each person, and the second explanatory data that describes the second feature amount of the corresponding person in association with the information of the attribute of the corresponding person and the area to which the corresponding person belongs Combining the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data in which the combinations of the attribute and the area match; including; The combining includes associating a weight-back value for each combination of area and attribute for the plurality of areas and the plurality of attributes with the combined data when combining the first feature data and the second feature data. The weight-back value is an information processing method based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data. [Item 14] A computer program for causing a computer to execute the information processing method according to Item 13.
Explanation of Signs
[0187] 10… Information processing system, 11… Processor, 13… Memory, 15… Storage, 17… User interface, 19… Communication interface, L1, L2, L3… Lifestyle table, Le2… Combined table, Le3… Extended table, Lp… Population table, Ls1, Ls12… Statistical table.
Claims
1. First explanatory data for explaining the characteristics of a plurality of individuals belonging to a first set, the first acquisition unit being configured to acquire, for each individual, the first explanatory data that describes the first feature amount of the corresponding individual in association with the attribute of the corresponding individual and the information of the area to which the corresponding individual belongs; A generation unit configured to generate statistical data by statistically processing the first explanatory data, the statistical data including, for a plurality of areas, area statistical data for each area regarding the set of individuals belonging to the corresponding area, the area statistical data including, for a plurality of attributes, first feature data for each attribute, the first feature data describing the statistical value of the first feature amount regarding the one or more individuals calculated by statistical processing on the one or more individuals having the corresponding attribute among the set of individuals in the corresponding area; Second explanatory data for explaining the characteristics of a plurality of people belonging to a second set, the second explanatory data including second feature data for each person, the second acquisition unit being configured to acquire the second explanatory data that describes the second feature amount of the corresponding person in association with the attribute of the corresponding person and the information of the area to which the corresponding person belongs; A combining unit configured to combine the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data where the combination of the attribute and the area matches; Comprising: When combining the first feature data and the second feature data, the combining unit is configured to associate a weight-back value for each combination of area and attribute for the plurality of areas and the plurality of attributes with the combined data; The weight-back value is an information processing system based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data.
2. The information processing system according to claim 1, wherein the weight-back value corresponds to a value obtained by dividing the population of the attribute corresponding to the combination in the area corresponding to the combination by the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data.
3. The information processing system according to claim 1 or claim 2, wherein the statistical value corresponds to a representative value of the first feature amount regarding the one or more individuals, or a ratio of individuals among the one or more individuals whose first feature amount satisfies a specific condition.
4. The information processing system according to claim 1 or claim 2, wherein the first feature data describes, as the first feature amount, a feature amount regarding at least one of the behavior and awareness of the corresponding individual.
5. The information processing system according to claim 1 or claim 2, wherein the first feature data describes, as the first feature amount, a feature amount regarding the behavior of the corresponding individual, and the behavior includes at least one of a viewing behavior, a purchasing behavior, an answering behavior to a questionnaire, and an online behavior.
6. The information processing system according to claim 4, wherein the behavior includes an answering behavior to a questionnaire, the first feature amount represents an answer of the corresponding individual to the questionnaire, and the statistical value corresponds to a ratio of individuals among the one or more individuals who gave a specific answer to the questionnaire.
7. The information processing system according to claim 4, wherein the behavior includes a viewing behavior, the first feature amount represents whether or not the corresponding individual viewed the target or the viewing amount, and the statistical value corresponds to the number or ratio of individuals among the one or more individuals who viewed the target, the number or ratio of individuals among the one or more individuals who viewed the target more than a reference amount, or a representative value of the viewing amount regarding the one or more individuals.
8. The information processing system according to claim 4, wherein the behavior includes a purchasing behavior, the first feature amount represents whether or not the corresponding individual purchased the target or the purchase amount, and the statistical value corresponds to the number or ratio of individuals among the one or more individuals who purchased the target, the number or ratio of individuals among the one or more individuals who purchased the target more than a reference amount, or a representative value of the purchase amount regarding the one or more individuals.
9. The information processing system according to claim 1 or claim 2, wherein the first feature data describes, as the first feature amount, a feature amount regarding the behavior of the corresponding individual, The generation unit calculates, for each combination of the area and the attribute, the estimated number of persons who take an action in the corresponding area in which the first feature amount satisfies the specific condition, based on the ratio of the individuals among the one or more individuals having the corresponding attribute in the corresponding area whose first feature amount satisfies the specific condition and the population of the corresponding attribute in the corresponding area, and generates first feature data describing the estimated number. An information processing system.
10. The information processing system according to claim 9, wherein the action is a response action to a questionnaire, and the estimated number is the estimated number of respondents who give a specific response to the questionnaire. An information processing system.
11. The information processing system according to claim 3, wherein the statistical value corresponds to the ratio of the individuals among the one or more individuals whose first feature amount satisfies the specific condition, and the information processing system further is configured to output, for at least one of the ratio of the individuals whose first feature amount satisfies the specific condition and the total number of the individuals whose first feature amount satisfies the specific condition in the group corresponding to the one or more combinations in the population, based on the statistical value and the population for the one or more combinations, on the condition that one or more combinations are specified for the combination of the area and the attribute. An output unit is provided. An information processing system.
12. The information processing system according to claim 11, wherein the one or more combinations are specified by specifying a condition related to at least one of the area, the attribute, the first feature amount, and the second feature amount. An information processing system.
13. An information processing method executed by a computer, acquiring first explanatory data that describes the characteristics of a plurality of individuals belonging to a first set, the first explanatory data describing, for each individual, a first feature amount of the corresponding individual in association with information on the attribute of the corresponding individual and the area to which the corresponding individual belongs, Generating statistical data by statistically processing the first explanatory data, wherein the statistical data includes, for a plurality of areas, area statistical data for each area regarding a set of individuals belonging to the corresponding area, the area statistical data includes, for a plurality of attributes, first feature data for each attribute, and the first feature data describes a statistical value of the first feature amount regarding the one or more individuals calculated by statistically processing the first feature amount of the one or more individuals having the corresponding attribute among the set of individuals within the corresponding area. Obtaining second explanatory data for explaining the characteristics of a plurality of people belonging to a second set, the second explanatory data including second feature data for each person, and the second feature data describes the second feature amount of the corresponding person in association with the attribute of the corresponding person and information on the area to which the corresponding person belongs. Combining the statistical data based on the first explanatory data and the second explanatory data so as to associate the first feature data and the second feature data where the combination of the attribute and the area matches. Including The combining includes, when combining the first feature data and the second feature data, associating a weight-back value for each combination of area and attribute for the plurality of areas and the plurality of attributes with the combined data. The weight-back value is an information processing method based on the population of the attribute corresponding to the combination in the area corresponding to the combination and the number of samples of the attribute corresponding to the combination in the area corresponding to the combination of the second explanatory data.
14. A computer program for causing a computer to execute the information processing method according to Claim 13.
Citation Information
Patent Citations
Analysis method for marketing information, information processor and medium
JP2002140490A
General-purpose data fusion system and general-purpose data fusion method
JP5638675B1
Information-processing system
WO2016021726A1
Character processor
JP1981038675A