Information processing method, information processing device, and program
By processing user behavior data with flag information-based mask processing, the method addresses the challenge of integrating diverse data attributes, enhancing user profile estimation accuracy and reducing outlier impact.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Existing user profile estimation methods face challenges in accurately integrating and normalizing data with different attributes, such as cyber and physical information, leading to biased data distribution and reduced estimation accuracy due to outliers and varying data ranges.
An information processing method that involves acquiring user behavior data, generating frequency distribution information, performing mask processing to assign flag information, and estimating user profiles based on selectively processed data to improve estimation accuracy.
The method enhances user profile estimation accuracy by selectively handling data with flag information, reducing the impact of outliers and aligning data attributes, thereby improving the precision of user preference and characteristic estimation.
Smart Images

Figure 2026069330000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing method, an information processing apparatus, and a program.
Background Art
[0002] In recent years, with the spread of IoT devices, technologies for profiling users themselves using data generated from user device operations and the like have become widespread. In this way, when analyzing data for user profiling, normalization processing is used for the purpose of aligning the scales of different data and making it easier to compare feature amounts. For example, in Patent Document 1, for the purpose of making non-experienced users and experienced users match each other more effectively to improve their motivation, normalization is performed so that all values fall within the range from 0 to 1 by dividing each feature amount of non-experienced users and fully experienced users by the maximum value among them. A technique is disclosed.
[0003] In Patent Document 2, in the statistical processing of biological data such as test results in the medical field, among various test items that act complexly, for the purpose of outputting a comprehensive evaluation such as the degree of aging, by comparing the measured value of biological data with a reference value automatically selected from the distribution pattern, a technique for calculating a score serving as an evaluation criterion is disclosed.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] According to the technology disclosed in Patent Document 1, since each feature is normalized by dividing it by the maximum value, if the maximum value is an outlier, each feature that is not an outlier will become a value near 0 after the division, which may cause the data distribution to be biased and the user's characteristics to not be properly reflected.
[0006] According to the technology disclosed in Patent Document 2, a score is calculated based on a reference value automatically selected from a predetermined distribution pattern for measured values of biological data. However, the trend of data distribution may differ depending on the data attribute, such as physical information or cyber information. Therefore, when fitting a distribution based on biological data, it may be difficult to maintain the user's characteristics for other data with different data attributes.
[0007] Therefore, the user profile estimation using the technology disclosed in Patent Document 1 or Patent Document 2 has low estimation accuracy.
[0008] This disclosure aims to provide an information processing method, an information processing device, and a program that can improve the estimation accuracy of user profile estimation. [Means for solving the problem]
[0009] An information processing method according to one aspect of the present disclosure involves an information processing device acquiring first data relating to user behavior, generating frequency distribution information relating to the first data, performing mask processing based on the frequency distribution information to generate second data to which flag information is selectively assigned, estimating a user profile including user preferences, tendencies, or characteristics based on the second data, and outputting the estimated results of the user profile. [Effects of the Invention]
[0010] This disclosure makes it possible to improve the estimation accuracy of user profile estimation. [Brief explanation of the drawing]
[0011] [Figure 1] It is a diagram showing a configuration example of an information processing system according to an embodiment of the present disclosure in a simplified manner. [Figure 2] It is a diagram showing the stored content of the storage unit. [Figure 3] It is a diagram showing the functional configuration of the processing unit in a simplified manner. [Figure 4] It is a diagram showing the functional configuration of the processing unit in a simplified manner. [Figure 5] It is a diagram showing the details of the function of the normalization unit. [Figure 6] It is a diagram showing the functional configuration of the processing unit in a simplified manner. [Figure 7] It is a diagram showing the functional configuration of the processing unit in a simplified manner. [Figure 8] It is a flowchart showing the processing content executed by the processing unit. [Figure 9] It is a diagram showing the first example of a histogram. [Figure 10] It is a diagram showing the second example of a histogram. [Figure 11] It is a diagram showing the third example of a histogram. [Figure 12] It is a diagram showing the fourth example of a histogram. [Figure 13] It is a flowchart showing the first example of the details of the deletion process of specific data. [Figure 14] It is a flowchart showing the second example of the details of the deletion process of specific data. [Figure 15] It is a diagram showing the fifth example of a histogram. [Figure 16] It is a diagram showing the sixth example of a histogram. [Figure 17] It is a flowchart showing the details of the normalization process. [Figure 18] It is a flowchart showing the first example of the details of the calculation process of the normalization parameter. [Figure 19] It is a flowchart showing the second example of the details of the calculation process of the normalization parameter. [Figure 20] It is a diagram for explaining the value range conversion process by the conversion unit. [Modes for carrying out the invention]
[0012] (Knowledge that forms the basis of this disclosure) Technologies that recommend products and content such as videos by estimating user profiles, including user preferences, based on website screen operations and screen transitions have been put into practical use. Thus, it is possible to estimate user profiles from information such as operation logs, and further accuracy is expected by using not only cyber information such as websites, but also physical information such as the number of steps taken and the range of activity of the user.
[0013] When estimating user profiles using data containing outliers, the accuracy of the estimation decreases. Data attributes differ depending on the data source, such as cyber information obtained through web activities and physical information obtained through real-world activities like step counts. Furthermore, if the value range or number of digits differs among multiple data sets, those data sets will have different data attributes. When integrating and using multiple data sets with different data attributes, even if there are no outliers when looking at the data individually, one data set may become an outlier when compared with other data sets. For example, cyber information such as the number of times a website screen was clicked is all of the same order of magnitude, around a few times. On the other hand, the value range of physical information such as step counts (four digits for thousands of steps) is significantly different from the value range of cyber information (single digit for a few times). Therefore, it is difficult to analyze user profiles using the relationships between data when data is simply combined from multiple data sets with different data attributes.
[0014] Outliers can be caused by differences in data sources, such as cyber information and physical information. The number of digits in the data can also be a contributing factor. In this disclosure, all indicators of data type, including data source and number of digits, are referred to as data attributes. Data preprocessing is necessary to remove the effects of outliers, integrate and handle multiple data sets with different data attributes, and appropriately estimate user profiles. Normalization, one type of preprocessing, and data processing based on estimated distributions to which the data fits are methods for integrating and processing multiple data sets, but prior art methods have the following limitations.
[0015] This paper describes the challenges of user profile estimation using data that has different properties depending on the source. Considering that websites are designed with a specific purpose in mind, there are relationships between information related to screen transitions and click operations. Therefore, conventional profiling techniques integrate multiple cyber information sources. However, such a self-evident relationship does not exist between cyber and physical information. Conventional integration of cyber and physical information merely enumerates data, making accurate profile estimation difficult. This paper discusses the nature of the data and its impact on profiling in relation to this challenge.
[0016] This section explains the differences in the nature of cyber information and physical information. Under conditions where user permission has been obtained, cyber information can be obtained without any gaps in data acquisition, such as time periods or locations where data cannot be obtained from the user, and for view counts and click operations, unambiguous counts and truth values can be obtained, respectively. On the other hand, with physical information such as step counts, the number of steps actually walked by the user does not match the step count of a smartphone, resulting in inaccuracies in observed values and missing step count data due to the user not having a smartphone with them. Therefore, unambiguous cyber information and ambiguous physical information have different data properties.
[0017] This section explains the impact of the differences between cyber and physical information on profile estimation. When the aforementioned complete cyber information is considered as a data distribution, as the amount of data increases, it converges to a simple data distribution that can be expressed with a few parameters, such as a normal distribution. However, physical information, which is ambiguous and contains missing data, does not converge in the same way as the normal distribution of cyber information. In other words, simply applying the above simple distribution to physical information does not accurately capture the data distribution, and as a result, the estimation accuracy of user profile estimation decreases. Furthermore, even with the aforementioned complete cyber information, log data is generated due to operations that differ from normal use of the application (hereinafter abbreviated as "app") during application installation or tutorial viewing. The distribution of data when this log of operations that differ from normal use is mixed with logs of normal use is not a normal distribution. As a result, the data from logs of operations that differ from normal use become outliers, which can reduce the estimation accuracy of user profile estimation.
[0018] The following example illustrates the challenges of normalization due to the nature of the number of digits: The mean of data containing outliers, where the value "100" occurs in 3σ (standard deviation σ=1) of the standard normal distribution, corresponding to 1% of the number of observations, is "1.2," which is a significant deviation from the standard normal distribution mean of "0." Furthermore, the standard deviation of data containing outliers is "9.9," which is a significant deviation from the standard normal distribution standard deviation of "1."
[0019] Furthermore, the mean of the data containing outliers, where the value "10" exists at 3σ (standard deviation σ=1) of the standard normal distribution, corresponding to 1% of the number of observations, is "0.3," which is slightly different from the mean "0" of the standard normal distribution. On the other hand, the standard deviation of the data containing outliers is "1.0," which is no different from the standard deviation "1" of the standard normal distribution.
[0020] Therefore, if the majority of the data can be considered to follow a standard normal distribution, the influence of the outlier on the standard deviation of the observation distribution cannot be ignored when a value of "10," which is two orders of magnitude larger, is observed around three times the standard deviation, which is generally treated as an outlier. Furthermore, if a value of "100," which is three orders of magnitude larger, is observed, the impact on both the mean and the standard deviation is significant. Thus, in data where outliers exist that are 100 times or more the representative value of the observed values (e.g., the mean), conventional methods cannot accurately normalize the data.
[0021] Furthermore, if the observed data follows a non-standard normal distribution, the values described above will simply be approximate constant multiples and constant additions, essentially remaining the same. Therefore, even if only outliers less than 100 times the representative value of the observed values are observed in the above explanation, conventional methods cannot accurately normalize the data if the data distribution containing the outliers has roughly the same shape.
[0022] The following describes the specific challenges in user profile estimation using data that contains a large number of data points with different properties.
[0023] One challenge is that data collected early in a user's app usage (for example, within the first week of use) includes app usage logs from installation and tutorial viewing. These logs may contain a large proportion of data collected for reasons unrelated to the app's intended use. If this data is used as is, the distribution of the data can be distorted by an excessive inclusion of certain parameters (e.g., "0" or "1"), obscuring characteristic data about the app user. As a result, the accuracy of user profile estimation using the obtained data decreases.
[0024] The second challenge is that app usage logs related to physical information and app usage logs that consist solely of cyber information are mixed together. In this case, the distribution of each data and the range of parameter values may differ, and a uniform normalization that divides all data obtained from the app by the maximum value of the data may leave data that should be excluded as outliers.
[0025] The third challenge is that the distribution trends of biometric data and data obtained from apps differ. Biometric data has fewer outliers than expected, and the proportion of outliers in the total is small (e.g., a few percent), so they can be ignored. On the other hand, in data obtained from apps, outliers indicate specific user behavior within the app and are important information that represents user characteristics. Furthermore, outliers near "0" indicate the existence of users who are not using specific functions of the app, and are data that cannot be ignored. Therefore, normalization methods for biometric data that ignore the presence of outliers make it difficult to normalize while preserving user characteristics.
[0026] This disclosure was made to address these challenges and aims to provide a masking technique for improving the estimation accuracy of user profile estimation.
[0027] Next, we will describe each aspect of this disclosure.
[0028] An information processing method according to a first aspect of this disclosure involves an information processing device acquiring first data relating to user behavior, generating frequency distribution information relating to the first data, performing mask processing based on the frequency distribution information to generate second data to which flag information is selectively assigned, estimating a user profile including user preferences, tendencies, or characteristics based on the second data, and outputting the estimated result of the user profile.
[0029] According to the first embodiment, the estimation accuracy of user profile estimation can be improved by estimating a user profile based on second data to which flag information has been selectively assigned by masking the first data.
[0030] In the information processing method according to a second aspect of this disclosure, in the first aspect, it is preferable to determine whether or not to use the second data for estimating the user profile based on the flag information.
[0031] According to the second embodiment, the second data used to estimate the user profile can be appropriately determined based on flag information.
[0032] In the third aspect of the information processing method of this disclosure, in the second aspect, in generating the second data, the flag information is added to the first data which includes a singular portion in the frequency distribution information, and in estimating the user profile, the second data to which the flag information has been added is not used for estimating the user profile.
[0033] According to the third embodiment, the estimation accuracy of user profile estimation can be improved by not using second data containing singular portions in the frequency distribution information for user profile estimation.
[0034] In the fourth aspect of the information processing method of this disclosure, in the third aspect, when generating the second data, it is preferable to determine a predetermined unique frequency distribution shape among the frequency distribution information as the unique portion.
[0035] According to the fourth aspect, by determining a predetermined unique frequency distribution shape among the frequency distribution information as a unique portion, the unique portion can be appropriately identified.
[0036] The information processing method according to the fifth aspect of this disclosure further generates third data by normalizing the second data in any one of the first to fourth aspects, estimates the user profile based on the third data in the estimation of the user profile, and determines whether or not to normalize the second data based on the flag information in the generation of the third data.
[0037] According to the fifth aspect, the second data to be normalized can be appropriately determined based on flag information.
[0038] In the sixth aspect of the present disclosure, the information processing method, in the fifth aspect, involves assigning the flag information to the first data, which includes a feature portion representing the user's characteristics, when generating the second data, and not normalizing the second data to which the flag information has been assigned when generating the third data.
[0039] According to the sixth aspect, by not normalizing the second data to which flag information has been assigned to the first data containing the user's characteristic features, the user's characteristic features can be emphasized in the normalized third data, and as a result, the estimation accuracy of user profile estimation can be improved.
[0040] In the information processing method according to the seventh aspect of this disclosure, in the fifth or sixth aspect, when generating the third data, it is preferable to estimate the data distribution of the second data from a plurality of data distributions, calculate parameters for normalization based on the data distribution of the second data, and normalize the second data using the parameters.
[0041] According to the seventh embodiment, by calculating parameters for normalization based on the data distribution of the second data, the second data can be appropriately normalized using the calculated parameters.
[0042] The information processing method according to the eighth aspect of this disclosure further generates fourth data by transforming the range of the second data in any one of the first to seventh aspects, estimates the user profile based on the fourth data in the estimation of the user profile, and determines whether or not to transform the range of the second data in the generation of the fourth data based on the flag information.
[0043] According to the eighth aspect, the second data for transforming the range can be appropriately determined based on flag information.
[0044] In the information processing method according to the ninth aspect of this disclosure, in the eighth aspect, when generating the second data, the flag information is added to the first data whose data range is different from the reference range, and when generating the fourth data, the range of the second data to which the flag information has been added is converted to the reference range.
[0045] According to the ninth aspect, the range of the second data, to which flag information has been assigned for the first data whose range differs from the reference range, can be converted to the reference range, thereby aligning the range of the second data. As a result, the estimation accuracy of user profile estimation can be improved.
[0046] The information processing method according to the tenth aspect of this disclosure may further perform matching processing between multiple users based on the estimated user profile results in any one of the first to ninth aspects.
[0047] According to the tenth embodiment, by performing matching processing based on the estimated results of the user profile, matching processing can be performed only on users whose preferences, tendencies, or characteristics are the same or similar, thereby optimizing matching and shortening processing time.
[0048] The information processing method according to the 11th aspect of this disclosure further involves estimating a representative user within a group to which multiple users belong, based on the user profile estimation result, and performing matching processing between the multiple representative users based on the user profile estimation result for the representative user.
[0049] According to the 11th embodiment, since matching can be performed on a group basis by matching representative users, the computational cost of the matching process can be reduced compared to user-based matching that targets all users.
[0050] An information processing device according to a twelfth aspect of this disclosure comprises a circuit configuration which acquires first data relating to user behavior, generates frequency distribution information relating to the first data, generates second data to which flag information is selectively assigned based on the first data by performing a masking process based on the frequency distribution information, estimates a user profile including the user's preferences, tendencies, or characteristics based on the second data, and outputs the estimated result of the user profile.
[0051] According to the twelfth embodiment, the estimation accuracy of user profile estimation can be improved by estimating a user profile based on second data to which flag information has been selectively assigned to first data by masking.
[0052] A program according to a 13th aspect of this disclosure is a program for causing an information processing device to perform processing, wherein the processing includes acquiring first data relating to user behavior, generating frequency distribution information relating to the first data, performing mask processing based on the frequency distribution information to generate second data to which flag information is selectively assigned, estimating a user profile including the user's preferences, tendencies, or characteristics based on the second data, and outputting the estimated result of the user profile.
[0053] According to the 13th embodiment, the estimation accuracy of user profile estimation can be improved by estimating a user profile based on second data to which flag information has been selectively assigned to first data by masking.
[0054] This disclosure can also be implemented as a program that causes a computer to execute each characteristic configuration included in such a method or apparatus, or as a system that operates using such a program. It goes without saying that such a computer program can be distributed via a computer-readable, non-temporary recording medium such as a CD-ROM, or via a communication network such as the Internet.
[0055] (Embodiments of the present disclosure) Embodiments of this disclosure will be described in detail below with reference to the drawings. Elements denoted by the same reference numeral in different drawings refer to the same or corresponding elements. Furthermore, the components, their arrangement, connection configurations, and operating sequences shown in the following embodiments are examples and are not intended to limit this disclosure. This disclosure is limited only by the claims. Therefore, among the components in the following embodiments, those not described in the independent claims representing the highest-level concepts of this disclosure are described as constituting a more preferable configuration, even though they are not necessarily required to achieve the object of this disclosure.
[0056] Figure 1 is a simplified diagram showing an example configuration of an information processing system 1 according to an embodiment of this disclosure. The information processing system 1 is configured to include a terminal 11 and an estimation device 12.
[0057] Terminal 11 is, for example, a storage device that stores data relating to the user's actions (hereinafter referred to as "first data D1"). First data D1 includes cyber information obtained through activities on the Web (such as user searches or computer games like e-sports) and physical information obtained through activities in the real environment, such as the number of steps taken. First data D1 is, for example, a data file that exists in the local environment, but its form is not particularly limited. Terminal 11 inputs the first data D1, which is stored in a predetermined format, to the estimation device 12. The estimation device 12 may read first data D1 from a data file that exists in the local environment, or it may read first data D1 from log data that exists in the cloud.
[0058] The estimation device 12 estimates a user profile based on a plurality of first data D1 input from the terminal 11. The user profile includes the user's preferences, tendencies, or characteristics. The estimation device 12 estimates the user profile, for example, by inputting the input data based on the first data D1 into a machine learning-based estimation model.
[0059] The estimation device 12 is configured using a computer that includes a processing unit 21, a storage unit 22, and a communication unit 23.
[0060] The processing unit 21 comprises the circuit configuration of an information processing device. The information processing device includes a processor such as a CPU.
[0061] The storage unit 22 is configured to include a storage device for storing data. The storage unit 22 includes a computer-readable non-volatile storage medium such as a hard disk drive, a solid-state drive, or flash memory.
[0062] The communication unit 23 is an arbitrary data input / output mechanism such as an interface circuit, and is configured to include, for example, a communication module that corresponds to the communication standard between the terminal 11 and the estimation device 12.
[0063] Figure 2 shows the contents stored in the memory unit 22. The memory unit 22 stores the first data D1 input from the terminal 11. The memory unit 22 also stores the program 30.
[0064] The first data, D1, includes cyber information and physical information. Cyber information includes log data related to user actions such as clicks acquired on the application. However, the first data, D1, is not limited to log data and may also include image data or audio data.
[0065] In the following section, the processing of the estimation device 12 will be explained using log data acquired on the app. However, the input data entered into the estimation device 12 from the terminal 11 is not limited to log data; the estimation device 12 should perform processing corresponding to the data format, such as image data or audio data.
[0066] If the first data D1 is image data, the upper limits of its parameters and its data distribution are affected by the sensor used to acquire the data. To illustrate with an example using 8-bit image data, since the pixel value range for 8-bit images is "0" to "255", there is a high probability that noise and outliers will converge (saturate) at the upper limit "255" or the lower limit "0". Therefore, when using 8-bit image data, from the perspective of reducing the influence of outliers, it is sufficient to remove the data at the upper limit "255" and the lower limit "0". This removes the influence of outliers concentrated at the upper or lower limit of the pixel value range, allowing for proper normalization and enabling the unified handling of information from multiple data with different properties.
[0067] If the first data D1 is audio data, then data processing should be based on the dynamic range. For example, in the case of audio data heard by humans, it is generally known that, due to human characteristics, sounds above the upper limit of the dynamic range (approximately 120 dB for humans) tend to be painful, and sounds below the lower limit of the dynamic range (approximately 30 dB for humans) tend to be heard as noise. Therefore, values that people find unpleasant or difficult to hear can be considered outliers, and by removing data corresponding to the upper or lower limit of the dynamic range as outliers, the impact of outliers can be mitigated.
[0068] Furthermore, since sound is a vibration of air, the audio data detected by sensors such as microphones will not be constant even when measuring sounds at the sensor's sensitivity limit. Therefore, when removing outliers from audio measurement data, it is sufficient to add a width of a few dB (for example, 2 dB) to the boundary where the outlier occurs and then remove the outlier. This removes the influence of outliers in audio data based on the human dynamic range, or outliers mixed into audio data acquired by sensors, and allows for proper normalization, enabling the unified handling of information from multiple data with different properties.
[0069] Figure 3 is a simplified diagram showing the functional configuration of the processing unit 21A as a first example of the functional configuration of the processing unit 21. The processing unit 21A has an acquisition unit 31, a generation unit 32, a mask processing unit 33, an estimation unit 34, and an output unit 35, which are functions realized by the processor executing the program 30 read from the storage unit 22. Note that the functions shown in Figure 3 may also be configured using dedicated hardware circuits such as ASICs.
[0070] The acquisition unit 31 acquires the first data D1 relating to the user's actions by reading the first data D1 from the storage unit 22.
[0071] The generation unit 32 generates frequency distribution information for the first data D1. The frequency distribution information is information that shows the distribution of the input data, and includes, for example, a histogram.
[0072] The mask processing unit 33 generates second data D2 in which flag information is selectively assigned to first data D1 by performing mask processing based on frequency distribution information. In this disclosure, mask processing means the process of selectively assigning flag information to data to be processed in order to select a specific part of the data to be processed and perform processing such as processing or editing. The flag information may be, for example, a 1-bit flag information of "1" or "0". Selectively assigning flag information means choosing whether to assign flag information (i.e., assigning flag information of "1") or not assigning flag information (i.e., assigning flag information of "0"). Therefore, second data D2 includes second data D2 in which flag information is assigned to first data D1 and second data D2 in which flag information is not assigned to first data D1. Note that the relationship between assigning or not assigning flag information and the flag information of "1" or "0" may be the reverse of the above example.
[0073] The estimation unit 34 estimates a user profile, including user preferences, tendencies, or characteristics, based on the second data D2. The estimation unit 34 estimates the user profile, for example, by inputting input data based on the second data D2 into a machine learning-based estimation model.
[0074] Here, the estimation unit 34 determines whether or not to use the second data D2 for user profile estimation based on flag information related to the second data D2. For example, the mask processing unit 33 generates the second data D2 by assigning flag information to each of the multiple first data D1s, for first data D1s that contain a singular portion in their frequency distribution information, and not assigning flag information to first data D1s that do not contain a singular portion in their frequency distribution information. In generating the second data D2, the mask processing unit 33 determines a predetermined singular frequency distribution shape in the frequency distribution information as a singular portion. The estimation unit 34 executes the user profile estimation process by not using the second data D2 with flag information (i.e., data containing a singular portion) for user profile estimation, and using the second data D2 without flag information (i.e., data that does not contain a singular portion) for user profile estimation.
[0075] The output unit 35 outputs data indicating the user profile estimation result by the estimation unit 34. For example, the data indicating the estimation result is input to a display device (not shown in the figure) via the communication unit 23, and the display device displays the user profile estimation result.
[0076] Figure 4 is a simplified diagram showing the functional configuration of the processing unit 21B as a second example of the functional configuration of the processing unit 21. The processing unit 21B has an acquisition unit 31, a generation unit 32, a mask processing unit 33, a normalization unit 41, an estimation unit 34, and an output unit 35, which are functions realized by the processor executing the program 30 read from the storage unit 22. Note that the functions shown in Figure 4 may also be configured using dedicated hardware circuits such as ASICs.
[0077] Similarly, the acquisition unit 31 acquires first data D1 related to the user's behavior. The generation unit 32 generates frequency distribution information related to the first data D1. The masking unit 33 generates second data D2 in which flag information is selectively attached to the first data D1 by performing masking based on the frequency distribution information.
[0078] The normalization unit 41 generates the third data D3 by normalizing the second data D2 input from the mask processing unit 33. Normalization includes the process of bringing the values of each of the multiple data (corresponding to the scale of the vertical axis of the histogram) to a specified range, for example, "0" to "1".
[0079] Figure 5 shows the details of the function of the normalization unit 41. The normalization unit 41 includes an estimation unit 411, a calculation unit 412, and a normalization unit 413. The estimation unit 411 estimates the data distribution of the second data D2 from a plurality of pre-set data distributions. That is, it applies one of the plurality of pre-set data distributions to the second data D2. The plurality of data distributions may include discrete distributions and continuous distributions. Alternatively, the plurality of data distributions may include at least two of the following: normal distribution, uniform distribution, Poisson distribution, binary data, monotonically increasing distribution, and monotonically decreasing distribution. The calculation unit 412 calculates parameters used in the normalization calculation (hereinafter referred to as "normalization parameters") based on the data distribution of the second data D2. The normalization parameters may include at least one of the following: a first threshold greater than the minimum value of the second data D2, and a second threshold less than the maximum value of the second data D2. The normalization unit 413 normalizes the second data D2 in a manner corresponding to the data distribution estimated by the estimation unit 411, and using the normalization parameters calculated by the calculation unit 412.
[0080] Here, the normalization unit 41 decides whether or not to normalize the second data D2 based on flag information related to the second data D2. For example, the mask processing unit 33 generates the second data D2 by assigning flag information to each of the multiple first data D1s, for first data D1s that include a feature portion representing the user's characteristics, and not assigning flag information to first data D1s that do not include a feature portion. The normalization unit 41 generates the third data D3 by not normalizing the second data D2 that has been assigned flag information (i.e., the data that includes the feature portion) and normalizing the second data D2 that does not have flag information (i.e., the data that does not include the feature portion).
[0081] Referring to Figure 4, the estimation unit 34 estimates the user profile based on the third data D3 input from the normalization unit 41. The output unit 35 outputs data showing the user profile estimation result by the estimation unit 34.
[0082] Figure 6 shows a simplified representation of the functional configuration of the processing unit 21C as a third example of the functional configuration of the processing unit 21. The processing unit 21C has an acquisition unit 31, a generation unit 32, a mask processing unit 33, a conversion unit 42, an estimation unit 34, and an output unit 35, which are functions realized by the processor executing the program 30 read from the storage unit 22. Note that the functions shown in Figure 6 may also be configured using dedicated hardware circuits such as ASICs.
[0083] Similarly, the acquisition unit 31 acquires first data D1 related to the user's behavior. The generation unit 32 generates frequency distribution information related to the first data D1. The masking unit 33 generates second data D2 in which flag information is selectively attached to the first data D1 by performing masking based on the frequency distribution information.
[0084] The transformation unit 42 generates the fourth data D4 by transforming the range of the second data D2 input from the mask processing unit 33. The range of the data represents the range in which the data is distributed and corresponds to the scale of the horizontal axis of the histogram.
[0085] Here, the conversion unit 42 determines whether or not to convert the range of the second data D2 based on flag information related to the second data D2. For example, the mask processing unit 33 generates the second data D2 by assigning flag information to each of the multiple first data D1s whose data range differs from a predetermined reference range, and by not assigning flag information to first data D1s whose data range is the same as the reference range. The reference range is, for example, the range of the main data used for user profile estimation. The conversion unit 42 generates the fourth data D4 by converting the range of the second data D2 to which flag information has been assigned to the reference range, and by not converting the range of the second data D2 to which flag information has not been assigned.
[0086] The estimation unit 34 estimates the user profile based on the fourth data D4 input from the conversion unit 42. The output unit 35 outputs data showing the user profile estimation result by the estimation unit 34.
[0087] Figure 7 shows a simplified representation of the functional configuration of the processing unit 21D as a fourth example of the functional configuration of the processing unit 21. The processing unit 21D, which is realized by the processor executing the program 30 read from the storage unit 22, has an acquisition unit 31, a generation unit 32, a determination unit 43, a deletion unit 44, a mask processing unit 33, a normalization unit 41, a conversion unit 42, an estimation unit 34, a matching processing unit 45, and an output unit 35. Note that the functions shown in Figure 7 may be configured using dedicated hardware circuits such as ASICs. In addition, the determination unit 43, the deletion unit 44, and the matching processing unit 45 may be added to processing units 21A to 21C.
[0088] Similarly, the acquisition unit 31 acquires first data D1 related to the user's behavior. The generation unit 32 generates frequency distribution information related to the first data D1.
[0089] The determination unit 43 determines the singular portion included in the frequency distribution information. The singular portion is a unique data portion in the distribution of the input data that is not expected from the original purpose of using the application, and includes outliers, abnormal values, or irregular portions. In determining the singular portion, the determination unit 43 determines a predetermined singular frequency distribution shape in the frequency distribution information as a singular portion. The singular frequency distribution shape includes the peak shape of the frequency distribution near the minimum value of the first data D1. The singular frequency distribution shape also includes a flat shape of the frequency distribution.
[0090] The deletion unit 44 generates the fifth data D5 by deleting data included in the singular portion of the first data D1 (hereinafter referred to as "singular data"). In generating the fifth data D5, the deletion unit 44 deletes singular data if predetermined deletion conditions are met. The deletion conditions include at least one of the following related to frequency distribution information: number of peaks, range, data attributes, and frequency distribution shape. The data attributes include physical information obtained from activities in real space and cyber information obtained from activities in virtual space.
[0091] The mask processing unit 33 generates second data D2 in which flag information is selectively assigned to the fifth data D5 by performing mask processing based on frequency distribution information. However, if the first data D1 does not contain singular parts, or if the first data D1 contains singular parts but does not satisfy the predetermined deletion conditions, the singular data will not be deleted, and the mask processing unit 33 may generate second data D2 in which flag information is selectively assigned to the first data D1.
[0092] The normalization unit 41 generates the third data D3 by normalizing the second data D2 input from the mask processing unit 33. Here, the normalization unit 41 decides whether or not to normalize the second data D2 based on the normalization flag included in the flag information related to the second data D2. For example, the mask processing unit 33 generates the second data D2 by assigning a normalization flag to each of the multiple fifth data D5 (or first data D1) that includes a feature portion representing the user's characteristics, and not assigning a normalization flag to the fifth data D5 that does not include a feature portion. The normalization unit 41 generates the third data D3 by not normalizing the second data D2 that has been assigned a normalization flag, and normalizing the second data D2 that has not been assigned a normalization flag.
[0093] The conversion unit 42 generates the fourth data D4 by converting the range of the third data D3. Here, the conversion unit 42 decides whether or not to convert the range of the third data D3 based on the conversion flag included in the flag information related to the third data D3. For example, the mask processing unit 33 generates the second data D2 by assigning a conversion flag to each of the multiple fifth data D5 (or first data D1) whose data range is different from a predetermined reference range, and not assigning a conversion flag to the fifth data D5 whose data range is the same as the reference range. The conversion unit 42 generates the fourth data D4 by converting the range of the third data D3 to the reference range for which the conversion flag has been assigned, and not converting the range of the third data D3 for which the conversion flag has not been assigned.
[0094] The estimation unit 34 estimates the user profile based on the fourth data D4 input from the conversion unit 42. Here, the estimation unit 34 decides whether or not to use the fourth data D4 for estimating the user profile based on the deletion flag included in the flag information related to the fourth data D4. For example, the mask processing unit 33 generates the second data D2 by assigning a deletion flag to each of the multiple fifth data D5 (or first data D1) that contains a singular portion in the frequency distribution information of the fifth data D5, and not assigning a deletion flag to the fifth data D5 that does not contain a singular portion in the frequency distribution information. The estimation unit 34 performs the user profile estimation process by not using the fourth data D4 that has the deletion flag assigned to it for estimating the user profile, and using the fourth data D4 that does not have the deletion flag assigned to it for estimating the user profile.
[0095] The matching processing unit 45 performs matching processing to match multiple users with each other, or matching users with content, based on the estimated user profile results. Alternatively, the matching processing unit 45 estimates a representative user within a group to which multiple users belong, based on the estimated user profile results, and performs matching processing to match multiple representative users with each other, based on the estimated user profile results for the representative user. While the matching processing can involve calculating Pearson's correlation coefficient and using the result as a similarity score, it is not limited to this method and other methods may be used. Furthermore, the matching processing unit 45 may perform matching processing on users with identical or similar user profiles, or it may perform matching processing on all users included in the input data.
[0096] The output unit 35 outputs data indicating the result of the matching process performed by the matching processing unit 45.
[0097] The configuration of the estimation device 12 is not particularly limited; for example, it may be configured using an edge server installed within a specific facility, or it may be configured using a cloud server. When the estimation device 12 is configured using an edge server, the terminal 11 and the estimation device 12 are connected via a local area network. When the estimation device 12 is configured using a cloud server, the terminal 11 and the estimation device 12 are connected via a wide-area communication network such as the Internet. Furthermore, a portion of the estimation device 12 may be configured using an edge server, while the other portion is configured using a cloud server.
[0098] Furthermore, the estimation device 12 does not necessarily have to be implemented using a single computer device, but may be implemented by a distributed processing system including a terminal device and a server device. In this case, for example, with respect to the processing unit 21D, the acquisition unit 31, generation unit 32, determination unit 43, deletion unit 44, and storage unit 22 may be provided in the terminal device, and the mask processing unit 33, normalization unit 41, conversion unit 42, estimation unit 34, matching processing unit 45, and output unit 35 may be provided in the server device. In this case, the transmission and reception of data between the components may be performed via a wide-area communication network.
[0099] The operation of the processing unit 21 (information processing device) according to the embodiment of this disclosure will be described below, using the processing unit 21D shown in Figure 7 as an example.
[0100] Figure 8 is a flowchart showing the processing steps performed by the processing unit 21D.
[0101] First, in step S1, the acquisition unit 31 acquires multiple first data D1s related to the user's actions by reading the first data D1 from the storage unit 22.
[0102] Next, in step S2, the generation unit 32 generates frequency distribution information for the first data D1. The frequency distribution information is information that shows the distribution of the input data, and includes, for example, a histogram.
[0103] Figure 9 shows the first example of a histogram. The horizontal axis represents the number of clicks on the user profile screen on the app, which is cyber information, and the vertical axis represents the number of users, which is the frequency for each bin. The data distribution shows a normal distribution with peak shapes corresponding to bins where the horizontal axis values are between "20" and "25".
[0104] Figure 10 shows a second example of a histogram. The horizontal axis represents the user's daily step count, which is physical information, and the vertical axis represents the number of users, which is the frequency for each bin. The data distribution shows a monotonically decreasing distribution with peak shapes corresponding to the bins near the minimum value where the horizontal axis value is between "0" and "2500".
[0105] Figure 11 shows a third example of a histogram. The horizontal axis represents the number of clicks on the user profile screen on the app, which is cyber information, and the vertical axis represents the number of users, which is the frequency for each bin. The data distribution shows a normal distribution with peak shapes corresponding to the bins near the minimum value where the horizontal axis value is between "0" and "5", and peak shapes corresponding to the bins where the horizontal axis value is between "20" and "25".
[0106] Figure 12 shows a fourth example of a histogram. The horizontal axis represents the number of clicks on the user profile screen on the app, which is cyber information, and the vertical axis represents the number of users, which is the frequency for each bin. The data distribution shows a normal distribution with a peak shape corresponding to the bins where the horizontal axis values are between "20" and "25", and a flat shape corresponding to the five bins where the horizontal axis values are between "35" and "60". The flat shape represents a distribution shape in which the data continues to be distributed at a constant frequency across multiple consecutive bins.
[0107] Referring to Figure 8, in step S3, the determination unit 43 determines the singular portion included in the histogram generated in step S2.
[0108] If the distribution of data obtained from an app contains an anomaly, the anomaly data within that anomaly distorts the overall distribution of the data. When using data containing this anomaly, the accuracy of user profile estimation decreases for users who use the app in accordance with its intended usage.
[0109] Furthermore, a given dataset may contain multiple data distributions based on different factors. In this case, the interaction of these multiple data distributions can cause the unique features of each data point to be obscured by singularities, leading to a decrease in the accuracy of user profile estimation.
[0110] Therefore, it is important to exclude data that negatively impacts the analysis from the data distribution, and the determination unit 43 has the effect of detecting data that should be excluded from the data distribution by determining the singular portion.
[0111] In determining the singular portion, the determination unit 43 determines that a predetermined singular frequency distribution shape in the histogram is a singular portion.
[0112] As a first example, the determination unit 43 determines that the peak shape of the frequency distribution near the minimum value of the first data D1 is a singular frequency distribution shape. The method for detecting peaks on the graph is to compare the frequency of a histogram bin with the frequencies of the preceding and succeeding bins. For example, if the difference between the frequency of a certain bin and the frequencies of the preceding and succeeding bins is greater than a set threshold (e.g., "5"), it can be determined to be a peak. For the bin containing the minimum value of the first data D1, if the difference between it and the frequency of the bin to its right is greater than a set threshold (e.g., "5"), it can be determined to be a peak. The determination unit 43 counts the number of peaks included in the histogram.
[0113] For the histogram shown in Figure 10, which has peaks only near the minimum value, the determination unit 43 determines that the peak shape of the frequency distribution near the minimum value of the first data D1, that is, the peak shape corresponding to the leftmost bin where the horizontal axis value is "0" to "2500", is a unique frequency distribution shape. The leftmost bin includes not only data of normal values such as 1000 steps or 2000 steps per day, but also data of abnormal values that are not reasonable for the number of steps per day, such as a few steps or a few tens of steps. The determination unit 43 determines that the peak shape corresponding to the bin near the minimum value of the first data D1 is a unique frequency distribution shape.
[0114] For the histogram with multiple peaks shown in Figure 11, the determination unit 43 determines that the peak shape of the frequency distribution near the minimum value of the first data D1, that is, the peak shape corresponding to the leftmost bin where the horizontal axis value is "0" to "5", is a unique frequency distribution shape. The histogram shown in Figure 11 contains data from users who continuously use the app and data from users who do not use the app after viewing the tutorial, and there are multiple peaks on the graph. In the data from users who continuously use the app, a peak may appear at a specific location other than near the minimum value. On the other hand, in the data from users who do not use the app, a peak appears near the minimum value. The determination unit 43 determines that the peak shape corresponding to the bin near the minimum value of the first data D1 is a unique frequency distribution shape.
[0115] As a second example, the determination unit 43 determines that a flat shape included in a normal distribution is a singular frequency distribution shape. Targeting a histogram with a flat shape as shown in Figure 12, the determination unit 43 determines that a flat shape included in a normal distribution is a singular frequency distribution shape. Log data of users who frequently used the app in the past but do not currently use the app may remain on the histogram as a fixed flat shape. When a flat shape exists, the distribution of the entire data is distorted towards the side where the flat shape exists. For example, if the distribution of user app usage data according to the expected app usage method follows a normal distribution, the data in the normal distribution may be distorted by the flat shape, potentially obscuring the characteristics of app users and reducing the estimation accuracy of user profile estimation. Therefore, the determination unit 43 determines that a fixed flat shape included in a normal distribution is a singular frequency distribution shape. A method for detecting a flat shape on a graph is to compare the frequency of the histogram bin with the frequencies of the preceding and succeeding bins. For example, if the difference between the frequency of a given bottle and the frequencies of the preceding and succeeding bottles is less than a set threshold (e.g., "2"), then that bottle and the preceding and succeeding bottles can be determined to have the same frequency. In this frequency comparison with the preceding and succeeding bottles, if the number of consecutive bottles that are determined to have the same frequency as the preceding and succeeding bottles is equal to or greater than a threshold (e.g., "5"), then the shape can be determined to be flat.
[0116] Referring to Figure 8, in step S4, the deletion unit 44 generates the fifth data D5 by deleting the singular data included in the singular portion of the first data D1.
[0117] By removing outlier data, the remaining fifth data point D5 can be fitted to a specific data distribution, enabling effective normalization. The fifth data point D5 is likely to contain a high proportion of data from users who use the app in its intended way. Calculating normalization parameters based on this fifth data point D5 and performing normalization improves the accuracy of user profile estimation.
[0118] Figure 13 is a flowchart showing the details of the process for deleting outlier data in step S4, as in the first example. The first example corresponds to the process of deleting the peak shape near the minimum value of the first data D1.
[0119] First, in step S4A1, the deletion unit 44 obtains the number of peaks included in the histogram, which were counted by the determination unit 43 in step S3.
[0120] Next, in step S4A2, the deletion unit 44 determines whether or not there is a bin near the minimum value of the first data D1 among the bins that represent peaks on the histogram detected in step S3. In other words, it determines whether or not a bin where the horizontal axis value of the first data D1 is near "0" is a peak.
[0121] The determination method involves checking whether the data in the bins near the minimum value contains "0". Alternatively, the determination can be made by checking whether the smallest data in the bin closest to "0" among the bins determined to be peaks is below a threshold (e.g., "1"). If none of the detected bins contain a peak near the minimum value (step S4A2: NO), the process of deleting outlier data is terminated.
[0122] If the detected bins include a bin that has a peak near the minimum value (step S4A2: YES), then in step S4A3, the deletion unit 44 determines whether there are two or more peaks on the graph based on the number of peaks obtained in step S4A1.
[0123] If there are no more than two peaks on the graph (step SS4A3: NO), then in step S4A4, the deletion unit 44 determines whether the frequency difference between the bin near the minimum value and the adjacent bin is greater than or equal to a predetermined threshold. In other words, for a data distribution where peaks exist only near the minimum value, it determines whether there is an excessive amount of data mixed into the bin near the minimum value compared to what would be expected from the app usage logs of continuous app users.
[0124] The deletion unit 44 determines, for example, that if all input data are positive values, and the frequency difference between the bin near "0" and the bin to its right is greater than or equal to a threshold (for example, a difference equivalent to twice the value), then the bin near "0" is excessively populated with data. On the other hand, if the frequency difference between the bin near "0" and the bin to its right is less than the threshold, the deletion unit determines that the bin near "0" is not excessively populated with data and terminates the deletion process for the outlier data. The deletion unit 44 may also compare the frequency difference between the bin near "0" and the bin to its left if all input data are negative values, or it may compare the frequency difference between the bin near "0" and both the bins to its right and left if the input data includes both positive and negative values.
[0125] If there are two or more peaks on the graph (step SS4A3: YES), or if the frequency difference between a bin near "0" and an adjacent bin is greater than or equal to a threshold (step SS4A4: YES), then in step S4A5, the deletion unit 44 determines whether the maximum value on the horizontal axis of the input data is greater than or equal to a predetermined threshold (e.g., "7"). If the maximum value is greater than or equal to the threshold, the input data is likely to be physical information; if the maximum value is less than the threshold, the input data is likely to be cyber information. The threshold may be set to an arbitrary value in advance depending on the input data.
[0126] If the maximum value is less than the threshold (step S4A5: NO), the process of deleting outlier data is terminated. Cyber information is, for example, the frequency of use of a specific function of an application over a week, and since the data distribution exists only within a narrow range, the impact on the overall data distribution is small even if outlier data is not deleted.
[0127] If the maximum value is greater than or equal to the threshold (step S4A5: YES), then in step S4A6, the deletion unit 44 determines whether the bins near "0" contain multiple values. For example, as shown in the histogram in Figure 10, if the bins near "0" contain multiple values across a wide range such as "0" to "2500", then the bins near "0" contain not only data of normal values such as 1000 steps or 2000 steps per day, but also data of abnormal values that are not appropriate for a daily step count, such as a few steps or a few tens of steps per day.
[0128] If the bins near "0" contain multiple values (step S4A6: YES), then in step S4A7, the deletion unit 44 deletes data below a predetermined threshold (e.g., "100") for the bins near "0", and terminates the process of deleting outlier data. The threshold may be set based on the maximum value of the input data, or an arbitrary value may be set in advance.
[0129] If the bins near "0" do not contain multiple values (step S4A6: NO), then in step S4A8, the deletion unit 44 deletes the data contained in the bins near "0" and terminates the process of deleting singular data.
[0130] Figure 14 is a flowchart showing a second example of the details of the process for deleting outlier data in step S4. The second example corresponds to the process of deleting flat shapes that are included in a normal distribution.
[0131] First, in step S4B1, the deletion unit 44 determines whether the data distribution of the histogram of the input data fits a uniform distribution.
[0132] If the data distribution is uniform (Step S4B1: YES), the process of deleting outlier data is terminated. When the data distribution is uniform, the frequencies of each bin on the histogram will be approximately the same, and it will be judged as having a flat shape. However, this indicates a characteristic of the input data and should not be deleted as outlier data.
[0133] If the data does not fit a uniform distribution (step S4B1: NO), the deletion unit 44 then determines in step S4B2 whether the maximum value on the horizontal axis of the input data is greater than or equal to a predetermined threshold (e.g., "7"). If the maximum value is greater than or equal to the threshold, the input data is likely to be physical information; if the maximum value is less than the threshold, the input data is likely to be cyber information. The threshold may be set to an arbitrary value in advance depending on the input data.
[0134] If the maximum value is less than the threshold (step S4B2: NO), the process of deleting outlier data is terminated. Cyber information is, for example, the frequency of use of a specific function of an application over a week, and since the data distribution exists only within a narrow range, the impact on the overall data distribution is small even if outlier data is not deleted.
[0135] If the maximum value is greater than or equal to the threshold (step S4B2: YES), then in step S4B3, the deletion unit 44 deletes flat shapes included in the histogram. Specifically, the deletion unit 44 deletes data from multiple consecutive bins where the difference between the frequency of a given bin and the frequencies of the preceding and succeeding bins is less than or equal to a set threshold (for example, "2"), as singular data.
[0136] Referring to Figure 8, in step S5, the mask processing unit 33 generates second data D2 to which flag information is selectively assigned to the fifth data D5 by performing mask processing based on frequency distribution information. Mask processing includes the process of assigning predetermined flag information to the data to be calculated. Data to which flag information is assigned is processed based on that flag information. For example, data to which flag information is assigned is not used in calculations, while data without flag information is used in calculations. Alternatively, data to which flag information is assigned is used in calculations, while data without flag information is not used in calculations. Note that data processing based on flag information is not limited to the above example, and other data processing methods may be used. Also, if the first data D1 does not contain singular parts, or if the first data D1 contains singular parts but does not satisfy predetermined deletion conditions, the singular data is not deleted in step S4, and therefore the mask processing unit 33 may generate second data D2 to which flag information is selectively assigned to the first data D1.
[0137] As a first example of the targets for flag information, the flag information includes the deletion flag mentioned above. The mask processing unit 33 assigns a deletion flag to each of the multiple fifth data D5 (or first data D1) if the fifth data D5 contains a singular portion in the frequency distribution information, and does not assign a deletion flag to the fifth data D5 if the frequency distribution information does not contain a singular portion. The mask processing unit 33 only needs to determine whether or not the frequency distribution information contains a singular portion by using the determination result from the determination unit 43 in step S3.
[0138] As a second example regarding the target of flag information, the flag information includes the normalization flag mentioned above. For each of the multiple fifth data D5 (or first data D1), the mask processing unit 33 assigns a normalization flag to the fifth data D5 that contains a feature portion representing the user's characteristics, and does not assign a normalization flag to the fifth data D5 that does not contain a feature portion. In other words, the mask processing unit 33 assigns a normalization flag to data that contains a portion indicating an individual's characteristics in user profile estimation.
[0139] The first example of a part that indicates individual characteristics is a parameter that becomes an outlier due to the user's characteristics at the upper limit, when there is no upper limit to the characteristics of the acquired data. For example, data with a standard score of 100 is an outlier in the high parameter range, and this is data that strongly indicates individual characteristics.
[0140] A second example concerning individual characteristics is a parameter that deviates from the overall trend of the data distribution, even if it does not deviate to the extent of being an outlier in the data distribution. For example, regarding the number of times a user presses a stamp to indicate a reaction in the app, this data represents a user who presses stamps a large amount, which is an outlier compared to the overall trend of users, because that user frequently presses stamps. In both the first and second examples above, the data distribution shape is either a shape with two peaks in the normal distribution, or a shape where outlier data is present after a peak in the normal distribution. The masking unit 33 should perform masking on such data.
[0141] One example of a feature on the graph of data to be masked is a bin that is isolated in a histogram. However, the features on the graph to be masked are not limited to bins that are isolated in a histogram; the masking unit 33 may determine which data to mask based on any other feature.
[0142] Figure 15 shows a fifth example of a histogram. The fifth example is a histogram that includes a single bin that is isolated. The histogram includes a single bin that is isolated in the range of click counts from "60" to "65". Data contained in such isolated bins is thought to represent strong individual user characteristics, such as user preferences, skills, or behaviors. For example, if a user's test score is significantly higher than the overall trend, that score data represents a characteristic of that user.
[0143] As a method for determining whether or not a single bin located in an isolated area is included, the mask processing unit 33 can determine that a bin is located in an isolated area if, when a bin on the histogram is not closely adjacent to the bins before and after it, the distance between that bin and the bins before and after it is greater than or equal to a threshold (for example, "10").
[0144] Figure 16 shows a sixth example of a histogram. The sixth example is a histogram that includes multiple bins in a segregated area. The histogram includes two bins in a segregated area in the range of click counts from "55" to "65".
[0145] As a method for determining whether or not a histogram contains multiple bins that are in a detached area, the mask processing unit 33 can determine that multiple bins are in a detached area if, when a bin on the histogram is closely connected to one of the preceding or succeeding bins but not to the other, the distance between that bin and the next bin on the side that is not closely connected is greater than or equal to a threshold (e.g., "10"), and the number of bins that are continuously closely connected to that bin on the side that is closely connected is less than or equal to a threshold (e.g., "1").
[0146] A third example related to the part that indicates individual characteristics is data where there is a large difference in characteristics between individuals, that is, data with a wide range of values or data with a large variance. In the case of data with a large difference in characteristics between individuals, there is a high probability that outliers exist on the maximum or minimum side of the data.
[0147] One example of a criterion for whether or not the mask processing unit 33 assigns a normalization flag to the fifth data D5 is, as will be described later, setting the maximum value threshold of the second data D2 instead of the maximum value as the maximum value normalization parameter, or setting the minimum value threshold of the second data D2 instead of the minimum value as the minimum value normalization parameter. In such cases, the normalization flag may be assigned to the fifth data D5. By normalizing using the minimum value threshold and maximum value threshold for the second data D2, characteristic data will have a value greater than "1.0" after normalization, and can be emphasized. The assignment of a normalization flag by the mask processing unit 33 further emphasizes characteristic data, which can be used in user profile estimation and matching processes.
[0148] Specifically, masking allows for the emphasis of data that tends to reveal individual characteristics during user profile estimation, thereby improving the accuracy of user profile estimation. In matching processes, increasing the weight of elements that are important allows for appropriate matching of users with each other or with specific content. For example, in matching users with similar physical activities, if all data is treated uniformly, cyber information and physical information are treated as having equal value. In this case, the data is normalized so that the range of values for cyber information and physical information are approximately the same. However, in reality, physical data such as "daily step count" is more important than cyber data, which is cyber information, from the perspective of physical activity. Therefore, by assigning a normalization flag to physical data through masking and not normalizing the physical data, it is possible to emphasize the physical data more than the normalized cyber data. As a result, the influence of physical data can be greatly enhanced, making it possible to appropriately match users with similar physical activities.
[0149] As a third example of the targets for flag information, the flag information includes the above-mentioned conversion flag. The mask processing unit 33 assigns a conversion flag to each of the multiple fifth data D5 (or first data D1) if the data range of the fifth data D5 is different from a predetermined reference range, and does not assign a conversion flag to the fifth data D5 if the data range of the fifth data D5 is the same as the reference range. The fifth data D5 whose data range is the same as the reference range includes the main data used for user profile estimation and is hereinafter referred to as "main data" in this specification. The fifth data D5 whose data range is different from the reference range includes non-main data used for user profile estimation and is hereinafter referred to as "sub-data" in this specification.
[0150] Referring to Figure 8, in step S6, the normalization unit 41 generates the third data D3 by normalizing the second data D2 input from the mask processing unit 33. The normalization unit 41 normalizes the second data D2 for which the mask processing unit 33 did not assign a normalization flag in step S5, but does not normalize the second data D2 for which the mask processing unit 33 assigned a normalization flag in step S5.
[0151] Figure 17 is a flowchart showing the details of the normalization process.
[0152] First, in step S61, the estimation unit 411 estimates the data distribution of the second data D2 from a plurality of pre-set data distributions. In other words, it applies one of the plurality of pre-set data distributions to the second data D2. The estimation of the distribution may involve determining whether it is a discrete distribution that is more likely to apply to cyber information (for example, the number of user accesses to a specific function) or a continuous distribution that is more likely to apply to physical information (for example, user biometric data). The determination of whether it is a discrete or continuous distribution may be made using a differential quantity based on the shape of the histogram or a determination method based on the probability of bin occurrence, but is not limited to these methods.
[0153] Furthermore, the estimation unit 411 may estimate whether the data distribution of the second data D2 fits a specific data distribution (e.g., a normal distribution). The specific data distribution may be arbitrarily set in advance, and there is no particular limit to the number of data distributions that can be set. To estimate whether it fits a specific data distribution, a test corresponding to each data distribution may be used, or the estimation may be based on the shape of the histogram. The target of the distribution estimation is the fifth data D5 if the singularity removal process has been performed, and the first data D1 if the singularity removal process has not been performed.
[0154] In this embodiment, five data distributions applicable to log data obtainable on the application were pre-set as specific data distributions. The first data distribution is a Poisson distribution, which is highly likely to apply to discrete user behaviors (such as access to specific functions) that occur frequently, such as the frequency of application use. The second data distribution is a normal distribution, which is highly likely to apply to biometric data as physical information. The third data distribution is a uniform distribution, which is highly likely to apply to data where a certain value appears with the same frequency. The fourth data distribution is binary data, which obtains binary information indicating whether or not a specific function was used. The fifth data distribution is a monotonically increasing distribution, which shows a monotonically increasing shape, such as the cumulative number of accesses to a specific function. A monotonically decreasing distribution, which shows a monotonically decreasing shape, may be used instead of a monotonically increasing distribution. Furthermore, the data distributions listed above are just examples, and other data distributions (e.g., exponential distributions) may be used as candidates for the applicable data distribution. As for methods for estimating these predefined data distributions, when determining whether or not the data fits a specific data distribution, a test corresponding to that data distribution (for example, the Shapiro-Wilk test for a normal distribution) may be used. Alternatively, when determining whether or not the data fits a binary data, monotonically increasing distribution, or monotonically decreasing distribution, the determination may be made based on the shape of the graph.
[0155] The data distributions obtained from cyber or physical information through early use of an application can be limited to at most a few types. If a suitable data distribution can be estimated, it is possible to appropriately perform normalization processing using a method or normalization parameters appropriate to that data distribution. For example, if it can be estimated that the data is binary, it can be immediately determined that the maximum and minimum values of the input data should be used as normalization parameters. Also, if it can be estimated that the data is a normal distribution, for example, when normalizing to the range of "0" to "1", the input data can be preprocessed so that the peak is "0.5", and at that time, the mean value of the second data D2 can be obtained.
[0156] Referring to Figure 17, in step S62, the calculation unit 412 calculates the normalization parameters to be used in the normalization calculation based on the data distribution of the second data D2.
[0157] Figure 18 is a flowchart illustrating the details of the calculation process for the normalization parameters in step S62, as in a first example. The first example corresponds to the process when the estimated data distribution is discrete or continuous.
[0158] First, in step S62A1, the calculation unit 412 obtains the minimum and maximum values of the horizontal axis of the second data D2.
[0159] Next, in step S62A2, the calculation unit 412 determines the normalization range based on the maximum and minimum values obtained in step S62A1. The calculation unit 412 can determine the normalization range based on the signs of the maximum and minimum values. For example, if the signs of the maximum value and the minimum value are the same, the range from "0" to "1" can be determined as the normalization range. On the other hand, if the signs of the maximum value and the minimum value are different, the range from "-1" to "1" can be determined as the normalization range. Note that the normalization range is not limited to the above example; for example, the range from "-1" to "0" may also be determined as the normalization range.
[0160] Next, in step S62A3, the calculation unit 412 determines whether the data distribution of the second data D2 is a discrete distribution. The estimation result from the estimation unit 411 in step S61 can be used to determine the data distribution.
[0161] If the data distribution of the second data D2 is a discrete distribution (step S62A3: YES), then in step S62A4, the calculation unit 412 determines whether the maximum value of the second data D2 is less than a predetermined threshold (e.g., "7").
[0162] If the maximum value of the second data D2 is less than the threshold (step S62A4: YES), it is considered that there are no outliers in the second data D2. Therefore, the maximum and minimum values of the second data D2 are set as normalization parameters, and the calculation process for the normalization parameters is terminated. If the second data D2 is a discrete distribution and the maximum value of the second data D2 is less than the threshold, it is highly likely that the data values fall within a certain range and the maximum value is limited, such as the number of days a specific function of an app was used within a week. Therefore, if the maximum value of the second data D2 is less than the threshold, it is considered unlikely that there are outliers in the second data D2.
[0163] If the data distribution of the second data D2 is not discrete (step S62A3: NO), or if the maximum value of the second data D2 is greater than or equal to the threshold (step S62A4: NO), then in step S62A5, the calculation unit 412 calculates the interquartile range of the second data D2 to calculate the minimum value threshold and the maximum value threshold as outlier thresholds for the second data D2.
[0164] For example, the calculation unit 412 calculates the minimum threshold by subtracting the first quartile from the value obtained by multiplying the interquartile range by 1.5. The calculation unit 412 also calculates the maximum threshold by adding the value obtained by multiplying the interquartile range by 1.5 to the third quartile. Note that the multiplier for the interquartile range is not limited to 1.5; any multiplier may be used.
[0165] Next, in step S62A6, the calculation unit 412 sets the minimum value side normalization parameter and the maximum value side normalization parameter based on the minimum value side threshold and the maximum value side threshold calculated in step S62A5, using the minimum value side threshold and the maximum value side normalization parameter of the second data D2.
[0166] The calculation unit 412 compares the minimum value of the second data D2 with the minimum value threshold calculated in step S62A5. If the minimum value of the second data D2 is less than the minimum value threshold, the calculation unit 412 sets the minimum value threshold as the minimum value normalization parameter. On the other hand, if the minimum value of the second data D2 is greater than or equal to the minimum value threshold, the calculation unit 412 sets the minimum value of the second data D2 as the minimum value normalization parameter.
[0167] The calculation unit 412 compares the maximum value of the second data D2 with the maximum value threshold calculated in step S62A5. If the maximum value of the second data D2 exceeds the maximum value threshold, the calculation unit 412 sets the maximum value threshold as the maximum value normalization parameter. On the other hand, if the maximum value of the second data D2 is less than or equal to the maximum value threshold, the calculation unit 412 sets the maximum value of the second data D2 as the maximum value normalization parameter.
[0168] The maximum value side of the second data set D2 may contain characteristic data indicating users who actively use the app. When the maximum value threshold of the second data set D2 is used as the maximum value normalization parameter, that characteristic data will have a value greater than "1.0" after normalization. This has the effect of highlighting information about characteristic users by displaying a value greater than "1.0".
[0169] Figure 19 is a flowchart illustrating a second example of the details of the calculation process for the normalization parameter in step S62. The second example corresponds to the process when the estimated data distribution is a Poisson distribution, a normal distribution, a uniform distribution, binary data, or a monotonically increasing distribution.
[0170] First, in step S62B1, the calculation unit 412 obtains the minimum and maximum values of the horizontal axis of the second data D2.
[0171] Next, in step S62B2, the calculation unit 412 determines the normalization range based on the maximum and minimum values obtained in step S62A1. The calculation unit 412 can determine the normalization range based on the signs of the maximum and minimum values. For example, if the signs of the maximum value and the minimum value are the same, the range from "0" to "1" can be determined as the normalization range. On the other hand, if the signs of the maximum value and the minimum value are different, the range from "-1" to "1" can be determined as the normalization range. Note that the normalization range is not limited to the above example; for example, the range from "-1" to "0" may also be determined as the normalization range.
[0172] Next, in step S62B3, the calculation unit 412 determines whether the data distribution of the second data D2 is a normal distribution. The estimation result from the estimation unit 411 in step S61 can be used to determine the data distribution.
[0173] If the data distribution of the second data D2 is a normal distribution (step S62B3: YES), then in step S62B4, the calculation unit 412 obtains the mean value of the second data D2. When the peak of the data after normalization is at the center of the normalization range, the estimation accuracy of user profile estimation using the normalization result is improved. In the case of a normal distribution, the peak of the data and the mean value coincide, so if the second data D2 is a normal distribution, the mean value of the data is obtained.
[0174] If the data distribution of the second data D2 is not a normal distribution (step S62B3: NO), the processing in step S62B4 is omitted.
[0175] Next, in step S62B5, the calculation unit 412 determines whether the data distribution of the second data D2 is binary data or a uniform distribution. The estimation result from the estimation unit 411 in step S61 can be used to determine the data distribution.
[0176] If the data distribution of the second data D2 is binary or uniform (step S62B5: YES), it is considered that there are no outliers in the second data D2. Therefore, the maximum and minimum values of the second data D2 are set as normalization parameters, and the calculation process for normalization parameters is terminated.
[0177] If the data distribution of the second data D2 is neither binary nor uniform (step S62B5: NO), then in step S62B6, the calculation unit 412 calculates the interquartile range of the second data D2 and calculates the maximum value threshold as the outlier threshold for the second data D2.
[0178] For example, the calculation unit 412 calculates the maximum value threshold by adding a value obtained by multiplying the interquartile range by 1.5 to the third quartile. Note that the multiplier for the interquartile range is not limited to 1.5, and any multiplier may be used.
[0179] Next, in step S62B7, the calculation unit 412 sets the maximum value side normalization parameter based on the maximum value of the second data D2 and the maximum value side threshold calculated in step S62B6.
[0180] The calculation unit 412 compares the maximum value of the second data D2 with the maximum value threshold calculated in step S62B6. If the maximum value of the second data D2 exceeds the maximum value threshold, the calculation unit 412 sets the maximum value threshold as the maximum value normalization parameter. On the other hand, if the maximum value of the second data D2 is less than or equal to the maximum value threshold, the calculation unit 412 sets the maximum value of the second data D2 as the maximum value normalization parameter.
[0181] The maximum value side of the second data set D2 may contain characteristic data indicating users who actively use the app. When the maximum value threshold of the second data set D2 is used as the maximum value normalization parameter, that characteristic data will have a value greater than "1.0" after normalization. This has the effect of highlighting information about characteristic users by displaying a value greater than "1.0".
[0182] Next, in step S62B8, the calculation unit 412 determines whether the data distribution of the second data D2 is a Poisson distribution or a monotonically increasing distribution. The estimation result from the estimation unit 411 in step S61 can be used to determine the data distribution.
[0183] If the data distribution of the second data D2 is a Poisson distribution or a monotonically increasing distribution (step S62B8: YES), the minimum value of the second data D2 is set as the minimum value-side normalization parameter, and the calculation process for the normalization parameter is terminated. Here, in the Poisson distribution, the distribution becomes sparser as you move towards the maximum value, and there is a possibility that data of users with characteristics who use the app only casually may exist as outliers. On the other hand, data is concentrated towards the minimum value, but there is no data of users who actively use the app. Also, in the case of monotonically increasing distributions, it is thought that data of users who use the app casually is more abundant as you move towards the maximum value. For the above reasons, for Poisson distributions and monotonically increasing distributions, the process of obtaining the normalization parameter is performed only on the maximum value side.
[0184] If the data distribution of the second data D2 is neither a Poisson distribution nor a monotonically increasing distribution (step S62B8: NO), then in step S62B9, the calculation unit 412 calculates the interquartile range of the second data D2 and calculates the minimum value threshold as the outlier threshold for the second data D2.
[0185] For example, the calculation unit 412 calculates the minimum threshold by subtracting the first quartile from a value obtained by multiplying the interquartile range by 1.5. Note that the multiplier for the interquartile range is not limited to 1.5; any multiplier may be used.
[0186] Next, in step S62B10, the calculation unit 412 sets the minimum value side normalization parameter based on the minimum value of the second data D2 and the minimum value side threshold calculated in step S62B9.
[0187] The calculation unit 412 compares the minimum value of the second data D2 with the minimum value threshold calculated in step S62B9. If the minimum value of the second data D2 is less than the minimum value threshold, the calculation unit 412 sets the minimum value threshold as the minimum value normalization parameter. On the other hand, if the minimum value of the second data D2 is greater than or equal to the minimum value threshold, the calculation unit 412 sets the minimum value of the second data D2 as the minimum value normalization parameter.
[0188] In the example shown in Figure 19, steps S62B9 and S62B10 are performed only if the second data D2 follows a normal distribution. Since a normal distribution may extend to the minimum value side, the minimum value normalization parameter is set by comparing it with the minimum value threshold.
[0189] Referring to Figure 17, in step S63, the normalization unit 413 performs a normalization operation on the second data D2 using the normalization parameters calculated in step S62.
[0190] If the data distribution of the second data D2 estimated in step S61 is a normal distribution, the normalization unit 413 performs processing according to the normalization range determined in step S62. If the normalization range is "0" to "1", the normalization unit 413 adds or subtracts the mean of the second data D2 obtained in step S62 to the second data D2 so that the mode after normalization is "0.5". Alternatively, if the normalization range is "-1" to "1", the normalization unit 413 adds or subtracts the mean of the second data D2 obtained in step S62 to the second data D2 so that the mode after normalization is "0".
[0191] The normalization unit 413 normalizes the second data D2 using the normalization parameters set in step S62.
[0192] If the normalization range is "0" to "1", the normalization unit 413 performs normalization using formula (1).
[0193]
number
[0194] If the normalization range is "-1" to "1", the normalization unit 413 performs normalization using equation (2).
[0195]
number
[0196] Here, a i a' refers to the i-th data point of the second data set, D2. i This represents the result of the normalization operation. max This represents the maximum-side normalization parameter. min This represents the minimum value normalization parameter.
[0197] By normalizing the data using minimum and maximum thresholds for the second data set D2, characteristic data points will exceed "1.0" after normalization. This has the effect of highlighting information about distinctive users by displaying values above "1.0". Furthermore, since the normalization results in a distribution shape closer to that of users who continuously use the app, rather than a distribution skewed towards "0", it has the effect of improving the estimation accuracy of user profile estimation.
[0198] Referring to Figure 8, in step S7, the transformation unit 42 generates the fourth data D4 by transforming the range of the third data D3 input from the normalization unit 41. The transformation unit 42 transforms the range of the third data D3, which is subdata to which a transformation flag was assigned by the mask processing unit 33 in step S5, into a reference range, but does not perform a range transformation on the third data D3, which is main data to which a transformation flag was not assigned by the mask processing unit 33 in step S5.
[0199] Figure 20 is a diagram illustrating the range conversion process performed by the conversion unit 42. Figure 20(A) shows a histogram of the main data, which represents the number of users relative to the number of stamp posts. The range of the data corresponding to the distribution range on the horizontal axis is "0" to "40", and this range is the reference range mentioned above. Figure 20(B) shows a histogram of the sub-data, which represents the number of users relative to the number of clicks. The range of the data is "0" to "400". In step S5, the mask processing unit 33 does not assign a conversion flag to the main data, but assigns a conversion flag to the sub-data. The conversion unit 42 converts the range of the sub-data, which has a conversion flag assigned to it, to the reference range of the main data, which does not have a conversion flag assigned to it. Figure 20(C) shows a histogram of the sub-data after the range conversion. The range of the data has been converted from the range of "0" to "400" before conversion shown in Figure 20(B) to the same range of "0" to "40" as the reference range of the main data.
[0200] The conversion unit 42 performs the range conversion calculation using equation (3).
[0201]
number
[0202] Here, d i d' refers to the i-th data point of the third data point, D3. i This represents the result of the range transformation operation. max This represents the maximum value of the third data point D3, and d min∫ represents the minimum value of the third data point D3. M represents the maximum value of the fourth data point D4 after range transformation, and N represents the minimum value of the fourth data point D4. Note that for M, the maximum value of the main data point with the largest range among multiple main data points may be used, or it may not be limited to the maximum value, but may be other representative values such as the mean, median, or mode of the maximum values. Similarly, for N, the minimum value of the main data point with the largest range among multiple main data points may be used, or it may not be limited to the minimum value, but may be other representative values such as the mean, median, or mode of the minimum values. Furthermore, the range transformation is not limited to the calculation using equation (3), but may be performed using any other method.
[0203] In this way, by converting the value range of sub-data with a conversion flag to the reference value range, the value range of the sub-data and the reference value range of the main data can be aligned. As a result, the estimation accuracy of user profile estimation can be improved. In other words, in user profile estimation, having the same value range across multiple data sets used has the effect of improving the estimation accuracy of the user profile. For example, in user profile estimation, using a first data set with a value range of "0.3" to "0.7" and a second data set with a value range of "0.0" to "1.2" is less effective than using a first data set with a value range of "0.0" to "1.2" and a second data set with a value range of "0.0" to "1.2".
[0204] The conversion unit 42 converts the value range of the sub-data to which a conversion flag has been assigned to it, to the reference value range of the main data that has not been assigned a conversion flag, so that the value ranges of the main data and sub-data used in user profile estimation become approximately the same. This eliminates the influence of the original data properties, such as cyber data and physical data, on multiple data used in user profile estimation, and has the effect of allowing each data to be used for user profile estimation with equal value. Specifically, when performing user profile estimation using main data of cyber information indicating the number of app function accesses, which is at most around two digits, and sub-data of physical information indicating the number of steps taken daily, which is assumed to be at least three digits, if the sub-data before value range conversion is used as is, the estimation accuracy of the user profile estimation will decrease due to the influence of the value range of the sub-data which is excessively large compared to the reference value range of the main data. Therefore, by converting the value range of the sub-data to be approximately the same as the reference value range of the main data, the influence of the main data and sub-data can be made equal in value, and as a result, the estimation accuracy of user profile estimation can be improved.
[0205] Referring to Figure 8, in step S8, the estimation unit 34 performs user profile estimation based on the fourth data D4 input from the conversion unit 42. Here, the estimation unit 34 decides whether or not to use the fourth data D4 for user profile estimation based on the deletion flag included in the flag information for the fourth data D4. For example, the mask processing unit 33 generates second data D2 by assigning a deletion flag to each of the multiple fifth data D5 (or first data D1) that contains a singular portion in the frequency distribution information of the fifth data D5, and not assigning a deletion flag to the fifth data D5 that does not contain a singular portion in the frequency distribution information. The estimation unit 34 does not use the fourth data D4 that has the deletion flag assigned to it for user profile estimation, and uses the fourth data D4 that does not have the deletion flag assigned to it for user profile estimation.
[0206] The relationship between mask processing and user profile estimation processing will be explained. For example, suppose the first data D1 contains five data: data 1, data 2, data 3, data 4, and data 5. If the mask processing unit 33 assigns a deletion flag to data 3, but does not assign a deletion flag to data 1, data 2, data 4, and data 5, the estimation unit 34 will perform user profile estimation using data 1, data 2, data 4, and data 5, without using data 3.
[0207] Furthermore, if the mask processing unit 33 assigns a normalization flag to data 3 but does not assign a normalization flag to data 1, data 2, data 4, and data 5, the estimation unit 34 performs user profile estimation using data 3 which has not been normalized by the normalization unit 41, and data 1, data 2, data 4, and data 5 which have been normalized by the normalization unit 41.
[0208] Furthermore, if the mask processing unit 33 assigns a conversion flag to data 3 but does not assign a conversion flag to data 1, data 2, data 4, and data 5, the estimation unit 34 uses data 3, whose range has been converted to the reference range by the conversion unit 42, and data 1, data 2, data 4, and data 5, whose ranges have not been converted by the conversion unit 42, to perform user profile estimation.
[0209] When the estimation unit 34 estimates user preferences as a user profile, it should estimate user preferences based on the user's frequency of use or number of uses. For example, when estimating the popular features of an app, the estimation unit 34 should identify the top few features (e.g., the top 5) that users access most frequently, and then estimate these top few features as popular features preferred by users.
[0210] Furthermore, when the estimation unit 34 estimates which tendency a user falls into for a predetermined category, it calculates the average value of the data that falls into each category and estimates the category with the highest average value as the user's tendency. For example, there may be categories of functions related to user interaction on the app and categories of functions linked to each user's own step count, and the estimation unit 34 wants to estimate which category of functions a user's tendency falls into. In this case, the estimation unit 34 calculates the average value of the data that falls into the category of functions related to user interaction and the average value of the data that falls into the category of functions linked to step count, and compares the average values of the two. Then, if the former average value is higher, the estimation unit 34 estimates that the user has a strong tendency to prefer functions related to user interaction, and if the latter average value is higher, it estimates that the user has a strong tendency to prefer functions linked to step count.
[0211] Furthermore, when the estimation unit 34 estimates the characteristics of a particular user for a specific item, it only needs to estimate the user profile based on the position of that user's data within the overall data distribution for that item. For example, when estimating the characteristics of a user for the item "walking" included in physical activities, the estimation unit 34 compares the median of a data distribution created based on the average daily step count of all users with the average daily step count of a particular user. Then, if the average step count of that user is smaller than the median of the data distribution, the estimation unit 34 estimates that user's characteristics as "a user who doesn't walk much," and if the average step count of that user is larger than the median of the data distribution, it estimates that user's characteristics as "a user who walks a lot."
[0212] In this embodiment, although the deletion unit 44 deletes outlier data in step S4, data that still contains values that could be outliers can be excluded from the data used for user profile estimation. In other words, data that could cause a decrease in the estimation accuracy of user profile estimation can be excluded from the data used for user profile estimation, and as a result, the estimation accuracy of user profile estimation can be improved.
[0213] Next, in step S9, the matching processing unit 45 performs matching processing to match multiple users with each other, or matching processing to match users with content, etc., based on the user profile estimation results by the estimation unit 34.
[0214] For the matching process, the matching processing unit 45 can match multiple users based on their similarity by calculating, for example, Pearson's correlation coefficient or cosine similarity for multiple data points corresponding to multiple users, regardless of the content of the masking process. When using Pearson's correlation coefficient, the matching processing unit 45 should match the users with the highest correlation coefficient values. When using cosine similarity, the matching processing unit 45 should match users whose cosine similarity value is "1.0", or users whose cosine similarity value is closest to "1.0". Note that the calculation method used in the matching process is not limited to the above examples, and other calculation methods may be used.
[0215] Furthermore, the matching processing unit 45 may perform matching processing only on users whose preferences, tendencies, or characteristics are the same or similar, based on the user profile estimation results in step S8. This reduces the computational load and shortens the processing time compared to performing matching processing on all users included in the input data. On the other hand, the matching processing unit 45 may perform matching processing on all users included in the input data. In this case, the number of candidate user combinations for matching increases, which has the effect of increasing the likelihood of achieving an optimal match.
[0216] The relationship between masking and matching is explained below. For example, suppose the first data D1 contains five data: data 1, data 2, data 3, data 4, and data 5. If the masking unit 33 assigns a deletion flag to data 3, but does not assign a deletion flag to data 1, data 2, data 4, and data 5, the matching unit 45 may perform the calculation for matching using data 1, data 2, data 4, and data 5, without using data 3. By assigning a deletion flag, data that distorts the result of the matching process (for example, data with a number of digits that differs by 5 digits compared to other data) can be excluded. In addition, data unrelated to matching for a specific item can be excluded, such as when physical data is mixed in when matching from a cyber data perspective. As a result, the matching accuracy can be improved.
[0217] Furthermore, if the mask processing unit 33 assigns a normalization flag to data 3 but does not assign a normalization flag to data 1, data 2, data 4, and data 5, the matching processing unit 45 may use data 3, which has not been normalized by the normalization unit 41, and data 1, data 2, data 4, and data 5, which have been normalized by the normalization unit 41, to perform calculations for matching. By assigning a normalization flag, the characteristics of the input data can be preserved without normalizing characteristic data, which has the effect of strengthening the influence of data that is likely to express the user's characteristics and enabling matching. As a result, the matching accuracy can be improved. For example, while most of the data used in matching becomes a value between "0" and "1" through normalization, the original value (e.g., between "0" and "50") of characteristic data that has not been normalized is used in the matching process, thereby emphasizing the influence of data with a characteristic distribution.
[0218] Furthermore, if the mask processing unit 33 assigns a conversion flag to data 3 but does not assign a conversion flag to data 1, data 2, data 4, and data 5, the matching processing unit 45 may perform calculations for matching using data 3, whose range has been converted to the reference range by the conversion unit 42, and data 1, data 2, data 4, and data 5, whose ranges have not been converted by the conversion unit 42. The accuracy of the matching process improves when the ranges of the multiple data used in the matching process are the same compared to when the ranges are not the same. By assigning a conversion flag, the ranges of the multiple data used in the matching process become approximately identical, which has the effect of improving the matching accuracy.
[0219] Referring to Figure 8, in step S10, the output unit 35 outputs data indicating the result of the matching process input from the matching processing unit 45. The data format and output destination of the data indicating the result of the matching process are arbitrary. The output destination may be a terminal in the local environment or a cloud server, etc.
[0220] As described above, according to this embodiment, the processing unit 21 (information processing device) acquires first data D1 relating to the user's behavior, generates frequency distribution information relating to the first data D1, performs masking based on the frequency distribution information to generate second data D2 to which flag information is selectively assigned to the first data D1, estimates a user profile including the user's preferences, tendencies, or characteristics based on the second data D2, and outputs the estimated result of the user profile. Therefore, according to this embodiment, it is possible to improve the estimation accuracy of the user profile by estimating the user profile based on second data D2 to which flag information is selectively assigned to the first data D1 by masking.
[0221] The following describes specific examples of how this disclosure can be applied.
[0222] One possible application example is to a social networking service (SNS) app that revitalizes local communities. Users within this app belong to a specific community (for example, residents of a town). In this case, log data originating from people with diverse attributes, such as differences in age from young people to the elderly, and differences in their occupations, is expected to be collected through the app. Therefore, individual characteristics are widely distributed in the log data obtained through the app, and distinctive behaviors of certain users may be detected as outliers. In particular, data collected early on, when users have not been using the app for very long, may contain outliers such as operation logs during app installation or operation logs that are not expected to be used in the intended way because the user is unfamiliar with the app.
[0223] If outliers exist in the data used for user profile estimation and matching, the accuracy of the user profile estimation and matching process using this data will be significantly affected by these outliers.
[0224] On the other hand, the masking process described in this disclosure removes the influence of outliers even in data acquired early, enabling highly accurate user profile estimation and matching.
[0225] For regional revitalization, for example, interaction among residents within a local community is important, and services that match residents with each other are being piloted to promote such interaction. However, generally speaking, it is difficult to match residents belonging to a local community with each other due to differences in preferences, tendencies, or characteristics.
[0226] In this disclosure, log data is processed by masking, and user profile estimation is performed based on the masked log data. At this time, even if unintended data that would reduce the estimation accuracy of user profile estimation is included in the input data, data that should clearly be excluded can be removed by masking. As a result, the estimation accuracy of user profile estimation can be improved.
[0227] Furthermore, by performing matching processing based on the estimation results of user profile estimation based on masked data, it is possible to match users who have the same or similar preferences, tendencies, or characteristics. For example, according to this disclosure, it is possible to match users who have viewed or commented on articles about restaurants posted on a social networking service (SNS) app with users who have viewed the articles. Then, by communicating the matching results to users through the app and promoting interaction among users, it becomes possible to revitalize local communities.
[0228] Thus, according to this disclosure, even if outliers due to characteristic user behavior are detected in the log data acquired by the app, or even if the data is acquired early through the app, the impact of these outliers that could negatively affect user profile estimation and matching can be removed, and the data can be processed. This enables effective user profile estimation and matching of users, and as a result, can promote interaction among local residents and contribute to regional revitalization.
[0229] A specific example of the processing in the first application example will be described in order, referring to the configuration of this disclosure. Assume that a user in a certain community is using a smartphone as terminal 11. Assume that an SNS application intended to revitalize the local community is already installed on terminal 11. The acquisition unit 31 acquires the user log of the user's SNS application. A database built on the cloud can be considered as the storage unit 22, and the user log acquired from the smartphone is stored in the database.
[0230] Taking the example of a normalization process that runs at 11 PM every day, this execution process is called as a batch process at 11 PM, and the determination unit 43 determines anomalies in the data obtained from the database that are not expected from the normal use of the application. In the case of the SNS application example above, certain operation logs that are not expected from normal application operation occur frequently due to the operation during installation, and this is detected as a peak in the data distribution. The deletion unit 44 deletes the data included in the anomaly determined by the determination unit 43.
[0231] The estimation unit 411 estimates which of several predetermined data distributions the distribution of the data from which singular parts have been removed by the deletion unit 44 corresponds. In this example, the distribution is set based on the characteristics of the data source being an SNS app used by residents of a community, such as a Poisson distribution from the perspective of the number of times the app was used in discrete time and the number of times stamps representing "likes" were pressed, and a normal distribution from the perspective of the distribution of a large number of users belonging to a community.
[0232] The calculation unit 412 determines the normalization parameters necessary for normalization based on the data distribution estimated by the estimation unit 411. At this time, from the perspective of user log data, outlier data is generated due to the user's behavior on the app. For example, log data of a user who communicates exceptionally actively on the app. This data would normally be removed as an outlier, but it represents the characteristics of a user who acts actively on the app. If normalization is simply performed using the maximum value of the data, data that could be considered outliers will be rounded to "1.0", while other data will be rounded to a format that is close to "0". At this time, the calculation unit 412 calculates parameters that allow normalization in a format that emphasizes the characteristics of the user, by leaving the data of the user who acts actively on the app at a value greater than "1.0", such as "1.3".
[0233] The masking unit 33 performs masking on data that contains outliers, or data that does not identify as outliers but contains values that originate from users characteristic of the app. For example, suppose the data representing the number of stamps ("likes") contains data representing an excessive number of stamps. An excessive number of stamps is characteristic user behavior through the app. To emphasize this data, it is excluded from the normalization operation, and the original values of the input data are used for user profile estimation and matching processing.
[0234] The normalization unit 413 performs normalization processing using the normalization parameters calculated by the calculation unit 412. The transformation unit 42 processes the data so that the value ranges are consistent across multiple data points in order to improve the accuracy of user profile estimation and matching processing.
[0235] The estimation unit 34 performs user profile estimation using data to which flag information has been added by the mask processing unit 33, normalized by the normalization unit 41, and whose value ranges have been aligned by the transformation unit 42. The estimation unit 34 performs user profile estimation based, for example, on actions such as pressing stamps like "Like" within the app, or on the frequency of use of app functions.
[0236] The matching processing unit 45 matches users who are compatible (i.e., have the same or similar preferences, tendencies, or characteristics) based on the user profile estimation results from the estimation unit 34. The output unit 35 outputs the results of the matching processing by the matching processing unit 45 through the app and presents them to the user. By presenting users with other users who are compatible with them, it is possible to match users who would not have been connected if they were not using the app, thereby promoting interaction within the local area.
[0237] A second application example is the application to a service that energizes office employees. The users of this service are employees working in a specific office. Therefore, similar to the first application example, it is expected that log data derived from diverse attributes will be obtained through the app, such as (1) age range from 20s to 60s, (2) sales work mainly involving fieldwork, or (3) desk work mainly, for each user's attributes such as age or work style. Consequently, the log data obtained through the app will have a wide distribution of each user's characteristics, and distinctive behaviors of certain users may be detected as outliers. In particular, data obtained early in the process, when the app has not been used for long, may include outliers such as operation logs during app installation, or operation logs that are not expected in the intended use of the app because the user is unfamiliar with its operation. If outliers are included in the data used for user profile estimation and matching, the accuracy of user profile estimation and matching using this data will decrease because it will be greatly affected by these outliers.
[0238] In contrast, according to this disclosure, even with early-acquired data that is susceptible to outliers, the influence of outliers can be removed by masking, thereby improving the accuracy of user profile estimation and matching processes.
[0239] To revitalize office employees, interaction among employees is crucial. Furthermore, for multiple departments or teams to collaborate on a project, interaction not only between employees but also within groups of multiple employees is important. However, depending on the size of the company, an office can have thousands or even tens of thousands of employees, resulting in a vast number of possible combinations. Consequently, performing a matching process for all employees would increase computational costs and processing time enormously.
[0240] Therefore, in this disclosure, the masking processing unit 33 performs masking on log data relating to all users acquired by the acquisition unit 31, and the estimation unit 34 performs user profile estimation based on log data relating to multiple users belonging to a specific group that has been narrowed down by the masking process. Furthermore, based on the results of the user profile estimation, the estimation unit 34 estimates a representative user who is considered to be a central figure within that specific group. For example, the estimation unit 34 estimates a representative user as a user who actively presses stamps such as "like" on other people's posts on the app, or a user who frequently exchanges comments with many users within the app, within that specific group. After estimating a representative user for each of the multiple groups present in the office, the estimation unit 34 performs user profile estimation based on log data relating to the multiple representative users that have been narrowed down by the masking process. The matching processing unit 45 performs matching processing between representative users based on the estimated user profiles of the multiple representative users. In other words, this disclosure performs matching processing between representative users, targeting multiple representative users who are central figures in each group, and uses the results of the representative user matching processing as the results of the group matching processing. In this way, by performing matching on a group basis rather than individual matching of all employees, computational costs can be reduced.
[0241] The matching processing unit 45 performs group matching by identifying combinations of groups that are compatible with representative users. The output unit 35 outputs the results of the matching processing by the matching processing unit 45 through the application and presents them to the user. By promoting interaction between groups through the application, interaction between groups within the office is promoted. Thus, according to this disclosure, representative users are estimated by highly accurate user profile estimation based on mask processing, and groups are matched by matching representative users. Compared to performing individual matching processing for all employees, computational costs and processing time can be reduced.
[0242] A specific example of the processing in the second application example will be explained with reference to the configuration of this disclosure. Assume that an employee working in an office is using a smartphone as terminal 11. Assume that an SNS application intended to revitalize interpersonal relationships within the office is already installed on terminal 11. The acquisition unit 31 acquires the user log of the user's SNS application. The storage unit 22 is thought to be a database built on the cloud, and the user log acquired from the smartphone is stored in the database.
[0243] Taking the example of a normalization process that runs at 11 PM every day, this execution process is called as a batch process at 11 PM, and the determination unit 43 determines anomalies in the data obtained from the database that are not expected from the normal use of the application. In the case of the SNS application example above, certain operation logs that are not expected from normal application operation occur frequently due to the operation during installation, and this is detected as a peak in the data distribution. The deletion unit 44 deletes the data included in the anomaly determined by the determination unit 43.
[0244] The estimation unit 411 estimates which of several predetermined data distributions the distribution of the data from which singular parts have been removed by the deletion unit 44 corresponds. In this example, the distribution is set based on the characteristics of the data source being log data from office employees, such as a Poisson distribution from the perspective of the number of times the app was used in discrete time and the number of times stamps representing "likes" were pressed, and a normal distribution from the perspective of the distribution of a large number of users working in an office.
[0245] The calculation unit 412 determines the normalization parameters necessary for normalization based on the data distribution estimated by the estimation unit 411. At this time, from the perspective of user log data, outlier data is generated due to the user's behavior on the app. For example, log data of a user who communicates exceptionally actively on the app. This data would normally be removed as an outlier, but it represents the characteristics of a user who acts actively on the app. If normalization is simply performed using the maximum value of the data, data that could be considered outliers will be rounded to "1.0", while other data will be rounded to a format that is close to "0". At this time, the calculation unit 412 calculates parameters that allow normalization in a format that emphasizes the characteristics of the user, by leaving the data of the user who acts actively on the app at a value greater than "1.0", such as "1.3".
[0246] The masking unit 33 assigns a normalization flag to data that contains values originating from characteristic users within the app by performing a masking process. The normalization unit 413 does not normalize data that has been assigned a normalization flag, but normalizes data that does not have a normalization flag. For example, suppose data regarding the number of stamps representing "likes" contains data representing an excessive number of stamps. An excessive number of stamps is a characteristic behavior of users through the app. To emphasize this data, the masking process excludes this data from the normalization calculation, and the original values of the input data are used for user profile estimation and matching.
[0247] The estimation unit 34 performs user profile estimation using data that was not normalized by the normalization unit 41 because a normalization flag was assigned by the mask processing unit 33, and data that was normalized by the normalization unit 41 because a normalization flag was not assigned by the mask processing unit 33. Furthermore, based on the results of the user profile estimation, the estimation unit 34 estimates representative users who are considered to be central figures within a particular group. After estimating representative users for each of the multiple groups present in the office, the estimation unit 34 performs user profile estimation based on log data related to multiple representative users that have been narrowed down by the masking process. The matching processing unit 45 performs matching processing between representative users based on the estimated user profiles of the multiple representative users.
[0248] The matching processing unit 45 performs group matching by identifying compatible group combinations between representative users. The output unit 35 outputs the results of the matching processing by the matching processing unit 45 to the user via the application. By encouraging interaction between groups through the application, communication between groups within the office is promoted. [Industrial applicability]
[0249] This disclosure is broadly applicable to user profile estimation or matching processes based on user behavior data. [Explanation of Symbols]
[0250] 1. Information Processing System 12 Estimation device Processing Units 21, 21A~21D 31 Acquisition Department 32 Generation part 33 Mask Processing 34 Estimation part 35 Output section 41 Normalization section 42 Conversion section 43 Judgment section 44 Deleted section 45 Matching Processing Unit
Claims
1. Information processing device, We obtain the first data regarding user behavior, Frequency distribution information is generated for the first data, By performing a masking process based on the frequency distribution information, a second data set is generated in which flag information is selectively assigned to the first data set. Based on the second data mentioned above, a user profile including the user's preferences, tendencies, or characteristics is estimated. Output the estimated results of the user profile. Information processing methods.
2. In the estimation of the user profile, it is determined whether or not to use the second data for the estimation of the user profile based on the flag information. The information processing method according to claim 1.
3. In generating the second data, the flag information is assigned to the first data which includes a singular portion in the frequency distribution information. In the estimation of the user profile, the second data to which the flag information is assigned is not used for estimating the user profile. The information processing method according to claim 2.
4. In generating the second data, a predetermined unique frequency distribution shape among the frequency distribution information is determined to be the unique portion. The information processing method according to claim 3.
5. Furthermore, by normalizing the second data, a third data is generated. In estimating the user profile, the user profile is estimated based on the third data. In generating the third data, whether or not to normalize the second data is determined based on the flag information. The information processing method according to claim 1.
6. In generating the second data, the flag information is added to the first data which includes a feature portion representing the user's characteristics. In generating the third data, the second data to which the flag information is attached is not normalized. The information processing method according to claim 5.
7. In generating the third data, The data distribution of the second data is estimated from multiple data distributions. Based on the data distribution of the second data mentioned above, the parameters for normalization are calculated, The second data is normalized using the aforementioned parameters. The information processing method according to claim 5.
8. Furthermore, by transforming the range of the second data, a fourth data is generated. In estimating the user profile, the user profile is estimated based on the fourth data. In generating the fourth data, whether or not to transform the range of the second data is determined based on the flag information. The information processing method according to claim 1.
9. In generating the second data, the flag information is assigned to the first data whose data range differs from the reference range. In generating the fourth data, the range of values of the second data to which the flag information is attached is converted to the reference range. The information processing method according to claim 8.
10. Furthermore, based on the estimated user profile results, a matching process is performed between multiple users. The information processing method according to any one of claims 1 to 9.
11. moreover, Based on the estimated user profile results, the representative user within a group to which multiple users belong is estimated. Based on the estimated user profile results for the aforementioned representative user, a matching process is performed between multiple representative users. The information processing method according to any one of claims 1 to 9.
12. Equipped with a circuit configuration, The aforementioned circuit configuration is, We obtain the first data regarding user behavior, Frequency distribution information is generated for the first data, By performing a masking process based on the frequency distribution information, a second data set is generated in which flag information is selectively assigned to the first data set. Based on the second data mentioned above, a user profile including the user's preferences, tendencies, or characteristics is estimated. Output the estimated results of the user profile. Information processing device.
13. A program that causes an information processing device to perform processing, The aforementioned process is, We obtain the first data regarding user behavior, Frequency distribution information is generated for the first data, By performing a masking process based on the frequency distribution information, a second data set is generated in which flag information is selectively assigned to the first data set. Based on the second data mentioned above, a user profile including the user's preferences, tendencies, or characteristics is estimated. Output the estimated results of the user profile. program.
Citation Information
Patent Citations
Inspection result data output system
JP2003067489A
Person matching device, method and program
JP2012078768A