Data processing apparatus, data processing method, and program

The data processing apparatus generates pseudo-sample data to analyze contact status with target content across multiple media, addressing the lack of contact frequency data in existing methods and enhancing analysis accuracy.

JP2025088678AActive Publication Date: 2025-06-11K K VIDEO RES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024024830
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-02-21
Publication Date
2025-06-11
Estimated Expiration
2044-02-21

AI Technical Summary

Technical Problem

Existing methods for analyzing contact status with content across multiple media lack data on contact frequency, which is essential for accurate analysis.

Method used

A data processing apparatus and method that acquire single-source data for users across multiple media, generate pseudo-sample data to maintain correlation coefficients, and calculate contact frequencies with target content using the pseudo-samples.

Benefits of technology

Enables the acquisition of pseudo-sample data useful for analyzing contact status with target content across multiple media, providing insights into contact frequencies and improving analysis accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025088678000001_ABST
    Figure 2025088678000001_ABST
Patent Text Reader

Abstract

To acquire pseudo-sample data useful for analyzing contacts with target content via a plurality of media.SOLUTION: A data processing apparatus includes: a real data acquisition unit which acquires single source data for a plurality of users including a first value indicating how a single user uses a first medium and a second value indicating how the single user uses a second medium; a pseudo-data generation unit which generates pseudo-samples of the single source data so that a correlation coefficient between the first and second values may not be different from the single source data for the multiple users; and a contact frequency allocation unit which calculates a first contact frequency of contacts with target content via the first medium, regarding each of the generated pseudo-samples. The contact frequency allocation unit uses data indicating contacts with the target content in the first medium, and calculates a first contact frequency based on the first value in each of the pseudo-samples.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data processing apparatus, a data processing method, and a program for analyzing the contact status with content using data amplified with a pseudo sample.

Background Art

[0002] In recent years, investigations have been conducted on the scale of the number of contacts on multiple media such as TV commercials and video site advertisements for a certain product's advertisement.

[0003] For example, Patent Document 1 discloses a method for calculating the number of contacts with at least one of a TV commercial and a digital advertisement using data on the number of contacts with a TV commercial, data on the number of contacts with a digital advertisement, and data (single-source data) indicating whether or not each of a plurality of subjects watched the TV commercial and the number of views of the site where the digital advertisement was posted.

[0004] Also, for example, as described in Patent Document 2, in an investigation of the contact status with content, it is known to amplify the number of data using pseudo sample data created based on actual sample data.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] In general, single-source data includes information indicating the contact status of a single individual with multiple media, but does not necessarily include information indicating the frequency of contact with a target advertisement via multiple media. However, in order to analyze the contact status with a target advertisement via multiple media, data including information indicating the contact status in accordance with the actual situation in each medium has been required.

[0007] An object of the present invention is to enable acquisition of pseudo-sample data useful for analyzing the contact status with target content via multiple media.

Means for Solving the Problems

[0008] A data processing apparatus according to the present invention includes an actual data acquisition unit that acquires single-source data for a plurality of users including a first value indicating the usage status of a first medium of a single user and a second value indicating the usage status of a second medium, and a pseudo-data generation unit that generates pseudo-samples so that the correlation coefficient between the first value and the second value remains unchanged with respect to the single-source data for the plurality of users, and a contact frequency allocation unit that calculates a first contact frequency of contacting target content via the first medium for each generated pseudo-sample. The contact frequency allocation unit uses data indicating the contact status with the target content in the first medium and calculates the first contact frequency based on the first value in each pseudo-sample.

[0009] The data processing method according to the present invention includes a step in which a processor acquires single-source data for a plurality of users including a first value indicating the usage status of a first medium of a single user and a second value indicating the usage status of a second medium, and a step in which the processor generates a pseudo-sample of the single-source data so that the correlation coefficient between the first value and the second value remains unchanged from the single-source data for the plurality of users, and a step in which the processor calculates a first contact frequency of contacting target content via the first medium for each generated pseudo-sample. In the step of calculating the first contact frequency, data indicating the contact status with the target content in the first medium is used, and the first contact frequency is calculated based on the first value in each pseudo-sample.

[0010] The program according to the present invention causes a computer to function as an actual data acquisition unit that acquires single-source data for a plurality of users including a first value indicating the usage status of a first medium of a single user and a second value indicating the usage status of a second medium, a pseudo-data generation unit that generates a pseudo-sample of the single-source data so that the correlation coefficient between the first value and the second value remains unchanged from the single-source data for the plurality of users, and a contact frequency allocation unit that calculates a first contact frequency of contacting target content via the first medium for each generated pseudo-sample. The contact frequency allocation unit uses data indicating the contact status with the target content in the first medium and calculates the first contact frequency based on the first value in each pseudo-sample.

Advantages of the Invention

[0011] According to the present invention, it is possible to acquire pseudo-sample data useful for analyzing the contact status with target content via a plurality of media.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0013] Next, embodiments for carrying out the present invention will be described in detail with reference to the drawings. (Embodiment 1) FIG. 1 is a block diagram showing the configuration of a data processing apparatus 1 according to Embodiment 1 of the present invention. The data processing apparatus 1 is composed of one or a plurality of computers connected by a communication line. The data processing apparatus 1 includes a processor 11, a main memory 12, an input / output interface 13, a communication interface 14, and a storage device 15. The storage device 15 is a computer-readable recording medium such as a semiconductor memory (e.g., a volatile memory or a non-volatile memory) or a disk medium (e.g., a magnetic recording medium or a magneto-optical recording medium). Programs to be executed by the processor 11 and various data are stored in the storage device 15. The programs are read from the storage device 15 into the main memory 12 and interpreted and executed by the processor 11, whereby various functions are executed.

[0014] FIG. 2 is a block diagram showing the functional modules of a program executed by the processor 11 of the data processing apparatus 1. As shown in FIG. 2, the functional modules executed by the processor 11 of the data processing apparatus 1 include an actual data acquisition unit 101, a pseudo data generation unit 102, a contact frequency assignment unit 103, and a totaling unit 104.

[0015] The storage device 15 stores measured single-source data (actual data) and pseudo-sample data generated based on the actual data. The single-source data is data including the results of measuring the contact status of a single user (the same individual) with a plurality of media. In this embodiment, as an example, data including the measurement results of the TV usage time, the number of times of contact with advertisements on TV, and the usage time of a video site (such as YouTube (registered trademark)) of the same individual is used as the single-source data.

[0016] Next, using the flowchart of FIG. 3, the data processing flow by the data processing apparatus 1 will be described. The data processing apparatus 1 generates data for analyzing the contact situation with advertisement C on TV (the second medium) and the contact situation with advertisement C on the web video site (the first medium) for a certain advertisement C (target content). Here, an example of analyzing the contact situation with the target advertisement in multiple media is given, but the target content for analyzing the contact situation is not limited to advertisements, and may be, for example, a specific program or video.

[0017] First, the measured data acquisition unit 101 acquires single-source data (actual data) regarding the usage history of TV (the second medium) and the usage history of the web video site (the first medium) (step S101). FIG. 4(A) is a diagram showing a specific example of the single-source data. As shown in FIG. 4(A), the single-source data includes the TV usage time (minutes) (the second value indicating the usage situation) of each surveyed user (Sno001, 002,...) during a predetermined survey period (for example, one week), the number of times of contacting advertisement C on TV (times), and the usage time (minutes) of the video site (such as Youtube) (the first value indicating the usage situation). Note that the single-source data does not include the number of times of contacting advertisement C on the video site.

[0018] Also, the single-source data may include user attribute information (such as gender and age). In the example of FIG. 4(A), it includes gender and age classification as attribute information, and as shown in the figure, single-source data regarding users of male (M18 - 24) aged 18 to 24 has been acquired.

[0019] Next, the pseudo-data generation unit 102 generates pseudo-sample data having data items similar to the acquired single-source data and having a similar distribution (step S102). FIG. 4(B) is a diagram illustrating pseudo-sample data generated based on the single-source data of FIG. 4(A). The pseudo-data generation unit 102 obtains a three-dimensional normal distribution for the three items (TV usage time, number of TV ad exposures, video site usage time) that make up the data based on the single-source data acquired in step S101. Further, pseudo-sample data is randomly generated according to the obtained three-dimensional normal distribution. The pseudo-data generation unit 102 generates pseudo-sample data such that the average of each item (TV usage time, number of TV ad exposures, video site usage time) and the correlation coefficient between the items in the generated pseudo-sample data are the same as the average and correlation coefficient in the original single-source data. In the example of FIG. 4(B), since a normal distribution random number is assigned to the numerical value of each item of the pseudo-sample, for example, the number of TV ad exposures also has a numerical value including a decimal point rather than a natural number.

[0020] Also, the number of pseudo-samples generated by the pseudo-data generation unit 102 can be set according to the purpose of the survey. In the example of FIG. 4(B), the number of pseudo-samples is determined based on the statistical data of the gender / age composition of the TV owner population illustrated in FIG. 5. FIG. 5 shows the TV owner population in each gender / age category when the number of pseudo-samples is 100,000. MF, M, and F represent men and women, men, and women, respectively, and the numbers on the horizontal axis represent age groups. FIG. 4(B) is a pseudo-sample generated based on the actual data of men aged 18 to 24 (M18-24). According to FIG. 5, when the total TV owner population is 100,000, the number of men aged 18 to 24 is 3,346. Therefore, in the example of FIG. 4(B), 3,346 pseudo-samples are generated. Here, data assuming the TV owner population for each gender / age category is used, but not only the TV owner population, but also, for example, the entire population for each gender / age category can be assumed.

[0021] Next, for each of the generated pseudo-samples, the contact frequency allocation unit 103 calculates the number of times of contact with advertisement C via the video site (the first contact frequency) (step S103).

[0022] Using FIG. 6, the method for calculating the number of times of contact with advertisement C on the video site by the contact frequency allocation unit 103 will be described. To calculate the number of times of contact, the distribution data of the number of times of contact with advertisement C on the video site provided as official data is used. In the second column of the table in FIG. 6(A), the distribution (official data) of the number of times of contact with advertisement C (from 0 times to 10 times or more) in a predetermined population is illustrated. In the third column, the number of samples (number of data) obtained by allocating the pseudo-samples generated in step S102 (the data for 3346 people in the example of FIG. 4(B)) to each number of times of contact (from 0 times to 10 times or more) according to the distribution in the second column is shown. Further, in the fourth column, the result of rounding off the decimal part of the numerical value in the third column and adjusting the number of people with 10 or more times of contact so that the total becomes 3346 people is shown.

[0023] FIG. 6(B) is a diagram showing an example in which the rank of the number of times of contact with the TV commercial (the third column of the table) and the rank of the usage time of the video site (the sixth column of the table) are assigned to each of the pseudo-samples generated in step S102. The rank of the number of times of contact with the TV commercial (the third column of the table) is assigned in ascending order of the number of times of contact with advertisement C on the TV in the fourth column of the table. On the other hand, the rank of the usage time of the video site (the sixth column of the table) is assigned in ascending order of the usage time of the video site.

[0024] Based on the distribution of the number of contacts with advertisement C on the video site shown in Fig. 6(A), the contact frequency assignment unit 103 calculates the number of contacts with advertisement C on the video site for each pseudo sample in Fig. 6(B). Referring to the fourth column in Fig. 6(A), out of the 3346 pseudo samples, for 1255 of them, the number of contacts with advertisement C on the video site is "0" times. Therefore, the contact frequency assignment unit 103 sets the number of contacts with advertisement C to "0" times for the pseudo samples up to the 1255th one in ascending order of the usage time of the video site among the pseudo samples in Fig. 6(B). Similarly, for the samples from the 1256th to the 1690th, the number of contacts with advertisement C is set to "1" time, from the 1691st to the 2008th is "2" times, from the 2009th to the 2319th is "3" times, and from the 2320th to the 2677th is "4" times. In the example of Fig. 6(B), since the samples Sno001,002 are included up to the 1255th one, the number of contacts with advertisement C is 0 times. On the other hand, since the sample Sno003 is included in the range from the 2320th to the 2677th, the number of contacts with advertisement C is 4 times. In the above way, the number of contacts with advertisement C on the video site in the pseudo sample data can be set.

[0025] Also, although the value of the number of TV commercial exposures is already included in the pseudo sample, it may be re-set based on the ranking of the number of TV commercial exposures. Specifically, similar to the number of exposures to Ad C on the video site, use the distribution data of the number of exposures to Ad C on TV provided as official data (the second column in Figure 7), allocate the data for 3346 people to each number of exposures (for example, 0 times to 10 times or more) (the third column in Figure 7), obtain the number of allocated data for each number of exposures (the fourth column in Figure 7), and allocate the number of exposures to Ad C on TV according to the ranking in the third column of Figure 6(B). As a result, for TV commercials as well, a pseudo sample with an exposure frequency distribution that matches the distribution of official data can be created. For example, in the example of Figure 6(B), for Sno001, the original number of TV commercial exposures shown in the pseudo sample is 5.6 times, but since the rank of the TV commercial is 2253rd, according to the distribution in Figure 7, the number of exposures becomes 2 times. Also, for Sno002, the original number of TV commercial exposures shown in the pseudo sample is 3.3 times, but since the rank of the TV commercial is 1521st, according to the distribution in Figure 7, the number of exposures becomes 0 times.

[0026] Through the procedures of the above steps S101 to S103, a pseudo sample of the desired number of cases including the number of exposures to Ad C on TV and the number of exposures to Ad C on the video site can be obtained from the limited number of single-source data (actual data) including TV usage time, the number of exposures to Ad C on TV, and video site usage time.

[0027] (Analysis of Integrated Reach and Duplicate Reach) The aggregating unit 104 estimates integrated reach and overlapping reach using the generated pseudo samples. The integrated reach is the ratio at which at least one of a plurality of events occurs. In the above embodiment, it indicates the ratio of users who have come into contact with at least one of the TV advertisement and the video site advertisement. The overlapping reach is the ratio at which all of a plurality of events occur. In the above embodiment, it indicates the ratio of users who have come into contact with both the TV advertisement and the video site advertisement. That is, in the above embodiment, the integrated reach and the overlapping reach can be calculated, for example, by the following formulas (1) and (2). In the following formulas (1) and (2), the integrated reach and the overlapping reach are calculated on the premise that a user who has come into contact even once is considered to have been reached. The definition of reach is not limited to this. For example, when it is determined that a user has been reached when they have come into contact two or more times, three or more times, etc., in the following formula, it can be calculated by replacing it with "number of contacts ≥ 2", "number of contacts ≥ 3".

[0028] Integrated reach = ([Number of users with ≥ 1 contact with TV advertisement] + [Number of users with ≥ 1 contact with video site advertisement] - [Number of users with ≥ 1 contact with both TV advertisement and video site advertisement]) / 3346 …(1) Overlapping reach = [Number of users with ≥ 1 contact with both TV advertisement and video site advertisement] / 3346 …(2)

[0029] By obtaining the integrated reach using the generated pseudo samples, the relationship between the contact rates with the TV advertisement and the video site advertisement respectively and the integrated reach can be analyzed and utilized for efficient advertisement deployment.

[0030] In the above embodiment, a single-source pseudo sample including the number of contacts with the TV advertisement and the video site advertisement is obtained. However, the items included in the pseudo sample can be adjusted according to the analysis purpose. For example, for advertisement C on the video site, it may be possible to distinguish between the case of contact on the TV screen and the case of contact on the smartphone. Also, for contact with advertisement C on TV, the number of contacts by station may be included. Also, the number of contacts in a specific time zone or at a specific site can be calculated in the same procedure.

[0031] As described above, according to this embodiment, using single-source data including the usage times of a plurality of media, pseudo samples are generated so that the correlation coefficients between items do not change. Further, using the distribution data of the number of contacts with the target advertisement C in each medium, the number of contacts with the advertisement C is assigned based on the usage time of the medium in the pseudo samples. As a result, it is possible to estimate the actual number of contacts using single-source data that contains only information on the usage time of the media. Thus, it is possible to generate pseudo sample data that can be used for analyzing the contact situation with advertisement C via a plurality of media. Also, it can be expected that even when analysis or the like is performed using the created pseudo samples, results that do not conflict with the results obtained when analyzing using measured data can be obtained.

[0032] In this embodiment, pseudo sample data showing the contact situation with TV commercials and video site advertisements is created, but the number and types of media are not limited to this, and it can be used to create pseudo samples regarding the contact situation with a plurality of media such as newspapers and radio in addition to TV and the web. Also, in addition to integrated reach and duplicate reach, various indicators and statistical data that can be analyzed and calculated based on single-source data can be created. Also, it is not limited to the integrated reach and duplicate reach of two types of media, and it can support the integrated reach and duplicate reach of any number of media and other analyses.

[0033] Also, the created pseudo sample data can be used not only for the analysis of integrated reach and duplicate reach, but also for, for example, the following uses. (1) Used for depicting the attribute profile of advertisement contacts. (2) By fusing with other data sources, it can be used for various other purposes. Specifically, the following examples can be given. (2)-1: Fused with the data of advertisement distributors to achieve effective distribution for complementing reach. (2)-2: Fused with brand evaluation data and used for analyzing the advertising effect on brand evaluation. (2)-3: Integrate with purchase history data and use it for analyzing the advertising effect on purchases. (2)-4: Integrate with the attribute profile data of consumers and use it for obtaining a detailed profile of the advertising contacts.

[0034] (Embodiment 2) The configuration of the data processing apparatus 1 according to Embodiment 2 of the present invention and the functional modules of the program executed by the processor 11 of the data processing apparatus 1 are the same as those in Embodiment 1 shown in FIGS. 1 and 2. Also, the data processing flow by the data processing apparatus 1 is the same as the flow shown in the flowchart of FIG. 3. That is, based on the single-source data exemplified in FIG. 4(A), pseudo-sample data exemplified in FIG. 4(B) is generated in the same manner as in Embodiment 1. Further, the contact frequency assignment unit 103 calculates the number of times of contact with the advertisement C (the first contact frequency) for each generated pseudo-sample via the video site. In Embodiment 2, the number of times of contact with the advertisement C via the video site is calculated by a method different from that in Embodiment 1.

[0035] In Embodiment 1, as official data, distribution data of the number of contacts with the advertisement C on the video site as shown in FIG. 6(A) is provided, and using this, the number of times of contact with the advertisement C via the video site in each pseudo-sample is calculated. On the other hand, on many video sites, the distribution data of the number of contacts with the advertisement C as described above is not provided. Instead, data indicating the ratio of the presence or absence of contact with the advertisement C on the video site may be provided. Specifically, in a predetermined population (for example, men aged 18 to 24 (M18-24)), a value defined as follows is provided. Ratio of having contact = Number of contacts with advertisement C on the video site / Number of people in the population Ratio of having no contact = 1 - (Ratio of having contact)

[0036] Also, in some cases, the average number of contacts in the group having contact with the advertisement C may be provided. Specifically, a value defined as follows is provided. Average number of contacts = Total number of displays of Ad C on the video site / Number of viewers who have been exposed to Ad C on the video site

[0037] In Embodiment 2, the number of times of contact with Ad C via the video site in each pseudo-sample is calculated using the data indicating the ratio of presence or absence of contact with Ad C on the video site and the average number of contacts in the group with contact.

[0038] First, the contact frequency assignment unit 103 assigns to each pseudo-sample whether or not there is contact with Ad C on the video site. The second column of the table in Fig. 8(A) is the data obtained as official data, and the ratio of presence or absence of contact with Ad C on the video site in a predetermined population (for example, men aged 18 to 24 (M18-24)) is illustrated. The third column shows the number of people who assigned the pseudo-samples (here, for 17,964 people) to no contact / contact according to the ratio in the second column. The fourth column shows the result of rounding the decimal part of the numerical value in the third column and adjusting the number of people with no contact so that the total becomes 17,964 people.

[0039] Fig. 8(B) is a diagram showing an example in which the pseudo-samples are given the rank of the usage time of the video site (the eighth column of the table). The rank of the usage time of the video site is given in ascending order of the usage time of the video site (the seventh column of the table). The contact frequency assignment unit 103 assigns whether or not there is contact with Ad C on the video site to each pseudo-sample in Fig. 8(B) based on the ratio of presence or absence of contact with Ad C on the video site shown in Fig. 8(A). Referring to the fourth column of Fig. 8(A), out of the 17,964 pseudo-samples, 15,719 people have no contact with Ad C on the video site. Therefore, the contact frequency assignment unit 103 assigns "no contact" with Ad C to the pseudo-samples up to the 15,719th in ascending order of the usage time of the video site among the pseudo-samples in Fig. 8(B). Similarly, for the samples from the 15,720th to the 17,964th, "contact" with Ad C is assigned.

[0040] Next, the contact frequency allocation unit 103 allocates the expected value of the number of ad contacts to the samples with "yes" contact with Ad C. The contact frequency allocation unit 103 allocates the expected value based on the relationship that satisfies the following three conditions. Condition 1: The expected value is proportional to the usage time of the video site. Condition 2: The average of the expected values matches the average number of contacts in the population with "yes" contact in the official data. Condition 3: Among the pseudo-samples assigned with "yes" contact, the expected value of the sample with the shortest usage time of the video site is "1".

[0041] The procedure for obtaining the expected value based on the relationship that satisfies Conditions 1 to 3 will be specifically described. First, the contact frequency allocation unit 103 obtains the equation Y = c + bX of the straight line (Condition 1) passing through the following two points in the plane defined by (X, Y) = (usage time, expected value of the number of contacts) as shown in FIG. 9. Point P1 (Condition 3): (the minimum value of the usage time in the samples with "yes" contact, 1) Point P2 (Condition 2): (the average usage time At calculated from the samples with "yes" contact, the average expected value Ar (where the average expected value Ar = "average number of ad contacts" in the official data))

[0042] Substitute the usage time (X) of each sample into the obtained equation (1) of the straight line to obtain the expected value Y of the number of ad contacts for each sample. Expected value of the number of ad contacts (Y) = c + b × usage time of the video site (X)... (1) (c, b are constants)

[0043] Furthermore, the contact frequency allocation unit 103 calculates the number of advertisement contacts for each sample by using the obtained expected value of each sample. For example, the contact frequency allocation unit 103 may generate one random number that follows a truncated Poisson distribution where the expected value matches the expected value of each sample, and use the generated random number as the number of advertisement contacts for the sample. Since the number of advertisement contacts is an integer of 1 or more, a truncated Poisson distribution with a domain of 1 or more may be used. When the expected value (λ) of the Poisson distribution before truncation is required to generate a random number of the truncated Poisson distribution, λ may be calculated individually according to the range of the expected value of each sample. The expected value E of the truncated Poisson distribution truncated at 1 or more and the expected value λ of the Poisson distribution before truncation have the following relationship. E = λ / (1 - exp(-λ))

[0044] According to the second embodiment, even when data on the distribution of the number of advertisement contacts on a video site cannot be obtained, if data on the ratio of the presence or absence of advertisement contacts and the average number of advertisement contacts can be obtained, the number of advertisement contacts that conforms to the actual state of the pseudo-samples can be estimated. As a result, similar to the first embodiment, pseudo-sample data that can be used for analyzing the contact situation with advertisement C via a plurality of media can be generated. Also, it can be expected that the results obtained by analyzing using the created pseudo-samples will not conflict with the results obtained by analyzing using actual measurement data.

[0045] Note that the probability distribution used to generate the number of advertisement contacts from the expected value is not limited to the truncated Poisson distribution. For example, a binomial distribution, a negative binomial distribution, a geometric distribution, a beta-binomial distribution, etc. can also be used. Also, similar to the first embodiment, the number of contacts with TV commercials may be re-set based on the ranking of the number of contacts with TV commercials.

[0046] Note that the present invention is not limited to the above-described embodiments, and can be implemented in various other forms without departing from the gist of the present invention. Therefore, the above embodiments are merely illustrative in all respects and should not be construed in a limiting sense. For example, the above-described processing steps can be arbitrarily changed in order or executed in parallel within a range that does not cause a contradiction in the processing content. Also, other steps can be added between the respective processing steps. Further, a step described as one step may be divided into multiple steps for execution, or multiple steps described separately may be grasped as one step.

Explanation of Reference Numerals

[0047] 1…Data processing device 11…Processor 12…Main memory 13…Input / output interface 14…Communication interface 15…Storage device 101…Actual data acquisition unit 102…Pseudo data generation unit 103…Contact frequency assignment unit 104…Aggregation unit

Claims

1. an actual data acquisition unit that acquires single source data for a plurality of users, the single user including a first value indicating a usage status of a first medium and a second value indicating a usage status of a second medium; a pseudo data generating unit configured to generate a pseudo sample of the single-source data such that a correlation coefficient between the first value and the second value is the same as that of the single-source data for the plurality of users; a contact frequency allocation unit that calculates a first contact frequency of the target content via the first medium for each of the generated pseudo samples, The contact frequency allocation unit A data processing device that uses data indicating a contact state with the target content in the first medium and calculates the first contact frequency based on the first value in each pseudo sample.

2. the data indicating the contact status with the target content is contact frequency distribution data, The contact frequency allocation unit The data processing apparatus according to claim 1 , further comprising: ranking each pseudo sample according to a length of time spent using the first medium; and allocating the first frequency of exposure to the target content based on distribution data of the frequency of exposure to the target content.

3. The data indicating the contact status with the target content is data indicating a ratio of contact presence / absence, The contact frequency allocation unit 2. The data processing device of claim 1, further comprising: ranking each pseudo sample according to the length of time the pseudo sample has spent using the first medium; assigning each pseudo sample a status of contact with the target content based on data indicating the ratio of contact with the target content to a status of contact; and assigning the first contact frequency to each pseudo sample that has been assigned a status of contact with the target content based on the length of time the pseudo sample has spent using the first medium.

4. The contact frequency allocation unit The data processing device according to claim 3 , wherein a random number according to a probability distribution having an expected value proportional to a length of time of using the first medium is assigned as the first contact frequency for a pseudo sample that is assigned a contact with the target content.

5. the single-source data includes a second frequency of exposure to the target content via the second medium; The contact frequency allocation unit The data processing device according to claim 1 or 3, wherein each pseudo sample is ranked according to the second contact frequency, and the second contact frequency is reallocated based on data indicating a situation regarding the target content in the second medium.

6. a processor obtaining single source data for a plurality of users, the single user including a first value indicative of a usage of a first medium and a second value indicative of a usage of a second medium; a processor generating a pseudo-sample of the single-source data such that a correlation coefficient between the first values ​​and the second values ​​is invariant to single-source data for the plurality of users; and calculating, by the processor, a first frequency of exposure to target content via the first medium for each of the generated pseudo samples; In the step of calculating the first contact frequency, A data processing method, comprising: utilizing data indicating an exposure state to the target content in the first medium; and calculating the first exposure frequency based on the first value in each pseudo sample.

7. Computer, an actual data acquisition unit that acquires single source data for a plurality of users, the single user including a first value indicating a usage status of a first medium and a second value indicating a usage status of a second medium; a pseudo data generating unit that generates a pseudo sample of the single-source data such that a correlation coefficient between the first value and the second value is the same as that of the single-source data for the plurality of users; a contact frequency allocation unit that calculates a first contact frequency of the target content via the first medium for each of the generated pseudo samples; The contact frequency allocation unit a program for calculating the first exposure frequency based on the first value for each pseudo sample by using data indicating an exposure state to the target content in the first medium;

Citation Information

Patent Citations

  • Data processing device, and data processing method

    JP2020160657A

  • Dummy sample making device, method for making dummy sample, and program

    JP2022028370A