Data processing device, data processing method, and program

The data processing device creates pseudo-sample data with matched average viewership ratings and correlation coefficients to address the challenge of predicting future broadcast frame suitability for advertising, improving advertising placement strategies.

JP2026069919APending Publication Date: 2026-04-27K K VIDEO RES
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
K K VIDEO RES
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing methods lack the ability to determine appropriate broadcast frames for advertising using pseudo samples, particularly in predicting audience ratings for future broadcast frames.

Method used

A data processing device and method that creates pseudo-sample data by matching average viewership ratings and correlation coefficients with past broadcast frames, ensuring reliability in predicting future viewership patterns.

Benefits of technology

Enables the creation of highly reliable pseudo-sample data for optimizing broadcast frame selection, enhancing the accuracy of advertising placement strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069919000001_ABST
    Figure 2026069919000001_ABST
Patent Text Reader

Abstract

Obtain reliable pseudo-sample data, including contact status with future broadcast slots. [Solution] The system comprises a predicted contact rate acquisition unit that acquires predicted contact rates for multiple future broadcast slots, a personal sample acquisition unit that acquires personal sample data for multiple individuals that includes contact status for multiple past broadcast slots corresponding to multiple future broadcast slots, and a pseudo-sample creation unit that creates pseudo-sample data for multiple individuals that includes contact status for multiple future broadcast slots based on the personal sample data for multiple individuals. The pseudo-sample creation unit creates pseudo-sample data such that the average contact rate for each future broadcast slot calculated from the created pseudo-sample data for multiple individuals matches the predicted contact rate, and the correlation coefficient between each future broadcast slot calculated from the pseudo-sample data matches the correlation coefficient between each corresponding past broadcast slot calculated from the personal sample data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data processing apparatus, a data processing method, and a program for optimizing the reach rate to content using pseudo samples.

Background Art

[0002] In order to increase the reach rate to television commercials, a method for determining which broadcast frame is appropriate for broadcasting commercials has been proposed.

[0003] For example, Patent Document 1 describes a technique for presenting, as candidates for commercial broadcast frames, program frames in which the predicted audience rating predicted by a predetermined algorithm satisfies a predetermined condition among a plurality of program frames.

[0004] Also, for example, as described in Patent Document 2, it is known to amplify the number of data using pseudo sample data created based on actual sample data in an investigation of the contact situation with content.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] Conventionally, since a method for creating pseudo samples assuming future broadcast frames has not been proposed, it has not been possible to determine broadcast frames for which advertisements or the like should be published using pseudo samples including the audience rating of future broadcast frames.

[0007] The present invention aims to obtain highly reliable pseudo-sample data, including contact status with future broadcast slots. [Means for solving the problem]

[0008] The data processing device according to the present invention comprises: a predicted contact rate acquisition unit that acquires predicted contact rates for multiple future broadcast slots; a personal sample acquisition unit that acquires personal sample data for multiple individuals, including contact status for multiple past broadcast slots corresponding to the multiple future broadcast slots; and a pseudo-sample creation unit that creates pseudo-sample data for multiple individuals, including contact status for the multiple future broadcast slots, based on the personal sample data for multiple individuals. The pseudo-sample creation unit creates the pseudo-sample data such that the average contact rate for each future broadcast slot calculated from the created pseudo-sample data for multiple individuals matches the predicted contact rate, and the correlation coefficient between each future broadcast slot calculated from the pseudo-sample data matches the correlation coefficient between each corresponding past broadcast slot calculated from the personal sample data.

[0009] The data processing method according to the present invention includes the steps of: a computer obtaining a predicted contact rate for multiple future broadcast slots; obtaining personal sample data for multiple individuals, including their contact status for multiple past broadcast slots corresponding to the multiple future broadcast slots; and creating pseudo-sample data for multiple individuals, including their contact status for the multiple future broadcast slots, based on the personal sample data for multiple individuals, wherein the step of creating pseudo-sample data for multiple individuals is to create the pseudo-sample data such that the average contact rate for each future broadcast slot calculated from the created pseudo-sample data for multiple individuals matches the predicted contact rate, and the correlation coefficient between each future broadcast slot calculated from the pseudo-sample data matches the correlation coefficient between each corresponding past broadcast slot calculated from the personal sample data.

[0010] The program according to the present invention causes a computer to function as a predicted contact rate acquisition unit that acquires predicted contact rates for multiple future broadcast slots, a personal sample acquisition unit that acquires personal sample data for multiple individuals that includes contact status for multiple past broadcast slots corresponding to the multiple future broadcast slots, and a pseudo-sample creation unit that creates pseudo-sample data for multiple individuals that includes contact status for multiple future broadcast slots based on the personal sample data for multiple individuals, wherein the pseudo-sample creation unit creates the pseudo-sample data such that the average contact rate for each future broadcast slot calculated from the created pseudo-sample data for multiple individuals matches the predicted contact rate, and the correlation coefficient between each future broadcast slot calculated from the pseudo-sample data matches the correlation coefficient between each corresponding past broadcast slot calculated from the personal sample data. [Effects of the Invention]

[0011] According to the present invention, highly reliable pseudo-sample data, including contact status with future broadcast slots, can be obtained. [Brief explanation of the drawing]

[0012] [Figure 1] A block diagram showing the configuration of a data processing device 1 according to an embodiment of the present invention. [Figure 2] A block diagram showing a functional module of a program executed by the processor 11 of a data processing device 1 according to an embodiment of the present invention. [Figure 3] A flowchart of the calculation process performed by the data processing device 1 according to an embodiment of the present invention. [Figure 4] A diagram illustrating predicted viewership data according to an embodiment of the present invention. [Figure 5] A diagram illustrating personal sample data according to an embodiment of the present invention. [Figure 6] A diagram illustrating pseudo-sample data according to an embodiment of the present invention. [Modes for carrying out the invention]

[0013] Next, embodiments for carrying out the present invention will be described in detail with reference to the drawings. (Embodiment) Figure 1 is a block diagram showing the configuration of a data processing device 1 according to an embodiment of the present invention. The data processing device 1 is composed of one computer or multiple computers connected by a communication line. The data processing device 1 includes a processor 11, a main memory 12, an input / output interface 13, a communication interface 14, and a storage device 15. The storage device 15 is a computer-readable recording medium such as a semiconductor memory (e.g., volatile memory or non-volatile memory) or a disk medium (e.g., a magnetic recording medium or a magneto-optical recording medium). The storage device 15 stores programs to be executed by the processor 11, as well as various data. The programs are read from the storage device 15 into the main memory 12, interpreted and executed by the processor 11, thereby executing various functions.

[0014] Figure 2 is a block diagram showing the functional modules of the program executed by the processor 11 of the data processing device 1. As shown in Figure 2, the functional modules executed by the processor 11 of the data processing device 1 include a predicted contact rate acquisition unit 101, a personal sample acquisition unit 102, a pseudo-sample creation unit 103, and an optimization calculation unit 104.

[0015] The storage device 15 may store data on predicted viewership ratings (predicted contact rates) for future broadcast slots, individual sample data showing viewership ratings (contact status) for each broadcast slot measured for individuals in the past, and pseudo-sample data created based on the individual sample data. Details of predicted viewership ratings, individual sample data, and pseudo-sample data will be described later.

[0016] The data processing device 1 creates pseudo sample data including the audience rating for each broadcast frame assuming future broadcast frames, and performs an optimization calculation using the created pseudo sample data to determine a combination of broadcast frames for achieving an audience rating that satisfies a predetermined condition. For example, it determines a combination of broadcast frames for achieving a target reach under a predetermined constraint condition (such as cost), or a combination of broadcast frames that maximizes the audience rating under a predetermined constraint condition.

[0017] Hereinafter, the calculation process (steps S1 to S4) by the data processing device 1 will be described using the flowchart of FIG. 3. Here, as an example, the optimization of the broadcast frame when advertising C is published on television is performed. The broadcast frame (slot) can be specified, for example, by a broadcast area, a broadcast station, a broadcast date, a start time, and an end time. One slot may be defined by a broadcast station, a broadcast date, and a time zone in this way, but the definition method is not limited to this. For example, a frame may be specified by a broadcast station and a time zone for broadcasting a certain program. Note that the content to be published is not limited to an advertisement, and may be, for example, a specific program, a video, or the like. Also, the medium is not limited to television, and may be a video site or the like.

[0018] (Step S1) First, in the predicted contact rate acquisition unit 101, the predicted audience ratings (predicted contact rates) of a plurality of broadcast frames that are candidates for broadcasting advertisement C in a future period from the current time are acquired. FIG. 4 is a diagram illustrating data of the predicted audience rating according to the present embodiment. In the example of FIG. 4, the broadcast time zone of a program (program code "CXP11", program name "XYZNews") broadcast from 6:10 to 8:00 on weekdays in the broadcast area "11" and the broadcast station "CX" is defined as a unit of broadcast frame (1 slot). Slots 1 to 10 correspond to broadcast frames from 6:10 (start time) to 8:00 (end time) on a plurality of broadcast days (the first half of weekdays, days 1 to 3, 6 to 10, 13 to 14) in the next month (November 2023) as seen from the current time. As shown in FIG. 4, the predicted audience ratings for slots 1 to 10 are shown.

[0019] "Predicted audience rating" is the predicted value of the audience rating for each future broadcast slot. The predicted audience rating can be calculated using any method such as statistical models or machine learning. For example, regression analysis, random forest method, XGBoost, LightGBM, CatBoost, etc. can be mentioned, but other methods may also be used. The predicted audience rating may utilize data provided from the survey period, etc.

[0020] (Step S2) Next, in the individual sample acquisition unit 102, for the broadcast slots in the period up to the current time (past) corresponding to the future broadcast slots in FIG. 4, individual sample data indicating the individual audience ratings are acquired. The individual sample data is data indicating the audience ratings (viewing status) of a plurality of broadcast slots for one viewer. The audience rating can be the length of time (seconds, minutes) that the viewer watched in the broadcast slot, whether there was contact with the channel, etc. (if there was contact, "1", if there was no contact, "0"), the ratio of the viewing time within the broadcast slot (length of viewing time / length of the broadcast slot time), etc.

[0021] FIG. 5 is a diagram illustrating the individual sample data according to the present embodiment. As shown in FIG. 5, for viewer 01, the audience ratings (here, the ratio of the viewing time in each broadcast slot) of a plurality of broadcast slots 101 to 110 are shown. The broadcast slots 101 to 110 correspond to the broadcast slots 1 to 10 shown in FIG. 4, and for example, correspond to the broadcast slots from 6:10 to 8:00 in the past (e.g., the first half of weekdays in October 2023) in the broadcast area "11" and the broadcast station "CX". The individual sample acquisition unit 102 acquires individual sample data similar to FIG. 5 for a plurality of viewers (here, viewers 01 to 10). The individual sample data can utilize viewing data provided from the survey period, etc.

[0022] (Step S3) Next, the pseudo-sample creation unit 103 creates pseudo-sample data for each future broadcast slot for which the predicted viewership rating was calculated in step S1, based on the individual sample data obtained in step S2. That is, it creates pseudo-sample data showing the viewing status of individuals when each broadcast slot in the individual sample data shown in Figure 5 is replaced with future broadcast slots 1 to 10.

[0023] The pseudo-sample generation unit 103 first determines a multidimensional normal distribution (in this case, a 10-dimensional normal distribution) for the viewership ratings (in this case, the proportion of viewing time within each broadcast slot) for each of the 10 broadcast slots. In this embodiment, the pseudo-sample generation unit 103 determines the multidimensional normal distribution so as to satisfy the following two conditions. (Condition 1) The average viewership rating for broadcast slots 1-10 matches the predicted viewership rating for broadcast slots 1-10 obtained in step S1. (Condition 2) The correlation coefficient of viewership between broadcast slots matches the correlation coefficient of viewership between corresponding broadcast slots calculated from the individual sample data obtained in step S2. Here, broadcast slots 101 to 110 in the individual sample data correspond to broadcast slots 1 to 10 in the pseudo-sample data, respectively. For example, the correlation coefficient of viewership between broadcast slot 1 and broadcast slot 2 should match the correlation coefficient of viewership between broadcast slot 101 and broadcast slot 102 in the individual sample data.

[0024] The pseudo-sample creation unit 103 creates pseudo-sample data using values ​​randomly extracted from the determined multidimensional normal distribution. That is, for each broadcast slot, it creates one pseudo-sample data (for one person) using a value extracted from the multidimensional normal distribution as the viewership rating. The pseudo-sample creation unit 103 creates multiple pseudo-sample data (for multiple people). Figure 6 is a diagram illustrating the pseudo-sample data according to this embodiment. As shown in Figure 6, 100 pseudo-sample data sets are created assuming viewers 01 to 100, and the viewership ratings for broadcast slots 1 to 10 are shown for each. Each column in the table in Figure 6 corresponds to the individual sample data of one viewer shown in Figure 5.

[0025] The pseudo-sample creation unit 103 creates each pseudo-sample data according to a multidimensional normal distribution determined to satisfy the above conditions 1 and 2. Therefore, the average viewership rating for broadcast slots 1 to 10 calculated from the pseudo-sample data (in the example in Figure 6, the average viewership rating for viewers 01 to 100) theoretically matches the predicted viewership rating for broadcast slots 1 to 10 obtained in step S1. Furthermore, the correlation coefficient of viewership ratings between broadcast slots calculated from the pseudo-sample data theoretically matches the correlation coefficient of viewership ratings between corresponding broadcast slots calculated from the individual sample data obtained in step S2. For example, the correlation coefficient of viewership ratings between broadcast slot 1 and broadcast slot 2 in the pseudo-sample data theoretically matches the correlation coefficient of viewership ratings between broadcast slot 101 and broadcast slot 102 in the individual sample data.

[0026] (Step S4) Next, the optimization calculation unit 104 uses the pseudo-sample data created in step S3 to perform an optimization calculation of the broadcast slots that should be used to achieve the target reach of advertisement C. Specifically, for example, the optimization (determination of the combination of broadcast slots to be used for advertisement C) may be performed with the objective function set to "maximize the reach of the broadcast slots to be used for advertisement C" and the constraint condition set to "total cost of use is less than or equal to X yen". Alternatively, the optimization (determination of the combination of broadcast slots to be used for advertisement C) may be performed with the objective function set to "reach of Y% or more" and the constraint condition set to "minimize cost". The objective function and constraint conditions for the optimization calculation are not limited to those described above. Note that the optimization calculation may also be performed using, for example, the solver function of a spreadsheet software.

[0027] For example, let's consider an example of optimizing which of the 10 slots (slot(i)) will maximize reach under the constraint that the total cost must be less than or equal to C yen, using a pseudo-sample data set (sample(x,i)) of 100 people. Here, x is the pseudo-sample number (x=1 to 100) and i is the slot number (i=1 to 10). The pseudo-sample data set (sample(x,i)) stores "1" if pseudo-sample x has viewed slot i, and "0" if it has not. Similarly, slot(i) is set to "1" if an ad is placed in slot i, and "0" if no ad is placed.

[0028] The optimization calculation unit 104 performs an optimization calculation to find a solution slot(i) (where i is 1 to 10) that maximizes the objective function P(reach) under the following constraints. (constraints) If the advertising cost for slot i is denoted as Cost(i), then the sum of Cost(i) for each i where Slot(i)=1 will be less than or equal to C. (Objective function) For each pseudo-sample, we ask for the following Q. The value obtained by performing an OR operation on Q = sample(x,i) × slot(i) for i = 1 to 10. The objective function P is obtained by summing Q for the pseudosamples x=1 to 100 and dividing the sum by the pseudosample size of 100.

[0029] The solution slot(i) that maximizes the value of the objective function P can be found using algorithms such as local search. For example, with local search, first, a randomly selected slot(i) that satisfies the constraints is set as the initial value, and the objective function P is calculated. Next, the operation of modifying slot(i) is performed within the neighborhood to satisfy the constraints (for example, modifying it so that the number of modified slots is less than or equal to a threshold), and the objective function P is calculated again. The value of P before and after the modification is evaluated, and if the value of P after the modification has increased, the modified slot(i) is adopted; otherwise, the modification is canceled. The operation of modifying slot(i) and evaluation of the objective function P are repeated until a predetermined termination condition is met (for example, the increment of P before and after the modification is less than or equal to a threshold, the number of iterations exceeds a threshold, etc.), and the finally obtained slot(i) is taken as the solution.

[0030] As described above, according to this embodiment, pseudo-sample data showing the viewership ratings for future broadcast slots was created using personal sample data showing viewership ratings for each broadcast slot that exist as actual data. In this case, the average viewership rating for each broadcast slot in the pseudo-sample data was theoretically made to match the predicted viewership rating separately predicted for each broadcast slot. Furthermore, the correlation coefficient between broadcast slots in the pseudo-sample data was theoretically made to match the correlation coefficient between corresponding broadcast slots in the personal sample data. This makes it possible to create highly reliable pseudo-sample data showing the viewership ratings for future broadcast slots. In other words, highly reliable pseudo-sample data can be created by using the values ​​from personal sample data for the correlation coefficient between data items (viewership ratings for each broadcast slot), while using predicted viewership ratings estimated by statistical methods for the average viewership rating. The correlation between data items is thought to easily reflect individual lifestyles and preferences, and is considered to be relatively unchanging even if the time period is different. For example, individual tendencies such as watching more often on Mondays at the beginning of the week and less often on Fridays at the end of the week are unlikely to change significantly, so the correlation between data items is expected to remain largely unchanged even when using past individual sample data. On the other hand, for average viewership ratings, the reliability of the resulting pseudo-sample data can be increased by using more reliable data obtained through statistical prediction.

[0031] Furthermore, by using the created pseudo-sample data to perform optimization calculations for broadcast slots where advertisements and other content will be placed, highly reliable optimization calculations can be performed.

[0032] Furthermore, in optimization calculations aimed at determining which broadcast slots will reach the largest audience for an advertisement, optimization is often performed by targeting viewers belonging to specific attributes (e.g., women in their 30s), such as gender or age group. When performing such attribute-based optimization calculations, two methods are possible. One method is to obtain individual sample data for each attribute (e.g., women in their teens, women in their 20s, etc.) as shown in Figure 5, and create pseudo-sample data in the same manner as in the example above. In this case, the predicted viewership ratings calculated for each attribute are used. The created pseudo-sample data is based on the assumption of viewers belonging to the same attribute, so by performing optimization calculations using this pseudo-sample data, it is possible to determine the optimal combination of broadcast slots targeting viewers of that attribute.

[0033] Another method involves including attribute information as a data item in both the individual sample data and the pseudo-sample data. In this case, when creating the pseudo-sample data, not only the correlation coefficient between viewership ratings for each broadcast slot, but also the correlation coefficient between attribute information and the viewership ratings for each broadcast slot is calculated to match that of the individual sample. The created pseudo-sample data includes attribute information as a data item, and when performing attribute-specific optimization, the pseudo-sample data for the target attribute can be extracted and the optimization calculation performed.

[0034] It should be noted that the present invention is not limited to the embodiments described above, and can be implemented in various other forms without departing from the spirit of the invention. For this reason, the above embodiments are merely illustrative in all respects and should not be interpreted restrictively. For example, the order of each processing step described above can be arbitrarily changed or executed in parallel, as long as there is no inconsistency in the processing content. In addition, other steps may be added between each processing step. Furthermore, a step described as one step may be divided into multiple steps and executed, and a step described as multiple steps may be considered as one step. [Explanation of Symbols]

[0035] 1…Data Processing Unit 11… Processor 12…Main memory 13… Input / Output Interface 14…Communication Interface 15...Storage device 101... Predicted Contact Rate Acquisition Unit 102…Personal Specimen Acquisition Department 103... Simulated Specimen Preparation Department 104...Optimization Calculation Unit

Claims

1. A predictive contact rate acquisition unit that acquires predicted contact rates for multiple future broadcast slots, A personal sample acquisition unit acquires personal sample data for multiple individuals, including their contact status with multiple past broadcast slots corresponding to multiple future broadcast slots. The system includes a pseudo-sample creation unit that creates pseudo-sample data for multiple individuals, including contact status with multiple future broadcast slots, based on the aforementioned individual sample data for multiple individuals, The aforementioned pseudo-specimen preparation unit is: A data processing device that creates pseudo-sample data such that the average contact rate to each future broadcast slot calculated from the created pseudo-sample data of multiple people matches the predicted contact rate, and the correlation coefficient between each future broadcast slot calculated from the pseudo-sample data matches the correlation coefficient between each corresponding past broadcast slot calculated from the individual sample data.

2. The data processing apparatus according to claim 1, further comprising an optimization calculation unit that uses the pseudo-sample data to calculate one or more future broadcast slots that are expected to have a reach that satisfies predetermined conditions.

3. The aforementioned personal sample data includes attribute information. The aforementioned pseudo-specimen preparation unit is: The data processing apparatus according to claim 1, wherein the pseudo-sample data is created such that the correlation coefficient between each future broadcast slot calculated from the pseudo-sample data and the attribute information matches the correlation coefficient between each corresponding past broadcast slot calculated from the personal sample data and the attribute information.

4. Computers The process of obtaining predicted contact rates for multiple future broadcast slots, A process to acquire personal sample data for multiple individuals, including their contact history with multiple past broadcast slots corresponding to multiple future broadcast slots, The process includes creating pseudo-sample data for multiple individuals, including their contact status with multiple future broadcast slots, based on the aforementioned individual sample data for multiple individuals. The process of creating the aforementioned pseudo-sample data for multiple individuals is as follows: A data processing method for creating pseudo-sample data such that the average contact rate to each future broadcast slot calculated from the created pseudo-sample data of multiple individuals matches the predicted contact rate, and the correlation coefficient between each future broadcast slot calculated from the pseudo-sample data matches the correlation coefficient between each corresponding past broadcast slot calculated from the individual sample data.

5. Computers, A predictive contact rate acquisition unit that acquires predicted contact rates for multiple future broadcast slots, A personal sample acquisition unit acquires personal sample data for multiple individuals, including their contact status with multiple past broadcast slots corresponding to multiple future broadcast slots. Based on the individual sample data of the aforementioned multiple individuals, it functions as a pseudo-sample creation unit that creates pseudo-sample data for multiple individuals, including their contact status with the aforementioned future broadcast slots. The aforementioned pseudo-specimen preparation unit is: A program that creates pseudo-sample data such that the average contact rate to each future broadcast slot calculated from the created pseudo-sample data of multiple individuals matches the predicted contact rate, and the correlation coefficient between each future broadcast slot calculated from the pseudo-sample data matches the correlation coefficient between each corresponding past broadcast slot calculated from the individual sample data.

Citation Information

Patent Citations

  • Information processing device, revision assistance method, and revision assistance program

    JP2021165927A

  • Dummy sample making device, method for making dummy sample, and program

    JP2022028370A