Learning device, estimation device, learning method, estimation method, and program
The learning device enhances the accuracy of advertising effect estimation for video advertisements by training a machine learning model with both training video and purchasing information, addressing the limitations of existing technologies.
Patent Information
- Application Number
- JP2024558298
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2043-08-29
AI Technical Summary
Existing technologies for estimating the advertising effect of video advertisements have insufficient accuracy due to not fully considering the content of actually distributed video advertisements.
A learning device that includes a training video information acquisition unit, a training purchasing information acquisition unit, and a learning unit that trains a machine learning model to estimate the advertising effect based on analyzed training video information and purchasing information.
The proposed solution increases the accuracy of estimating the advertising effect of video advertisements by considering both behavioral data and the content of distributed video advertisements.
Smart Images

Figure 0007672590000001 
Figure 0007672590000002 
Figure 0007672590000003
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a learning device, an estimation device, a learning method, an estimation method, and a program. [Background technology]
[0002] Conventionally, there are known techniques for estimating the advertising effect of video advertising. For example, Patent Document 1 describes a method in which behavioral data and physiological data of a viewer who has viewed media content such as a video advertisement are processed to obtain data points of a time-series emotional state, and a classification model that maps between the effect data and a prediction parameter of the data points of the time-series emotional state outputs predicted effect data indicating the predicted effect of the media content. The prediction parameter in Patent Document 1 is a quantitative index of a relative change in the viewer's response to the video advertisement. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2020-501260 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the classification model of Patent Document 1 outputs predicted effect data by considering only data based on the behavior of the viewer of the video advertisement, so the content of the video advertisement actually delivered is not fully taken into consideration in estimating the advertising effect of the video advertisement. Therefore, the technology of Patent Document 1 sometimes does not have sufficient accuracy in estimating the advertising effect of the video advertisement. This point is the same for other conventional technologies other than the technology of Patent Document 1.
[0005] One of the objectives of the present disclosure is to improve the accuracy of estimating the advertising effectiveness of video advertisements. [Means for solving the problem]
[0006] The learning device of the present disclosure includes a training video information acquisition unit that acquires training video information acquired by analyzing a training video advertisement that is a video advertisement for training, a training purchasing information acquisition unit that acquires training purchasing information regarding purchases of training viewers who are viewers of the training video advertisement, and a learning unit that learns a machine learning model that estimates the advertising effectiveness of an estimated video advertisement from estimated video information acquired by analyzing an estimated video advertisement that is a video advertisement for estimation based on the training video information and the training purchasing information.
[0007] The estimation device of the present disclosure includes an estimated video information acquisition unit that acquires estimated video information acquired by analyzing an estimated video advertisement that is a video advertisement for estimation, a model memory unit that stores a machine learning model that has been trained based on training video information acquired by analyzing a training video advertisement that is a video advertisement for training and training purchase information regarding purchases of training viewers who are viewers of the training video advertisement, and an estimation unit that estimates the advertising effect of the estimated video advertisement based on the estimated video information and the machine learning model. Effect of the Invention
[0008] According to the present disclosure, it is possible to improve the accuracy of estimating the advertising effectiveness of a video advertisement. [Brief description of the drawings]
[0009] [Figure 1] FIG. 2 is a diagram illustrating an example of a hardware configuration of each of a learning device and an estimation device. [Diagram 2] A figure showing an example of a training video advertisement distributed via a live distribution service. [Diagram 3] FIG. 2 is a diagram illustrating an example of an overview of a learning device and an estimation device. [Figure 4] FIG. 2 is a diagram illustrating an example of functions realized by each of a learning device and an estimation device. [Diagram 5] FIG. 2 is a diagram illustrating an example of a training database. [Figure 6] FIG. 13 is a diagram illustrating an example of an estimation database. [Figure 7] FIG. 2 is a diagram illustrating an example of a video advertisement database. [Figure 8] FIG. 2 illustrates an example of a process executed by a learning device. [Figure 9] FIG. 11 is a diagram illustrating an example of a process executed in S10. [Figure 10] FIG. 4 is a diagram illustrating an example of a process executed by the estimation device. [Figure 11] FIG. 11 is a diagram illustrating an example of a process executed in S20. [Figure 12] FIG. 13 is a diagram illustrating an example of a function realized in a modified example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] [1. Hardware configuration of the learning device and the estimation device] An example of an embodiment of a learning device, an estimation device, a learning method, an estimation method, and a program according to the present disclosure will be described. In this embodiment, an example is given in which the learning device and the estimation device are different devices, but the learning device and the estimation device may be the same device. In other words, a certain device may function as both a learning device and an estimation device.
[0011] In this embodiment, the learning device learns a machine learning model that estimates the advertising effectiveness of a video advertisement, which is an advertisement that uses a video. The estimation device performs estimation based on the trained machine learning model. Hereinafter, a training (learning) video advertisement used in learning the machine learning model is referred to as a training video advertisement. The performers of the training video advertisement are referred to as training performers. The viewers of the training video advertisement are referred to as training viewers. A video advertisement for estimation that is the subject of processing by the trained machine learning model is referred to as an estimated video advertisement. The performers of the estimated video advertisement are referred to as estimated performers. The viewers of the estimated video advertisement are referred to as estimated viewers.
[0012] In this embodiment, when there is no need to distinguish between training video advertisements and estimated video advertisements, they may simply be referred to as video advertisements. When there is no need to distinguish between training performers and estimated performers, they may simply be referred to as performers. When there is no need to distinguish between training viewers and estimated viewers, they may simply be referred to as viewers. In addition, when there is no need to distinguish between terms such as training XX and estimated XX, they may simply be referred to as XX.
[0013] Fig. 1 is a diagram showing an example of the hardware configuration of each of a learning device and an estimation device. In this embodiment, an example of a live distribution system 1 including a learning device 10 and an estimation device 20 will be described. In the example of Fig. 1, the live distribution system 1 includes a server 30, a training performer device 40, a training viewer device 50, an estimated performer device 60, and an estimated viewer device 70 in addition to the learning device 10 and the estimation device 20. Each of the learning device 10, the estimation device 20, the server 30, the training performer device 40, the training viewer device 50, the estimated performer device 60, and the estimated viewer device 70 is connected to a network N such as the Internet or a LAN.
[0014] The learning device 10 is a computer that learns a machine learning model. For example, the learning device 10 is a personal computer, a server computer, a tablet, or a smartphone. For example, the learning device 10 includes a control unit 11, a memory unit 12, a communication unit 13, an operation unit 14, and a display unit 15. The control unit 11 includes at least one processor. The memory unit 12 includes at least one of a volatile memory such as a RAM and a non-volatile memory such as a flash memory. The communication unit 13 includes at least one of a communication interface for wired communication and a communication interface for wireless communication. The operation unit 14 is an input device such as a touch panel. The display unit 15 is a liquid crystal or organic EL display.
[0015] The estimation device 20 is a computer that performs estimation based on a trained machine learning model. For example, the estimation device 20 is a personal computer, a server computer, a tablet, or a smartphone. For example, the estimation device 20 includes a control unit 21, a storage unit 22, a communication unit 23, an operation unit 24, and a display unit 25. The hardware configurations of the control unit 21, the storage unit 22, the communication unit 23, the operation unit 24, and the display unit 25 may be similar to those of the control unit 11, the storage unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively.
[0016] Server 30 is a server computer managed by an operator of a live streaming service in which videos are distributed in real time to an unspecified number of people. In this embodiment, an example is given in which the operator of the live streaming service also manages learning device 10 and estimation device 20, but learning device 10 and estimation device 20 may be managed by someone else. For example, server 30 includes a control unit 31, a storage unit 32, and a communication unit 33. The hardware configurations of control unit 31, storage unit 32, and communication unit 33 may be similar to those of control unit 11, storage unit 12, and communication unit 13, respectively.
[0017] The training performer device 40 is a computer of the training performer. For example, the training performer device 40 is a personal computer, a tablet, or a smartphone. For example, the training performer device 40 includes a control unit 41, a memory unit 42, a communication unit 43, an operation unit 44, and a display unit 45. The hardware configurations of the control unit 41, the memory unit 42, the communication unit 43, the operation unit 44, and the display unit 45 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively. A photographing unit 46 is connected to the training performer device 40. The photographing unit 46 includes at least one camera. The photographing unit 46 may be included inside the training performer device 40.
[0018] The training viewer device 50 is a computer of the training viewer. For example, the training viewer device 50 is a personal computer, a tablet, or a smartphone. For example, the training viewer device 50 includes a control unit 51, a memory unit 52, a communication unit 53, an operation unit 54, and a display unit 55. The hardware configurations of the control unit 51, the memory unit 52, the communication unit 53, the operation unit 54, and the display unit 55 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively.
[0019] The estimated performer device 60 is a computer of the estimated performer. For example, the estimated performer device 60 is a personal computer, a tablet, or a smartphone. For example, the estimated performer device 60 includes a control unit 61, a storage unit 62, a communication unit 63, an operation unit 64, and a display unit 65, and is connected to a shooting unit 66. The hardware configurations of the control unit 61, the storage unit 62, the communication unit 63, the operation unit 64, the display unit 65, and the shooting unit 66 are similar to those of the control unit 11, the storage unit 12, the communication unit 13, the operation unit 14, the display unit 15, and the shooting unit 46, respectively. The shooting unit 66 may be included inside the estimated performer device 60.
[0020] The estimated viewer device 70 is a computer of the estimated viewer. For example, the estimated viewer device 70 is a personal computer, a tablet, or a smartphone. For example, the estimated viewer device 70 includes a control unit 71, a memory unit 72, a communication unit 73, an operation unit 74, and a display unit 75. The hardware configurations of the control unit 71, the memory unit 72, the communication unit 73, the operation unit 74, and the display unit 75 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 14, and the display unit 15, respectively.
[0021] The programs stored in the memories 12, 22, 32, 42, 52, 62, 72 may be supplied to the learning device 10, the estimation device 20, the server 30, the training performer device 40, the training viewer device 50, the estimated performer device 60, or the estimated viewer device 70 via the network N. Also, the programs stored in a computer-readable information storage medium may be supplied to the learning device 10, the estimation device 20, the server 30, the training performer device 40, the training viewer device 50, the estimated performer device 60, or the estimated viewer device 70 via a reading unit (e.g., an optical disk drive or a memory card slot) that reads the information storage medium, or an input / output unit (e.g., a USB port) that inputs and outputs data to and from an external device.
[0022] Furthermore, the live distribution system 1 only needs to include at least one computer, and is not limited to the example of Fig. 1. For example, the live distribution system 1 may include a learning device 10, an estimation device 20, and a server 30, and may not include a training performer device 40, a training viewer device 50, an estimated performer device 60, and an estimated viewer device 70. In this case, the training performer device 40, the training viewer device 50, the estimated performer device 60, and the estimated viewer device 70 exist outside the live distribution system 1. As another example, the learning device 10 and the estimation device 20 may not be included in the live distribution system 1, and may exist outside the live distribution system 1.
[0023] [2. Overview of the learning device and the estimation device] 2 is a diagram showing an example of a training video advertisement distributed by a live distribution service. For example, a training performer advertises a product or service that is the subject of electronic commerce in front of the filming unit 46. The training performer device 40 transmits the training video advertisement generated by the filming unit 46 to the server 30. The server 30 distributes the training video advertisement received from the training performer device 40 to the training viewer devices 50 of an unspecified number of training viewers in real time. A screen SC showing the training video advertisement is displayed on the display unit 55 of the training viewer device 50.
[0024] For example, when the training viewer selects button B1, the training viewer device 50 accesses a page showing details of the product or service introduced in the training video advertisement. The training viewer can purchase the product or service from the page. The details may be displayed on the screen SC. When the training viewer selects button B2, the training viewer device 50 accesses a page showing coupons for the product or service introduced in the training video advertisement. The training viewer can obtain a coupon for the product on the page. The coupon may be obtainable on the screen SC.
[0025] For example, the training viewer can input a comment on the training video advertisement in the input form F. The comments input by the training viewer and the other training viewers are displayed in the display area A of the screen SC. One of the goals of the above-mentioned live distribution service is to maximize the advertising effect. The advertising effect can also be called conversion. For example, the advertising effect is the number of sales of a product or service, the sales amount, the number of additions to a shopping cart, the number of additions to a bookmark, the number of users who registered the store as a favorite from the video advertisement page, or a combination of these. It is considered that the advertising effect is influenced by various factors. In this embodiment, the learning device 10 learns a machine learning model that estimates the advertising effect in the live distribution service. The estimation device 20 performs estimation using the trained machine learning model.
[0026] FIG. 3 is a diagram showing an example of an overview of each of the learning device 10 and the estimation device 20. In this embodiment, the learning device 10 acquires training video information indicating characteristics of the training video advertisement by performing various analyses on the training video advertisement that has already been distributed. The training video information indicates characteristics that are thought to have a causal relationship with the advertising effect of the training video advertisement. Although the method of acquiring the training video information will be described in detail later, for example, the learning device 10 analyzes the facial expressions of the training performers, the responses by the training performers, the explanations by the training performers, and the like, and acquires multiple pieces of training video information.
[0027] In this embodiment, the learning device 10 acquires training viewer information related to purchases made by the training viewer by tracking the behavior of the training viewer. The training viewer information corresponds to an index showing the advertising effectiveness of the training video advertisement. Details of a method for acquiring training purchase information will be described later, but for example, the training viewer information is acquired by tracking whether or not a purchase has been made by the training viewer. Although one piece of training viewer information is shown in FIG. 3, the learning device 10 may acquire multiple pieces of training viewer information.
[0028] In the example of FIG. 3, the learning device 10 generates training data for the machine learning model M to learn based on multiple pieces of training video information and one piece of training viewer information. Details of the training data will also be described later. For example, the learning device 10 may execute a similar process for each of multiple training video advertisements to generate training data one after another. The learning device 10 learns the machine learning model M based on the training data. Details of the learning will also be described later. The learning device 10 transmits the trained machine learning model M to the estimation device 20. The estimation device 20 records the trained machine learning model M received from the learning device 10 in the memory unit 22.
[0029] For example, the estimation device 20 acquires estimated video information by performing various analyses on the estimated video advertisement in order to retroactively analyze the advertising effectiveness of the estimated video advertisement that has already been distributed. In this embodiment, the estimated video advertisement is distributed in a similar manner to the training video advertisement described in FIG. 2. For example, the estimated performer operates the estimated performer device 60, films himself / herself with the filming unit 66, and introduces the product or service. The server 30 distributes the estimated video advertisement to the estimated viewer devices 70 of an unspecified number of estimated viewers. Details of the estimated video information will be described later, but the method of acquiring the estimated video information may be the same as the method of acquiring the training video information.
[0030] For example, the estimation device 20 inputs a plurality of pieces of estimated video information to the machine learning model M. The machine learning model M calculates features (embedded expressions) based on the plurality of pieces of estimated video information. The machine learning model M estimates the advertising effectiveness of the estimated video advertisement based on the features. For example, the machine learning model M estimates the probability that an estimated viewer who has viewed the estimated video advertisement will purchase a product or service.
[0031] For example, the estimation device 20 obtains an estimation result output from the machine learning model M. The estimation device 20 can use the estimation result of the machine learning model M for any purpose. For example, the estimation device 20 may display the estimation result of the machine learning model M on the display unit 25. The estimation device 20 may transmit the estimation result of the machine learning model M to the estimated performer device 60 in order to provide feedback on the advertising effect of the estimated video advertisement to the estimated performer.
[0032] As described above, the learning device 10 of this embodiment creates a machine learning model M with high estimation accuracy of the advertising effectiveness by learning the machine learning model M based on the training video information and the training purchase information. The estimation device 20 can accurately estimate the advertising effectiveness of the estimated video advertisement by performing estimation based on the machine learning model M learned by the learning device 10. Hereinafter, the details of this embodiment will be described.
[0033] [3. Functions Realized by the Learning Device and the Estimation Device] Fig. 4 is a diagram showing an example of functions realized by each of the learning device 10 and the estimation device 20. Fig. 4 also shows functions realized by the server 30.
[0034] [3-1. Functions realized by the learning device] For example, the learning device 10 includes a data storage unit 100, a model storage unit 101, a training video information acquisition unit 102, a training purchase information acquisition unit 103, and a learning unit 104. The data storage unit 100 and the model storage unit 101 are realized by the storage unit 12. The training video information acquisition unit 102, the training purchase information acquisition unit 103, and the learning unit 104 are realized by the control unit 11.
[0035] [Data storage section] The data storage unit 100 stores data necessary for learning the machine learning model M. For example, the data storage unit 100 stores a training database DB1.
[0036] 5 is a diagram showing an example of the training database DB1. The training database DB1 is a database in which training data to be learned by the machine learning model M is stored. In this embodiment, each unit of data used when learning the machine learning model M is called training data. A collection of training data in which multiple training data are stored is called the training database DB1. All or a part of the multiple training data stored in the training database DB1 is learned by the machine learning model M.
[0037] In the example of FIG. 5, "No" is the record number of the training database DB1. In this embodiment, an example is taken of one training data being generated from one training video advertisement. That is, the training video advertisement and the training data have a one-to-one relationship. Therefore, "No" in FIG. 5 can also be said to be information that can identify the training video advertisement. Note that multiple training data may be generated from one training video advertisement, or one training data may be generated from multiple training video advertisements.
[0038] For example, the training data includes an input portion that is input to the machine learning model M during learning, and an output portion that should be output from the machine learning model M during learning. The input portion of the training data has the same format as the input data that is input to the machine learning model M during estimation. The input portion of the training data is sometimes called an explanatory variable. In this embodiment, the input portion of the training data is a plurality of training video information. The output portion of the training data has the same format as the output data that is output from the machine learning model M during estimation. The output portion of the training data corresponds to the correct answer during learning. The output portion of the training data is sometimes called a target variable. In this embodiment, an example is given in which the output portion of the training data is training purchasing information.
[0039] The data storage unit 100 can store any data. The data stored in the data storage unit 100 is not limited to the training database DB1. For example, the data storage unit 100 may store a learning program indicating a series of processes in learning the machine learning model M. The data storage unit 100 may store video data of a training video advertisement downloaded by the learning device 10 from the server 30. The data storage unit 100 may store data associated with the video data of a training video advertisement (for example, a category of a product or service introduced in the training video advertisement, a description input in advance, coupon information, or a combination of these, etc.).
[0040] For example, the data storage unit 100 may store a program for natural language processing, which will be described later. The data storage unit 100 may store a dictionary database in which keywords used in natural language processing are stored. The data storage unit 100 may store a machine learning model M for natural language processing. The data storage unit 100 may store a program for image analysis processing, which will be described later. The data storage unit 100 may store a machine learning model M for image analysis processing.
[0041] [Model memory section] The model storage unit 101 stores the machine learning model M. The machine learning model M includes a program portion indicating processing such as calculation of feature amounts (embedded representation) of input data input to the model M itself, and a parameter portion (e.g., weighting coefficients and biases) referenced by the program portion. The parameter portion of the machine learning model M is changed by learning by the learning unit 104. The model storage unit 101 stores the machine learning model M whose parameter portion has initial values. The machine learning model M whose parameter portion has initial values is the machine learning model M before learning. The parameter portion whose parameter portion has initial values is overwritten by processing by the learning unit 104.
[0042] The machine learning model M is a model that uses a machine learning technique. The machine learning itself can use various known techniques. For example, the machine learning model M may be a model that uses any of supervised learning, semi-supervised learning, or unsupervised learning. In this embodiment, the machine learning model M is a neural network, but the machine learning model M may be a model that uses other techniques and is not limited to a neural network. For example, the machine learning model M may be a model that uses linear regression, a model that uses logistic regression, a model that uses a support vector machine, or a model that uses a decision tree.
[0043] [Training video information acquisition department] The training video information acquisition unit 102 acquires training video information acquired by analyzing a training video advertisement, which is a video advertisement for training. In this embodiment, a case where a video advertisement that has been distributed in the past corresponds to the training video advertisement is taken as an example, but a video advertisement that has not been distributed may also correspond to the training video advertisement. In addition, a case where a video of a live distribution service corresponds to the training video advertisement is taken as an example, but the training video advertisement may be any video advertisement used for learning, and is not limited to a video of a live distribution service. For example, a web advertisement, a television advertisement, an outdoor advertisement, a movie advertisement, or other advertisements may correspond to the training video advertisement.
[0044] The analysis of the training video advertisement is a process of extracting features related to the training video advertisement. For example, voice analysis processing, natural language processing, image analysis processing, or a combination of these corresponds to the analysis of the training video advertisement. The training video information indicates the features extracted by the analysis of the training video advertisement. Hereinafter, an example of the features acquired by the training video information acquisition unit 102 as the training video information will be described. Note that the training video information acquisition unit 102 may acquire, as the training video information, features that a person who creates the machine learning model M (e.g., an operator of a live streaming service) believes to have some causal relationship with the advertising effectiveness. The features acquired by the training video information acquisition unit 102 as the training video information are not limited to the example of this embodiment.
[0045] In this embodiment, the training video information acquisition unit 102 acquires training video information acquired by performing natural language processing on the training text acquired by analyzing the audio of the training video advertisement. For example, the audio of the training video advertisement is audio produced by the training performer. The audio of the training video advertisement may be any audio included in the training video advertisement, and is not limited to audio produced by the training performer. For example, the audio of the training video advertisement may be audio produced by a staff member other than the training performer, an audience watching the training video advertisement on-site, a third party interviewed by the training performer, or an artificial audio prepared in advance.
[0046] The training text is a text generated from the voice of the training video advertisement. For example, the training video information acquisition unit 102 generates the training text by performing a voice analysis process on the voice of the training video advertisement. The algorithm of the voice analysis process itself may be a known algorithm. For example, the training information acquisition unit generates the training text from the voice of the training video advertisement based on a method using a hidden Markov model, a method using a predetermined waveform pattern, a machine learning method such as a neural network, or other methods. The training text may be generated by a computer other than the learning device 10. In this case, the training video information acquisition unit 102 acquires the training text generated by the other computer from the other computer, another computer, or an information storage medium.
[0047] For example, the training video information acquisition unit 102 acquires the training video information by performing natural language processing on the training text. Natural language processing is processing for a computer to recognize language used by humans. The natural language processing may be executed by a computer other than the learning device 10. In this case, the training video information acquisition unit 102 acquires the training text generated by the natural language processing executed by the other computer from the other computer, from another computer, or from an information storage medium.
[0048] For example, the natural language processing may be at least one of a sentiment analysis process that analyzes the emotions of a training performer who appears in the training video advertisement, a response detection process that detects a response by a training performer who appears in the training video advertisement, an explanation detection process that detects an explanation regarding a product or service featured in the training video advertisement, and a promotion detection process that detects a promotion related to the training video advertisement.
[0049] For example, the training video information acquisition unit 102 acquires training video information related to the emotions of the training performers by performing an emotion analysis process on the training text. The emotion analysis process may be a known method. For example, the training video information acquisition unit 102 may analyze the emotions of the training performers based on a rule approach that determines whether or not a keyword indicating an emotion is included in the training text. As another example, the training video information acquisition unit 102 may analyze the emotions of the training performers based on a machine learning approach that calculates the feature amount of the training text and outputs an emotion classification result.
[0050] For example, the training video information acquisition unit 102 acquires training video information indicating an analysis result of the emotions of the training performers by emotion analysis processing. Assuming that the four emotions of joy, anger, sadness, and happiness are analyzed, the training video information acquisition unit 102 acquires training video information indicating the degree of the emotion of "joy" (for example, the number of occurrences or frequency of occurrence of keywords indicating "joy"), the degree of the emotion of "anger" (for example, the number of occurrences or frequency of occurrence of keywords indicating "anger"), the degree of the emotion of "sadness" (for example, the number of occurrences or frequency of occurrence of keywords indicating "sadness"), and the degree of the emotion of "happy" (for example, the number of occurrences or frequency of occurrence of keywords indicating "happy"). The training video information acquisition unit 102 may analyze emotions other than joy, anger, sadness, and happiness.
[0051] For example, the training video information acquisition unit 102 acquires training video information related to the responses of the training performers by executing a response detection process on the training text. The method of the response detection process may be a known method. For example, the training video information acquisition unit 102 may detect the responses of the training performers based on a rule approach that determines whether or not a keyword indicating a response is included in the training text. As another example, the training video information acquisition unit 102 may detect the responses of the training performers based on a machine learning approach that calculates the feature amount of the training text and outputs a classification result of the response.
[0052] The response detected by the response detection process can also be called a reaction. The response detected by the response detection process can be a response of a training performer to a statement made by another training performer. In this case, the response can also be called a dialogue between the training performers. The training video information acquisition unit 102 acquires training video information indicating the detection result of the response of the training performer by the response detection process. For example, when a statement such as "I see" is included in the training text, the training video information acquisition unit 102 detects the response of the other training performer to the statement made by a training performer.
[0053] For another example, the response detected by the response detection process may be a response of the training performer to a comment input by the training viewer. When a statement such as "Thank you for your comment" is included in the training text, the training video information acquisition unit 102 detects the response of the training performer to the comment input by the training viewer. The training video information acquisition unit 102 may acquire training video information indicating the number of responses detected by the response detection process, the frequency of the responses, the type of the responses (for example, what the response is in response to, or the specific content of the response, etc.), or a combination thereof.
[0054] For example, the training video information acquisition unit 102 acquires training video information related to the explanation of the training performer by executing an explanation detection process on the training text. The explanation detection process may be performed by a known method. For example, the training video information acquisition unit 102 may detect the response of the training performer based on a rule approach that determines whether or not a keyword indicating a description of a product or service is included in the training text. As another example, the training video information acquisition unit 102 may detect the explanation by the training performer based on a machine learning approach that calculates the feature amount of the training text and outputs a classification result of the explanation of the product or service.
[0055] For example, the training video information acquisition unit 102 acquires training video information indicating the detection result of the explanation by the training performer through the explanation detection process. If sweets are introduced in the training video advertisement, the training video information acquisition unit 102 counts the number of occurrences or frequency of occurrence of a keyword indicating the product itself, such as "chocolate," in the training text. The training video information acquisition unit 102 counts the number of occurrences or frequency of occurrence of a keyword indicating the quality of the product, such as "delicious," in the training text. The training video information acquisition unit 102 acquires training video information indicating these count results.
[0056] For example, the promotion detection process may be a keyword detection process that detects at least one of a keyword related to the price of a product or service introduced in the training video advertisement, a keyword related to sales of the product or service introduced in the training video advertisement, and a keyword related to a description associated with the training video advertisement. The promotion detection process may also be a process of detecting words that lead to the promotion of a product or service. The training video information acquisition unit 102 may detect the response of the training performer based on a rule approach that determines whether a keyword indicating the price of a product or service is included in the training text.
[0057] For example, the training video information acquisition unit 102 determines whether or not a keyword stored in a dictionary database storing keywords related to price (e.g., "1000 yen," "bargain," "discount," or "coupon") is included in the training text, and detects the keyword. The training video information acquisition unit 102 determines whether or not a keyword stored in a dictionary database storing keywords related to sales (e.g., "purchase" or "only a few left") is included in the training text, and detects the keyword.
[0058] The explanation associated with the training video advertisement is an explanation regarding the product or service introduced in the training video advertisement. For example, information that the training viewer can refer to as appropriate from the screen SC may correspond to the explanation associated with the training video advertisement. In the example of FIG. 2, the detailed explanation that the training viewer can refer to by selecting button B1 is an example of an explanation associated with the training video advertisement. Coupon information that the training viewer can refer to by selecting button B2 is an example of an explanation associated with the training video advertisement.
[0059] For example, the training video information acquisition unit 102 determines whether or not a keyword stored in a dictionary database that stores explanatory keywords (e.g., "limited time only," "delicious," or "topical") related to a product or service introduced in a training video advertisement is included in the training text, and detects the keyword. Based on the result of the detection of such keywords, the training video information acquisition unit 102 acquires training video information indicating the number of detections or the detection frequency of the keywords.
[0060] The promotion detection process is not limited to the above example. For example, the training video information acquisition unit 102 may detect a promotion based on a machine learning approach that calculates the feature amount of a training text and outputs an estimated result of the promotion, instead of a rule approach using keywords. The model used in the machine learning approach is assumed to have learned various words that are related to the promotion of products or services.
[0061] Furthermore, the training video information acquisition unit 102 can perform any natural language processing on the training text. The natural language processing performed by the training video information acquisition unit 102 is not limited to the example of this embodiment. For example, the training video information acquisition unit 102 may perform morphological analysis, syntactic analysis, semantic analysis, feature extraction, or a combination of these as natural language processing on the training text, and acquire training video information based on the execution result of the natural language processing. The training video information acquisition unit 102 may perform natural language processing on the training text using a machine learning technique such as a transformer.
[0062] For example, the training video information acquisition unit 102 may acquire training video information related to facial expressions of a training performer who is a performer in the training video advertisement, acquired by analyzing the video of the training video advertisement. The video of the training video advertisement includes a plurality of frames (still images). The training video information acquisition unit 102 acquires the training video information by executing an image analysis process on each frame. The image analysis process may be executed by a computer other than the learning device 10. In this case, the training video information acquisition unit 102 acquires the training video information generated by the image analysis process executed by the other computer from the other computer, yet another computer, or an information storage medium.
[0063] The facial expression of the training performer can be said to be the emotion of the training performer. For this reason, the training video information acquisition unit 102 may identify an emotion similar to that of the emotion analysis process by image analysis processing. The method of analyzing human facial expressions by image analysis processing may be a known method. For example, the training video information acquisition unit 102 may analyze the facial expression of the training performer in the training video advertisement based on template matching using a template image showing basic human facial expressions. The training video information acquisition unit 102 may analyze the facial expression of the training performer in the training video advertisement based on an arrangement pattern of feature points detected from the video of the training video advertisement. The training video information acquisition unit 102 may analyze the facial expression of the training performer in the training video advertisement based on a machine learning model in which various human facial expressions have been learned.
[0064] The training video information acquisition unit 102 may analyze features other than the facial expressions of the training performers by image analysis processing. For example, the training video information acquisition unit 102 may analyze the camera work, the brightness of the lighting, the gestures of the training performers, the effects in the video, the contents of the subtitles, or other features in the training video advertisement by image analysis processing. The training video information acquisition unit 102 may acquire training video information indicating these features.
[0065] In addition, in the present embodiment, the case where the training video information acquisition unit 102 acquires a plurality of pieces of training video information is taken as an example, but the training video information acquisition unit 102 only needs to acquire at least one piece of training video information. The training video information acquisition unit 102 may acquire only one piece of training video information. For example, the training video information acquisition unit 102 may execute only one of the above-described methods and acquire only one piece of training video information.
[0066] [Training and purchasing information acquisition department] The training purchase information acquisition unit 103 acquires training purchase information regarding purchases made by training viewers who are viewers of the training video advertisement. A purchase by a training viewer is the purchase of a product or service introduced in a training video advertisement. A training viewer may make a purchase while viewing a training video advertisement, or may make a purchase after viewing a training video advertisement. A training viewer may make a purchase not during live distribution of a training video advertisement, but by viewing an archived training video advertisement after the live distribution has ended.
[0067] 2, the training viewer may select button B1 on screen SC and then make a purchase from a page displayed on the training viewer device 50, or may make a purchase by an operation other than selecting button B1. For example, the training viewer may close screen SC for the time being, select a link to a page for a product or service introduced in the training video advertisement, and then make a purchase from a page displayed on the training viewer device 50.
[0068] In this embodiment, the distribution unit 301 of the server 30 tracks the behavior of the training viewers and generates training purchasing information. The training purchasing information acquisition unit 103 acquires the training purchasing information from the server 30. The training purchasing information may be generated by a computer other than the server 30. The training purchasing information acquisition unit 103 may acquire the training purchasing information from the other computer, a further computer, or an information storage medium. The training purchasing information acquisition unit 103 may acquire the training purchasing information by generating the training purchasing information itself.
[0069] In this embodiment, the training purchase information acquisition unit 103 acquires training purchase information indicating the ratio of training viewers who made a purchase among the training viewers who viewed the training video advertisement (total number of training viewers who made a purchase / total number of training viewers). The training purchase information may be any information related to purchases by the training viewers, and is not limited to the example of this embodiment. For example, the training purchase information acquisition unit 103 may acquire training purchase information related to at least one of the presence or absence of a purchase by the training viewer and information related to the sales of a product or service introduced in the training video advertisement.
[0070] The training purchasing information regarding the presence or absence of a purchase is information generated based on the presence or absence of a purchase by the training viewer. The above-mentioned ratio is also an example of the training purchasing information regarding the presence or absence of a purchase, but the training purchasing information regarding the presence or absence of a purchase may be other information. For example, the training purchasing information regarding the presence or absence of a purchase may be information indicating whether or not each training viewer made a purchase, the total number of training viewers who made a purchase, the total number of training viewers per unit time, or a combination of these. The unit time may be any length, for example, 1 minute or 5 minutes. For example, the server 30 generates the training purchasing information by aggregating this information.
[0071] The information on sales of the product or service introduced in the training video advertisement is information generated based on the sales of the product or service. For example, the information on sales may be the number of sales (sales volume) of the product or service, the sales amount, the number of sales per unit time, the sales amount per unit time, or a combination of these. For example, the server 30 generates the training purchase information by aggregating this information.
[0072] [Learning Department] The learning unit 104 learns a machine learning model M that estimates the advertising effectiveness of an estimated video advertisement from estimated video information acquired by analyzing the estimated video advertisement, which is a video advertisement for estimation, based on the training video information and the training purchase information. The learning unit 104 generates training data based on the training video information and the training purchase information. In this embodiment, the learning unit 104 generates training data such that the training video information is an input portion of the training data and the training purchase information is an output portion of the training data.
[0073] The learning unit 104 may use the training video information as the input part of the training data after performing some processing such as normalization or aggregation on the training video information, rather than using the training video information as the input part of the training data as it is. Similarly, the learning unit 104 may use the training purchase information as the output part of the training data after performing some processing such as normalization or aggregation on the training purchase information, rather than using the training purchase information as the output part of the training data as it is.
[0074] For example, the learning unit 104 learns the machine learning model M based on training data. In this embodiment, the learning unit 104 successively generates training data based on each of a plurality of training video advertisements and stores the training data in the training database DB1. The learning unit 104 learns the machine learning model M based on all or part of the plurality of training data stored in the training database DB1. A portion of the training data may be used for verifying the learned machine learning model M. Note that the learning unit 104 may generate training data based on only one training video advertisement instead of multiple training video advertisements.
[0075] Learning is a process of adjusting the parameter part of the machine learning model M. The learning method may be a known method adopted in known machine learning. For example, the learning unit 104 learns the machine learning model M based on a learning method such as backpropagation or gradient descent so that an output part of training data is output when an input part of training data is input. The learning unit 104 calculates a loss based on a known loss function and performs learning until the loss becomes small to a certain extent. The learning unit 104 may complete learning when learning based on a certain number of training data is performed.
[0076] For example, when training video information is acquired by executing natural language processing on a training text, similar data is input to the machine learning model M at the time of estimation. For this reason, the learning unit 104 learns the machine learning model M that estimates the advertising effect from the estimated video information acquired by executing processing similar to the natural language processing on the training text on the estimated text acquired by analyzing the audio of the estimated video advertisement.
[0077] The estimated text is a text generated from the audio of the estimated video advertisement. In this embodiment, the process at the time of estimation is executed by the estimation device 20, and therefore the generation of the estimated text is executed by the estimation device 20. The method of generating the estimated text may be the same as the method of generating the training text. For details of the method of generating the estimated text, "training" in the description of the method of generating the training text can be read as "estimation". For natural language processing of the estimated text, "training" in the description of the natural language processing of the training text can be read as "estimation".
[0078] For example, when training video information is acquired by analyzing a video of a training video advertisement, at the time of estimation, similar data is input to the machine learning model M. For this reason, the learning unit 104 learns the machine learning model M that estimates the advertising effect from estimated video information related to facial expressions of the estimated performers who are performers of the estimated video advertisement, which is acquired by analyzing the video of the estimated video advertisement.
[0079] The facial expression of the estimated performer is an expression identified by an analysis of the video of the estimated video advertisement. In this embodiment, the process at the time of estimation is performed by the estimation device 20, and therefore the analysis of the facial expression of the estimated performer is performed by the estimation device 20. The method of analyzing the facial expression of the estimated performer may be similar to the method of analyzing the facial expression of the training performer. For details of the method of analyzing the facial expression of the estimated performer, "training" in the explanation of the method of analyzing the facial expression of the training performer can be read as "estimation".
[0080] For example, when at least one of the information on the presence or absence of a purchase by an estimated viewer who is a viewer of an estimated video advertisement and the information on the sales of a product or service introduced in the estimated video advertisement is acquired as the training purchase information, the learning unit 104 learns a machine learning model M that estimates the at least one of the information as the advertising effect. The learning unit 104 acquires an output portion of the training data based on the at least one of the information. The learning method based on the training data is as described above.
[0081] [3-2. Functions realized by the estimation device] For example, the estimation device 20 includes a data storage unit 200, a model storage unit 201, an estimated moving image information acquisition unit 202, and an estimation unit 203. The data storage unit 200 and the model storage unit 201 are realized by the storage unit 22. The estimated moving image information acquisition unit 202 and the estimation unit 203 are realized by the control unit 21.
[0082] [Data storage section] The data storage unit 200 stores data necessary for estimation based on the machine learning model M. For example, the data storage unit 200 stores an estimation database DB2.
[0083] 6 is a diagram showing an example of the estimated database DB2. The estimated database DB2 is a database that stores data related to estimated video advertisements for causing the machine learning model M to estimate advertising effectiveness. In this embodiment, an example is given in which video data of each of a plurality of estimated video advertisements is stored in the estimated database DB2, but the estimated database DB2 may store estimated video information generated by the estimated video information acquisition unit 202 from each of the plurality of estimated video advertisements. The estimated database DB2 may store data indicating advertising effectiveness estimated for the estimated video advertisement. It is assumed that the video data also includes audio information.
[0084] For example, the estimation device 20 downloads all or part of the video data of the video advertisement stored in a video advertisement database DB3 described below from the server 30 as video data of the estimated video advertisement. The estimation device 20 stores the downloaded video data in the estimation database DB2. The estimation device 20 may obtain the video data of the estimated video advertisement from a computer or information storage medium other than the server 30.
[0085] The data storage unit 200 can store any data. The data stored in the data storage unit 200 is not limited to the estimated database DB2. For example, the data storage unit 200 may store data associated with the video data of the estimated video advertisement (for example, a category of a product or service introduced in the estimated video advertisement, a description input in advance, coupon information, or a combination thereof, etc.).
[0086] For example, the data storage unit 200 may store a program for natural language processing. The data storage unit 200 may store a dictionary database in which keywords used in natural language processing are stored. The data storage unit 200 may store a machine learning model M for natural language processing. The data storage unit 200 may store a program for image analysis processing. The data storage unit 200 may store a machine learning model M for image analysis processing.
[0087] [Model memory section] The model storage unit 201 stores a machine learning model M that has been trained based on training video information acquired by analyzing a training video advertisement that is a video advertisement for training, and training purchase information regarding purchases of training viewers who are viewers of the training video advertisement. The estimation device 20 acquires the machine learning model M that has been trained by the learning unit 104 from the learning device 10. The estimation device 20 records the machine learning model M acquired from the learning device 10 in the model storage unit 201.
[0088] [Estimated video information acquisition section] The estimated video information acquisition unit 202 acquires estimated video information acquired by analyzing an estimated video advertisement, which is a video advertisement for estimation. For example, the estimated video information acquisition unit 202 acquires estimated video information of all or a part of an estimated video advertisement whose video data is stored in the estimated database DB2. The estimated video advertisement from which the estimated video information acquisition unit 202 acquires estimated video information may be specified by a person who operates the estimation device 20, or may be determined by a predetermined method.
[0089] In this embodiment, a case where a video advertisement that has been distributed in the past corresponds to the estimated video advertisement is taken as an example, but a video advertisement that has not been distributed may also correspond to the estimated video advertisement. In addition, a case where a video of a live distribution service corresponds to the estimated video advertisement is taken as an example, but the estimated video advertisement may be any video advertisement that is the subject of estimation of advertising effectiveness, and is not limited to a video of a live distribution service. For example, a web advertisement, a television advertisement, an outdoor advertisement, a movie advertisement, or other advertisements may correspond to the estimated video advertisement.
[0090] The analysis of the estimated video advertisement is a process of extracting features related to the estimated video advertisement. For example, voice analysis processing, natural language processing, image analysis processing, or a combination of these corresponds to the analysis of the estimated video advertisement. The estimated video information indicates the features extracted by the analysis of the estimated video advertisement. The estimated video information acquisition unit 202 may acquire the estimated video information based on the same acquisition method as the training video information. Therefore, for details of the process by which the estimated video information acquisition unit 202 acquires the estimated video information, "training" in the explanation of the training video information acquisition unit 102 can be read as "estimation."
[0091] For example, the estimated video information acquisition unit 202 may acquire estimated video information acquired by performing a process similar to the natural language processing performed on the training text on estimated text acquired by analyzing the audio of the estimated video advertisement. The process may be at least one of the above-mentioned emotion analysis process, response detection process, explanation detection process, and promotion detection process. In these details, too, "training" in the explanation of the training video information acquisition unit 102 can be read as "estimation". As with the natural language processing on the training text, any natural language processing may be performed on the estimated text, and the natural language processing may be performed by a computer other than the estimation device 20.
[0092] For example, the estimated video information acquisition unit 202 may acquire estimated video information related to facial expressions of estimated performers who are performers of the estimated video advertisement, acquired by analyzing the video of the estimated video advertisement. The video of the estimated video advertisement includes a plurality of frames (still images). The estimated video information acquisition unit 202 acquires the estimated video information by executing image analysis processing on each frame. This is similar to the image analysis processing on the video of the training video advertisement in that any image analysis processing may be executed on the video of the estimated video advertisement and that the image analysis processing may be executed by a computer other than the estimation device 20.
[0093] [Estimation part] The estimation unit 203 estimates the advertising effect of the estimated video advertisement based on the estimated video information and the machine learning model M. For example, the estimation unit 203 inputs the estimated video information to the machine learning model M. The estimation unit 203 may perform some processing such as normalization or aggregation processing, rather than inputting the estimated video information as is, and then input the estimated video information on which the processing has been performed to the machine learning model M. The machine learning model M calculates the feature amount of the estimated video information input thereto based on the parameter portion adjusted by learning, and outputs an estimation result according to the feature amount. The internal processing of the machine learning model M may be processing adopted in a known machine learning method.
[0094] In this embodiment, an example is taken in which the output portion of the training data is training purchase information indicating the proportion of training viewers who have made a purchase among the training viewers who have viewed the training video advertisement, and therefore the estimation result output from the machine learning model M is similar information. For example, the machine learning model M outputs estimated purchase information indicating the proportion of estimated viewers who have viewed the estimated video advertisement and are estimated to make a purchase (the proportion of estimated viewers expected to make a purchase, i.e., the probability of purchase) as the estimation result. The estimation unit 203 acquires the estimation result output from the machine learning model M. The estimation unit 203 may store the estimation result of the machine learning model M in the estimation database DB2.
[0095] In addition, when the training purchase information is the other information described above, the estimated purchase information is the same as the other information. For example, the machine learning model M may output, as an estimation result, information indicating whether or not each estimated viewer has made a purchase, the total number of estimated viewers estimated to make a purchase, the total number of viewers per unit time, or a combination thereof. As another example, the machine learning model M may output, as an estimation result, information regarding estimated sales for a product or service introduced in the estimated video advertisement. The information regarding sales may be the estimated number of sales (sales volume), sales amount, number of sales per unit time, sales amount per unit time, or a combination thereof, estimated for the product or service.
[0096] [3-3. Functions realized by the server] For example, the server 30 includes a data storage unit 300 and a distribution unit 301. The data storage unit 300 is realized by the storage unit 32. The distribution unit 301 is realized by the control unit 31.
[0097] [Data storage section] The data storage unit 300 stores data necessary for the live distribution service. For example, the data storage unit 300 stores a video advertisement database DB3.
[0098] 7 is a diagram showing an example of the video advertisement database DB3. The video advertisement database DB3 is a database in which data related to video advertisements in a live distribution service is stored. In this embodiment, the video advertisement data stored in the video advertisement database DB3 can be at least one of a training video advertisement and an estimated video advertisement. The video advertisement data stored in the video advertisement database DB3 may be used as both a training video advertisement and an estimated video advertisement.
[0099] For example, the video advertisement database DB3 stores a video advertisement ID capable of identifying a video advertisement, performer data on performers, product / service data on a product or service introduced in the video advertisement, coupon data on a coupon for the product or service, video data of the video advertisement, and comment data on a viewer's comment. Not only the video advertisement currently being distributed, but also an archive of video advertisements distributed in the past may be stored in the video advertisement database DB3. In addition, for example, data on reaction stamps may be stored in the video advertisement database DB3. The reaction stamp is a function that allows a viewer to react by selecting at least one of a plurality of stamps (e.g., hard or thumbs up) at any timing while watching the video advertisement. The reaction stamp is an example of a user's response (reaction).
[0100] For example, when a performer performs an operation to start distributing a new video advertisement, the distribution unit 301 issues a video advertisement ID for the video advertisement. The distribution unit 301 stores data such as performer data in the video advertisement database DB3 in association with the video advertisement ID. The performer data may indicate any information about the performer, for example, the performer's account, name, sex, age, other characteristics, or a combination of these. The product / service data may indicate any information about the product or service, for example, the product or service category, the product or service title, the product or service description, a link to a detailed page of the product or service, or a combination of these.
[0101] The coupon data may indicate any information related to a coupon for a product or service, such as a discount amount, a discount rate, a total number of coupons, coupon application conditions, coupon acquisition conditions, coupon acquisition status, or a combination thereof. The video data is generated by the distribution unit 301 recording a video advertisement. The video data may be in any data format. The video data may indicate only video and not include audio. The video data may include basic information of the video advertisement, such as the playback time of the video.
[0102] The comment data is data indicating the content of the comment. For example, the comment data indicates the account of the viewer who posted the comment, text indicating the content of the comment, the posting date and time of the comment, or a combination of these. The viewer may make a reaction other than a comment. In this case, data indicating the viewer's reaction is stored in the video advertisement database DB3. The video advertisement database DB3 may also store data regarding sales of a product or service introduced in a video advertisement.
[0103] The data storage unit 300 can store any data. The data stored in the data storage unit 300 is not limited to the video advertisement database DB3. For example, the data storage unit 300 may store data indicating the purchasing status of a product or service introduced in each video advertisement. The data may indicate the account of a viewer who purchased the product or service, the number of sales of the product or service, the sales amount, the number of sales per unit time, the sales amount per unit time, or a combination of these. These data may be stored in the video advertisement database DB3. The video advertisement database DB3 may store data on viewers who viewed the video advertisement, regardless of whether or not a comment was posted. Based on the data, information such as how many viewers viewed the video advertisement and how long each viewer viewed the video advertisement may be analyzed.
[0104] [Distribution Department] The distribution unit 301 distributes video advertisements to viewers. The processing executed by the distribution unit 301 may be similar to processing employed in known live distribution services. The distribution unit 301 distributes video advertisements to each of the training viewer device 50 and the estimated viewer device 70. The distribution unit 301 accepts input of comments from each of the training viewer device 50 and the estimated viewer device 70. The distribution unit 301 updates the video advertisement database DB3 based on the content of communication with each of the training performer device 40, the training viewer device 50, the estimated performer device 60, and the estimated viewer device 70.
[0105] [4. Processing Executed by the Learning Device and the Estimation Device] Fig. 8 is a diagram showing an example of processing executed by study device 10. The processing in Fig. 8 is executed by control unit 11 operating in accordance with a program stored in storage unit 12. When the processing in Fig. 8 is executed, it is assumed that study device 10 has downloaded video data of a training video advertisement from server 30 in advance. As shown in Fig. 8, study device 10 acquires training video information by analyzing the training video advertisement (S10).
[0106] FIG. 9 is a diagram showing an example of the processing executed in S10. The learning device 10 extracts the audio of the training video advertisement (S100). The learning device 10 converts the audio into text to obtain the training text (S101). The learning device 10 performs natural language processing on the training text (S102). The learning device 10 analyzes the video of the training video advertisement and analyzes the facial expressions of the training performers (S103). The learning device 10 obtains training video information based on each analysis result (S104). Details of each process of S100 to S104 are as described as the processing of the training video information obtaining unit 102.
[0107] Returning to FIG. 8, the learning device 10 acquires training purchase information (S11). Details of the processing of S11 are as described as the processing of the training purchase information acquisition unit 103. It is assumed that the learning device 10 has downloaded data necessary for acquiring the training purchase information from the server 30 in advance. The learning device 10 learns the machine learning model M based on the training video information and the training purchase information (S12). Details of the processing of S12 are as described as the processing of the learning unit 104. The learning device 10 transmits the trained machine learning model M to the estimation device 20 (S13), and the processing of FIG. 8 ends.
[0108] Fig. 10 is a diagram showing an example of processing executed by the estimation device 20. The processing in Fig. 10 is executed by the control unit 21 operating in accordance with a program stored in the storage unit 22. It is assumed that, when the processing in Fig. 10 is executed, the estimation device 20 has downloaded video data of the estimated video advertisement from the server 30 in advance. As shown in Fig. 10, the estimation device 20 acquires estimated video information by analyzing the estimated video advertisement (S20).
[0109] FIG. 11 is a diagram showing an example of the process executed in S20. The estimation device 20 extracts the sound of the estimated video advertisement (S200). The estimation device 20 converts the sound into text to obtain the estimated text (S201). The estimation device 20 executes natural language processing on the estimated text (S202). The estimation device 20 analyzes the video of the estimated video advertisement and analyzes the facial expressions of the estimated performers (S203). The estimation device 20 obtains estimated video information based on each analysis result (S204). Details of each process of S200 to S204 are as described as the process of the estimated video information obtaining unit 202.
[0110] Returning to Fig. 10, the estimation device 20 performs estimation using the machine learning model M based on the estimated video information and the machine learning model M (S21), and the processing in Fig. 10 ends. Details of the processing in S21 are as described above as the processing of the estimation unit 203.
[0111] [5. Summary of the embodiment] The learning device 10 of the present embodiment learns the machine learning model M based on the training video information and the training purchase information. The learning device 10 allows the machine learning model M to learn the training video information and the purchases of the training viewers, so that the machine learning model M can estimate the advertising effectiveness from the characteristics of the estimated video advertisement, and therefore the machine learning model M with high estimation accuracy of the advertising effectiveness can be created. For example, the learning device 10 allows the machine learning model M to learn the training video information having a causal relationship with the advertising effectiveness, so that the machine learning model M can learn the characteristics of the training video advertisement itself. As a result, the machine learning model M can estimate the advertising effectiveness from the characteristics of the estimated video advertisement itself, and therefore the estimation accuracy of the advertising effectiveness is improved.
[0112] The learning device 10 also acquires training video information acquired by executing natural language processing on the training text. The learning device 10 trains a machine learning model M that estimates advertising effectiveness from estimated video information acquired by executing processing similar to natural language processing on the estimated text. The learning device 10 trains the machine learning model M on the training video information analyzed by natural language processing, so that the machine learning model M can estimate advertising effectiveness from the characteristics of estimated text in which the audio of the estimated video advertisement has been converted into text. For this reason, the learning device 10 can create a machine learning model M with high estimation accuracy of advertising effectiveness. For example, the learning device 10 converts audio into text, making it easier to perform analysis based on various characteristics.
[0113] Furthermore, the natural language processing is at least one of a sentiment analysis process, a response detection process, an explanation detection process, and a promotion detection process. For example, the learning device 10 can cause the machine learning model M to learn the causal relationship between the emotions of the training performer and the purchases of the training viewers through the sentiment analysis process. The learning device 10 can cause the machine learning model M to learn the causal relationship between the responses of the training performer and the purchases of the training viewers through the response detection process. The learning device 10 can cause the machine learning model M to learn the causal relationship between the explanation by the training performer and the purchases of the training viewers through the explanation detection process. The learning device 10 can cause the machine learning model M to learn the causal relationship between the promotion in the training video advertisement and the purchases of the training viewers through the promotion detection process. As a result, the learning device 10 can create a machine learning model M with high estimation accuracy of the advertising effect.
[0114] The promotion detection process is a keyword detection process that detects at least one of keywords related to the price of a product or service introduced in the training video advertisement, keywords related to sales of the product or service introduced in the training video advertisement, and keywords related to a description associated with the training video advertisement. The learning device 10 can detect promotions that can be effectively used in estimating the advertising effectiveness using these keywords. As a result, the learning device 10 can improve the estimation accuracy of the machine learning model M that estimates the advertising effectiveness from the promotion in the video advertisement. For example, the learning device 10 can obtain a wider range of information by detecting keywords such as a coupon acquisition rate as keywords related to price and analyzing parameters other than words of the product or service itself introduced in the video advertisement.
[0115] Furthermore, the learning device 10 acquires training video information related to facial expressions of the training performers who are performers of the training video advertisement, which is acquired by analyzing the video of the training video advertisement. The learning device 10 learns a machine learning model M that estimates the advertising effectiveness from the estimated video information related to the facial expressions of the estimated performers who are performers of the estimated video advertisement, which is acquired by analyzing the video of the estimated video advertisement. The learning device 10 can cause the machine learning model M to learn the causal relationship between the facial expressions of the training performers and the purchases of the training viewers. As a result, the learning device 10 can create a machine learning model M with high estimation accuracy of the advertising effectiveness. For example, the learning device 10 can analyze individual features of the training performers that cannot be analyzed by natural language processing.
[0116] Further, the learning device 10 acquires training purchase information related to at least one of whether or not a training viewer has made a purchase and information on the sales of a product or service introduced in the training video advertisement. The learning device 10 learns a machine learning model M that estimates at least one of whether or not a purchase has been made by an estimated viewer who is a viewer of the estimated video advertisement and information on the sales of a product or service introduced in the estimated video advertisement as an advertising effect. The learning device 10 can cause the machine learning model M to learn the causal relationship between the training video information and the at least one of them. As a result, the learning device 10 can create a machine learning model M with high estimation accuracy of the advertising effect. For example, since the unit price differs depending on the product or service, if an analysis is performed using only the parameter of the sales amount, a product with a high unit price may be advantageous. In this regard, the learning device 10 can create a more accurate machine learning model M by analyzing various features such as the sales number.
[0117] The estimation device 20 of the present embodiment estimates the advertising effectiveness of the estimated video advertisement based on the estimated video information and the machine learning model M. The estimation device 20 can estimate the advertising effectiveness from the characteristics of the estimated video advertisement itself, thereby improving the estimation accuracy of the advertising effectiveness.
[0118] [6. Modifications] The present disclosure is not limited to the above-described embodiment, and may be modified as appropriate without departing from the spirit and scope of the present disclosure.
[0119] Fig. 12 is a diagram showing an example of functions realized in a modified example. As shown in Fig. 12, in the modified example described below, the learning device 10 includes a contribution degree calculation unit 105 and a training viewer information acquisition unit 106. The contribution degree calculation unit 105 and the training viewer information acquisition unit 106 are realized by the control unit 11. The estimation device 20 includes an estimated viewer information acquisition unit 204. The estimated viewer information acquisition unit 204 is realized by the control unit 21.
[0120] [6-1. Variation 1] For example, as described in the embodiment, the training video information acquisition unit 102 may acquire multiple pieces of training video information. The learning unit 104 may further learn the machine learning model M based on the multiple pieces of training video information. These processes are as described in the embodiment. When multiple pieces of training video information are learned by the machine learning model M, some pieces of training video information contribute strongly to the advertising effect, while other pieces of training video information do not contribute much to the advertising effect. Therefore, the contribution of each piece of training video information to the advertising effect may be calculated.
[0121] The learning device 10 of the first modification includes a contribution degree calculation unit 105. The contribution degree calculation unit 105 calculates the contribution degree of each of the multiple training video information to the advertising effect. The contribution degree is the degree to which the training video information contributes to the advertising effect. In other words, the contribution degree is the degree to which the training video information contributes to the estimation of the machine learning model M. The contribution degree may be called importance, score, impact, or other names. In the first modification, the contribution degree is expressed by a numerical value between 0 and 100 as an example, but the contribution degree may be expressed by a numerical value in another range, or may be expressed by a character or a symbol instead of a numerical value. In the first modification, a high numerical value indicated by the contribution degree means that the training video information strongly contributes to the advertising effect.
[0122] The method of calculating the contribution degree may be a known method known in the field of machine learning. For example, the contribution degree calculation unit 105 calculates the contribution degree of each of the multiple training video information based on a calculation method such as random forest, gradient boosting, SHAP (SHapley Additive exPlanations), LIME (Local Interpretable Model-agnostic Explanations), L1 normalization, or L2 normalization. For example, when the estimation result of the machine learning model M changes significantly when the value of a certain training video information changes, the contribution degree calculation unit 105 calculates the contribution degree of the training video information so that the contribution degree of the training video information becomes higher. For example, the contribution degree calculation unit 105 calculates the contribution degree of the training video information so that the contribution degree of the training video information becomes higher as the increase in the numerical value indicated by the estimation result of the machine learning model M relative to the increase in the numerical value of the certain training video information increases.
[0123] The contribution degree calculation unit 105 records the contribution degree of each of the multiple training video information in the data storage unit 100. The contribution degree may be used for any purpose. For example, the learning device 10 may display the contribution degree of each of the multiple training video information on the display unit 15. The learning device 10 may output data indicating the contribution degree of each of the multiple training video information to another computer or an information storage medium. The learning unit 104 may select at least one training video information having a relatively high contribution degree from among the multiple training video information, and may learn a new machine learning model M based only on the selected at least one training video information.
[0124] The learning device 10 of the first modification calculates the contribution degree of each of the multiple training video information to the advertising effect. The learning device 10 can identify the training video information that contributes to the advertising effect based on the contribution degree of each of the multiple training video information. For example, the learning device 10 can present which of the multiple training video information contributes to the advertising effect to a person who creates the machine learning model M, a performer, or other person. For example, if a performer creates a future video advertisement with reference to the contribution degree, the performer can create a video advertisement with a higher advertising effect. For example, the learning device 10 can identify which element of the training video information used in learning the machine learning model M is effective in the advertising effect based on the contribution degree, and create a machine learning model M that can estimate the advertising effect with higher accuracy.
[0125] [6-2. Variation 2] For example, the machine learning model M may learn not only the characteristics of the training video advertisement but also the characteristics of the training viewers. The learning device 10 of the second modification includes a training viewer information acquisition unit 106. The training viewer information acquisition unit 106 acquires training viewer information related to the training viewers. For example, the training viewer information acquisition unit 106 acquires the training viewer information from the server 30. The training viewer information acquisition unit 106 may generate the training viewer information based on data acquired from the server 30. The training viewer information acquisition unit 106 may generate the training viewer information based on data acquired from the training viewer device 50. The training viewer information acquisition unit 106 only needs to acquire training viewer information of at least one training viewer, and may acquire training viewer information of each of a plurality of training viewers.
[0126] The training viewer information is information indicating some characteristic related to the training viewer. For example, the training viewer information is behavior by the training viewer, demographic information of the training viewer, or a combination of these. For example, the training viewer information acquisition unit 106 may acquire training viewer information related to at least one of the age of the training viewer, the gender of the training viewer, a purchase history in a store by the training viewer, a training comment input by the training viewer for a training video advertisement, acquisition of a coupon related to the training video advertisement by the training viewer, a viewing status of the training video advertisement by the training viewer, an access status by the training viewer, and a search by the training viewer.
[0127] Demographic information of the training viewer, such as the age and gender of the training viewer, may be stored in the data storage unit 300 of the server 30, or may be input by the training viewer when viewing the training video advertisement. The training viewer information acquisition unit 106 may acquire the training viewer information by acquiring the age and gender stored in advance in the data storage unit 300 of the server 30, or by acquiring the age and gender input by the training viewer.
[0128] The purchase history of the training viewer at a store is a history of online or offline purchases at the store. The purchase history can also be referred to as a history of the training viewer's use of a store. For example, the purchase history is information about the store used by the training viewer, information about products or services purchased by the training viewer, the date and time when the training viewer used the store, or a combination of these. The purchase history of the training viewer at a store may be stored in the data storage unit 300 of the server 30, or may be stored in another computer or information storage medium. The training viewer information acquisition unit 106 may acquire the training viewer information by acquiring the purchase history of the training viewer at a store from the data storage unit 300, another computer, or an information storage medium.
[0129] For example, comment data indicating training comments is stored in the video advertisement database DB3. The training viewer information acquisition unit 106 acquires training viewer information by acquiring the training comments from the video advertisement database DB3. The training viewer information acquisition unit 106 may perform natural language processing on the training comments and acquire the training viewer information based on the execution result of the natural language processing. The natural language processing may be similar to the processing on the training text described in the embodiment. For example, the training viewer information acquisition unit 106 may perform natural language processing such as sentiment analysis processing or response detection processing on the training comments.
[0130] For example, coupon data indicating the acquisition status of coupons related to training video advertisements by training viewers is stored in the video advertisement database DB3. The training viewer information acquisition unit 106 acquires the training viewer information by acquiring the coupon acquisition status from the video advertisement database DB3. The coupon acquisition status may be the total number of coupons acquired by the training viewer, the total number per unit time, a time-series transition of the total number, or a combination thereof.
[0131] For example, the viewing status of the training video advertisement by the training viewer is stored in the video advertisement database DB3. The training viewer information acquisition unit 106 acquires the training viewer information by acquiring the viewing status from the video advertisement database DB3. The viewing status of the training video advertisement may be any information related to the viewing by the training viewer, and may be, for example, the viewing time by the training viewer (the time during the playback time of the video advertisement that the training viewer actually viewed), the ratio of the viewing time to the playback time, the time period during which the training viewer viewed the video advertisement, or a combination of these.
[0132] For example, the access status of the training viewers is stored in the video advertisement database DB3. The training viewer information acquisition unit 106 acquires the training viewer information by acquiring the access status from the video advertisement database DB3. The access status may be the access status for any page, for example, the access status for a page of a live distribution service, or the access status for an electronic commerce page. The access status may be the presence or absence of access to a certain page, the number of floors accessed, the frequency of access, the time period of access, or a combination of these.
[0133] For example, search results by training viewers are stored in the video advertisement database DB3. The training viewer information acquisition unit 106 acquires training viewer information by acquiring access status from the video advertisement database DB3. The search by the training viewer may be a search in a live distribution service, an electronic commerce transaction, or another service. The search by the training viewer may be a query entered at the time of the search, a time period when the search was performed, a user's behavior after the search, or a combination thereof.
[0134] The training viewer information may be any characteristic related to the training viewer, and is not limited to the above examples. For example, the training viewer information acquisition unit 106 may acquire, as the training viewer information, the elapsed time since the training viewer started using the live streaming service, the product or service purchased by the training viewer via the live streaming service, the number of video advertisements viewed by the training viewer via the live streaming service, the annual income of the training viewer, the preferences of the training viewer, or other information. The training viewer information acquisition unit 106 may acquire multiple pieces of training viewer information.
[0135] The learning unit 104 of the second modification learns the machine learning model M that estimates the advertising effectiveness by further using estimated viewer information on the estimated viewer who is a viewer of the estimated video advertisement, further based on the training viewer information. For example, the learning unit 104 learns the machine learning model M that estimates the advertising effectiveness by further using estimated viewer information on at least one of the age of the estimated viewer, the sex of the estimated viewer, the purchase history at a store by the estimated viewer, the estimated comment input by the estimated viewer for the estimated video advertisement, the acquisition of a coupon by the estimated viewer for the estimated video advertisement, the viewing status of the estimated video advertisement by the estimated viewer, the access status by the estimated viewer, and the search by the estimated viewer. The learning unit 104 may learn the machine learning model M in the same manner when other estimated viewer information is acquired.
[0136] For example, the learning unit 104 generates an input portion of the training data based on the training video information and the training viewer information. The learning unit 104 may not use the training viewer information as it is as the input portion of the training data, but may perform some processing, such as normalization or aggregation, on the training viewer information and then use the training viewer information as the input portion of the training data. For example, the learning unit 104 may aggregate the training video information of each of a plurality of training viewers and use the aggregation result as the input portion of the training data. The learning unit 104 may aggregate the gender ratio or number of training viewers who viewed the training video advertisement and use the ratio or number of training viewers as the input portion of the training data.
[0137] For another example, the learning unit 104 may compile an age distribution of training viewers who viewed a training video advertisement, and use the distribution as an input part of the training data. The learning unit 104 may compile a ratio or number of training viewers who purchased a specific product, and use the ratio or number as an input part of the training data. Similarly, for other training viewer information, the learning unit 104 may execute a compilation process or the like and then use the information as an input part of the training data. Although the input part of the training data is different from that of the embodiment, the process in which the learning unit 104 causes the machine learning model M to learn the training data is as described in the embodiment.
[0138] The learning device 10 of the second modification learns a machine learning model M that estimates the advertising effectiveness by further utilizing estimated viewer information on an estimated viewer who is a viewer of an estimated video advertisement, further based on the training viewer information. The learning device 10 can create a machine learning model M with high estimation accuracy of the advertising effectiveness by having the machine learning model M learn not only the training video information but also the training viewer information. For example, when there is a causal relationship between the characteristics of the training viewer and the advertising effectiveness (for example, when the advertising is effective for a specific age group), the machine learning model M can estimate the advertising effectiveness according to the characteristics of the training viewer. For example, the learning device 10 can widen the scope of learning by having the machine learning model M learn the characteristics of individual training viewers that cannot be learned from training purchase information alone.
[0139] The learning device 10 also acquires training viewer information related to at least one of the following: the age of the training viewer, the gender of the training viewer, the purchase history of the training viewer at a store, the training comment input by the training viewer for the training video advertisement, the acquisition of a coupon related to the training video advertisement by the training viewer, the viewing status of the training video advertisement by the training viewer, the access status by the training viewer, and the search by the training viewer. The learning device 10 further uses the estimated viewer information related to the at least one of the items to learn a machine learning model M that estimates the advertising effectiveness. The learning device 10 can create a machine learning model M with high estimation accuracy of the advertising effectiveness by having the machine learning model M learn not only the training video information but also the at least one of the items. For example, when there is a causal relationship between the at least one of the items and the advertising effectiveness, the machine learning model M can estimate the advertising effectiveness according to the characteristics of the training viewer.
[0140] [6-3. Variation 3] For example, in variant example 2, the training viewer information acquisition unit 106 may acquire training viewer information regarding the training comments by executing at least one of an emotion analysis process for analyzing the emotions of the training viewer, a comment classification process for classifying the training comments, and a keyword detection process for detecting keywords from the training comments.
[0141] The emotion analysis process may be the same as the emotion analysis process for the training text described in the embodiment. The training viewer information acquisition unit 106 executes emotion analysis process for the training comments to analyze emotions such as joy, anger, sadness, and happiness of the training viewers, and acquires training viewer information indicating the analysis results. In the description of the emotion analysis process for the training text in the embodiment, "training text" may be read as "training comments", and "training performers" may be read as "training viewers". Similar to the emotion analysis process for the training text, various emotion analysis processes may be executed for the training comments.
[0142] The comment classification process is a process for classifying the contents of the training comments. The comment classification process can utilize a method used in classifying the meaning of words in natural language processing. For example, the training viewer information acquisition unit 106 may classify the training comments based on a rule approach that determines whether or not a keyword corresponding to each of a plurality of classifications is included in the training comment. As another example, the training viewer information acquisition unit 106 may classify the training comments based on a machine learning approach that calculates the feature amount of the training comment and outputs a classification result. For example, in the comment classification process, the training comments are classified as either positive or negative. The classification of the training comments may be other classifications.
[0143] The keyword detection process is a process for detecting keywords included in the training comments. The training viewer information acquisition unit 106 determines whether or not a keyword stored in a dictionary database storing keywords related to prices (e.g., "cheap" or "expensive") is included in the training comments, and detects the keyword. The training viewer information acquisition unit 106 determines whether or not a keyword stored in a dictionary database storing keywords related to opinions on products or services (e.g., "cute" or "cool") is included in the training comments, and detects the keyword.
[0144] The learning unit 104 performs learning of a machine learning model M that estimates advertising effectiveness by further utilizing estimated viewer information regarding the estimated comment acquired by performing at least one of a sentiment analysis process for analyzing the sentiment of the estimated viewer, a comment classification process for classifying the estimated comment, and a keyword detection process for detecting a keyword from the estimated comment. The learning unit 104 may tally up the execution results of the sentiment analysis process, the comment classification process, and the keyword detection process, and generate an input portion of the training data based on the tallying result (e.g., 100 people who made positive comments, 50 people who made negative comments, etc.). This point is also as described in the second modification.
[0145] In addition, in the details of the processing for estimated comments, "training" in the description of the processing for training comments can be read as "estimation." In the details of the natural language processing for estimated comments, "training" in the description of the natural language processing for training comments can be read as "estimation." Since the processing at the time of estimation is performed by the estimation device 20, the processing for estimated comments is performed by the estimation device 20.
[0146] The learning device 10 of the third modification acquires training viewer information regarding the training comments by performing at least one of a sentiment analysis process, a comment classification process, and a keyword detection process on the training comments. The learning device 10 further utilizes the estimated viewer information acquired by performing at least one of a sentiment analysis process, a comment classification process, and a keyword detection process on the estimated comments to learn a machine learning model M that estimates the advertising effectiveness. The learning device 10 can create a machine learning model M with high estimation accuracy of the advertising effectiveness by having the machine learning model M learn the analysis results of at least one of the sentiment analysis process, the comment classification process, and the keyword detection process. For example, the learning device 10 can further improve the estimation accuracy of the machine learning model M by performing an analysis based on information regarding the reactions of the training viewers.
[0147] [6-4. Variation 4] For example, as described in the embodiment, the training video information acquisition unit 102 may acquire a plurality of pieces of training video information. As described in Modification 2 and Modification 3, the training viewer information acquisition unit 106 may acquire a plurality of pieces of training viewer information. In this case, the learning unit 104 learns a machine learning model M that estimates advertising effectiveness from a plurality of pieces of estimated video information and a plurality of pieces of estimated viewer information, based on the plurality of pieces of training video information and the plurality of pieces of training viewer information. These processes are as described in the embodiment, Modification 2, and Modification 3.
[0148] The learning device 10 of the fourth modification includes a contribution degree calculation unit 105. The contribution degree calculation unit 105 calculates the contribution degree of each of the plurality of training video information and each of the plurality of training viewer information to the advertising effect. Among the functions of the contribution degree calculation unit 105, the configuration for calculating the contribution degree of each of the plurality of training video information may be the same as that of the first modification. The contribution degree calculation unit 105 of the fourth modification calculates not only the contribution degree of each of the plurality of training video information, but also the contribution degree of each of the plurality of training viewer information. That is, the contribution degree calculation unit 105 calculates the contribution degree of each of the plurality of training viewer information, such as the emotions of the training viewer and the training comments of the training viewer, in order to analyze which training viewer information contributes to the advertising effect.
[0149] The meaning of the contribution degree of the training viewer information is the same as that of the contribution degree of the training video information. Regarding the meaning of the contribution degree of the training viewer information, "video" in the description of the contribution degree of the training video information in the first modification can be read as "viewer". The contribution degree calculation unit 105 may calculate the contribution degree of the training viewer information by a calculation method similar to that of the calculation method of the contribution degree of the training video information. Regarding the calculation method of the contribution degree of the training viewer information, "video" in the description of the calculation method of the contribution degree of the training video information in the first modification can be read as "viewer". The contribution degree calculation unit 105 in the fourth modification records the contribution degree of each of the multiple pieces of training video information and each of the multiple pieces of training viewer information in the data storage unit 100. As in the first modification, the contribution degree may be used for any purpose.
[0150] The learning device 10 of the fourth modification calculates the contribution degree of each of the plurality of training video information and each of the plurality of training viewer information to the advertising effect. The learning device 10 can identify the training video information and the training viewer information that contribute to the advertising effect by the contribution degree of each of the plurality of training video information and each of the plurality of training viewer information. For example, the learning device 10 can present which of the plurality of training video information and the plurality of training viewer information contributes to the advertising effect to a person who creates the machine learning model M, a performer, or other person. For example, if a performer creates a future video advertisement with reference to the contribution degree, the performer can create a video advertisement with a higher advertising effect. For example, the learning device 10 can identify which element of the training video information and the training viewer information used in learning the machine learning model M works effectively for the advertising effect by the contribution degree, and create a machine learning model M that can estimate the advertising effect with higher accuracy.
[0151] [6-5. Variation 5] For example, in variant example 4, the learning unit 104 may select at least one of the multiple training video information and the multiple training viewer information based on the contribution degree of each of the multiple training video information and the contribution degree of each of the multiple training viewer information, and learn a new machine learning model M based on the selected at least one.
[0152] The learning unit 104 selects at least one training video information having a relatively high contribution degree from among the multiple training video information. For example, the learning unit 104 selects at least one training video information having a contribution degree equal to or greater than a threshold from among the multiple training video information. The learning unit 104 may select a predetermined number of training video information from among the multiple training video information in descending order of contribution degree. The learning unit 104 uses only the selected at least one training video information from among the multiple training video information in learning the new machine learning model M.
[0153] The learning unit 104 selects at least one piece of training viewer information having a relatively high contribution degree from among the multiple pieces of training viewer information. For example, the learning unit 104 selects at least one piece of training viewer information having a contribution degree equal to or greater than a threshold from among the multiple pieces of training viewer information. The learning unit 104 may select a predetermined number of pieces of training viewer information from among the multiple pieces of training viewer information in descending order of contribution degree. The learning unit 104 uses only the selected at least one piece of training viewer information from among the multiple pieces of training viewer information in learning the new machine learning model M.
[0154] For example, the learning unit 104 generates an input portion of training data based on the selected at least one piece of training video information and the selected at least one piece of training viewer information. The learning unit 104 learns a new machine learning model M based on the training data. Although the training data is different, the learning method itself may be the same as that of the old machine learning model M (the machine learning model M created temporarily to calculate the contribution degree). The old machine learning model M here is the machine learning model M created by the method described in the embodiment and the first to fourth modifications.
[0155] The learning device 10 of the fifth modification selects at least one of the plurality of training video information and the plurality of training viewer information based on the contribution degree of each of the plurality of training video information and the contribution degree of each of the plurality of training viewer information, and learns a new machine learning model M based on the selected at least one. This allows the learning device 10 to reduce the amount of information input to the machine learning model M at the time of estimation, and thus generates a machine learning model M that can reduce the processing load at the time of estimation. Furthermore, training video information and training viewer information with a relatively low contribution degree are not learned by the machine learning model M, and therefore the accuracy of the machine learning model M is improved.
[0156] [6-6. Variation 6] For example, the estimation device 20 may perform estimation based on the machine learning model M generated by the method described in Modification 2 to Modification 5. The machine learning model M in Modification 6 is further trained based on training viewer information on the training viewer. The machine learning model M is as described in Modifications 2 to 5.
[0157] The estimation device of the sixth modification includes an estimated viewer information acquisition unit 204. The estimated viewer information acquisition unit 204 acquires estimated viewer information related to estimated viewers who are viewers of the estimated video advertisement. The estimated viewer information acquisition unit 204 may acquire the estimated viewer information based on the same acquisition method as that for the training viewer information. Therefore, in the details of the process in which the estimated viewer information acquisition unit 204 acquires the estimated viewer information, "training" in the explanation of the training viewer information acquisition unit 106 can be read as "estimation."
[0158] For example, the estimated viewer information is behavior of the estimated viewer, demographic information of the estimated viewer, or a combination of these. These pieces of information may be stored in the data storage unit 300 of the server 30. The estimated viewer information acquisition unit 204 may acquire estimated viewer information related to at least one of the age of the estimated viewer, the gender of the estimated viewer, the purchase history of the estimated viewer in a store, an estimated comment input by the estimated viewer for the estimated video advertisement, acquisition of a coupon related to the estimated video advertisement by the estimated viewer, the viewing status of the estimated video advertisement by the estimated viewer, the access status by the estimated viewer, and a search by the estimated viewer.
[0159] For example, the estimated viewer information acquisition unit 204 may acquire estimated viewer information regarding an estimated comment by executing at least one of a sentiment analysis process for analyzing the sentiment of an estimated viewer, a comment classification process for classifying the estimated comment, and a keyword detection process for detecting a keyword from the estimated comment. The estimated viewer information acquisition unit 204 may acquire multiple pieces of estimated viewer information.
[0160] The estimation unit 203 of the sixth modification estimates the advertising effectiveness further based on the estimated viewer information. The estimation unit 203 inputs not only the estimated video information but also the estimated viewer information to the machine learning model M. Although the input to the machine learning model M is different from that of the embodiment, the process executed by the machine learning model M is similar to that of the embodiment. The machine learning model M calculates a feature amount based on not only the estimated video information but also the estimated viewer information, and outputs an estimation result according to the feature amount. The estimation unit 203 may aggregate the estimated viewer information and then input the aggregation result to the machine learning model M.
[0161] The estimation device 20 of the sixth modification estimates the advertising effectiveness further based on the estimated viewer information. The estimation device 20 can improve the estimation accuracy of the advertising effectiveness by inputting not only the estimated video information but also the estimated viewer information to the machine learning model M. For example, when there is a causal relationship between the characteristics of the estimated viewer and the advertising effectiveness, the machine learning model M can estimate the advertising effectiveness according to the characteristics of the estimated viewer. For example, the estimation device 20 can improve the estimation accuracy by having the machine learning model M learn the characteristics of each estimated viewer that cannot be estimated based on the estimated purchase information alone.
[0162] [6-7. Other variations] For example, the above-described modifications may be combined.
[0163] For example, the learning device 10 and the estimation device 20 may be the same device. The functions described as being realized by the learning device 10 may be shared by multiple computers. The functions described as being realized by the estimation device 20 may be shared by multiple computers. The functions described as being realized by the server 30 may be realized by the learning device 10 or the estimation device 20.
[0164] [7. Notes] For example, the learning device may have the following configuration.
[0165] (1) a training video information acquisition unit that acquires training video information acquired by analyzing a training video advertisement that is a video advertisement for training; a training purchase information acquisition unit that acquires training purchase information regarding purchases of training viewers who are viewers of the training video advertisement; A learning unit that learns a machine learning model that estimates the advertising effectiveness of an estimated video advertisement from estimated video information acquired by analyzing the estimated video advertisement, which is a video advertisement for estimation, based on the training video information and the training purchase information; A learning device comprising: (2) The training video information acquisition unit acquires the training video information acquired by performing natural language processing on a training text acquired by analyzing a voice of the training video advertisement, The learning unit learns the machine learning model that estimates the advertising effectiveness from the estimated video information obtained by performing a process similar to the natural language processing on the estimated text obtained by analyzing the audio of the estimated video advertisement. A learning device as described in (1). (3) The natural language processing is at least one of a sentiment analysis process for analyzing the sentiment of a training performer who is a performer in the training video advertisement, a response detection process for detecting a response by a training performer who is a performer in the training video advertisement, an explanation detection process for detecting an explanation regarding a product or service introduced in the training video advertisement, and a promotion detection process for detecting a promotion related to the training video advertisement. A learning device as described in (2). (4) The promotion detection process is a keyword detection process that detects at least one of a keyword related to a price of a product or service introduced in the training video advertisement, a keyword related to sales of the product or service introduced in the training video advertisement, and a keyword related to a description associated with the training video advertisement. A learning device as described in (3). (5) The training video information acquisition unit acquires the training video information related to facial expressions of training performers who are performers of the training video advertisement, the training video information being acquired by analyzing a video of the training video advertisement; The learning unit learns the machine learning model that estimates the advertising effectiveness from the estimated video information related to facial expressions of the estimated performers who are performers of the estimated video advertisement, the facial expressions being acquired by analyzing the video of the estimated video advertisement. A learning device according to any one of (1) to (4). (6) The training purchase information acquisition unit acquires the training purchase information regarding at least one of the presence or absence of a purchase by the training viewer and information regarding sales of a product or service introduced in the training video advertisement, The learning unit learns the machine learning model that estimates, as the advertising effect, at least one of the presence or absence of a purchase by an estimated viewer who is a viewer of the estimated video advertisement and information regarding sales of a product or service introduced in the estimated video advertisement. A learning device according to any one of (1) to (5). (7) The training video information acquisition unit acquires a plurality of pieces of training video information, The learning unit learns the machine learning model further based on the plurality of pieces of training video information, The learning device further includes a contribution degree calculation unit that calculates a contribution degree of each of the plurality of training video information to the advertising effect. A learning device according to any one of (1) to (6). (8) the learning device further includes a training viewer information acquisition unit that acquires training viewer information related to the training viewer, the learning unit performs learning of the machine learning model that estimates the advertising effectiveness by further utilizing estimated viewer information on an estimated viewer who is a viewer of the estimated video advertisement, based on the training viewer information; A learning device according to any one of (1) to (7). (9) the training viewer information acquisition unit acquires the training viewer information relating to at least one of the age of the training viewer, the sex of the training viewer, the purchase history of the training viewer at a store, the training comment input by the training viewer in response to the training video advertisement, the acquisition of a coupon by the training viewer in relation to the training video advertisement, the viewing status of the training video advertisement by the training viewer, the access status by the training viewer, and a search by the training viewer; the learning unit learns the machine learning model that estimates the advertising effectiveness by further utilizing the estimated viewer information related to at least one of the age of the estimated viewer, the sex of the estimated viewer, the purchase history of the estimated viewer in a store, the estimated comment input by the estimated viewer for the estimated video advertisement, the acquisition of a coupon related to the estimated video advertisement by the estimated viewer, the viewing status of the estimated video advertisement by the estimated viewer, the access status by the estimated viewer, and the search by the estimated viewer; A learning device as described in (8). (10) the training viewer information acquisition unit acquires the training viewer information regarding the training comments by executing at least one of a sentiment analysis process for analyzing sentiments of the training viewer, a comment classification process for classifying the training comments, and a keyword detection process for detecting keywords from the training comments, the learning unit performs learning of the machine learning model that estimates the advertising effect by further utilizing the estimated viewer information regarding the estimated comment acquired by performing at least one of a sentiment analysis process that analyzes the sentiment of the estimated viewer, a comment classification process that classifies the estimated comment, and a keyword detection process that detects a keyword from the estimated comment, on the estimated comment; A learning device as described in (9). (11) The training video information acquisition unit acquires a plurality of pieces of training video information, the training viewer information acquisition unit acquires a plurality of pieces of training viewer information, The learning unit learns the machine learning model that estimates the advertising effectiveness from the plurality of pieces of estimated video information and the plurality of pieces of estimated viewer information based on the plurality of pieces of training video information and the plurality of pieces of training viewer information; The learning device further includes a contribution degree calculation unit that calculates a contribution degree of each of the plurality of training video information and each of the plurality of training viewer information to the advertising effect. A learning device according to any one of (8) to (10). (12) the learning unit selects at least one of the plurality of training video information and the plurality of training viewer information based on the contribution degree of each of the plurality of training video information and the contribution degree of each of the plurality of training viewer information, and performs learning of a new machine learning model based on the selected at least one. A learning device as described in (11).
Claims
1. a training video information acquisition unit that acquires training video information acquired by performing natural language processing on a training text acquired by analyzing a voice of a training video advertisement that is a video advertisement for training; a training purchase information acquisition unit that acquires training purchase information regarding purchases of training viewers who are viewers of the training video advertisement; A learning unit that learns a machine learning model that estimates the advertising effectiveness of the estimated video advertisement from estimated video information obtained by performing a process similar to the natural language processing on estimated text obtained by analyzing the audio of the estimated video advertisement, which is a video advertisement for estimation, based on the training video information and the training purchase information; Including, The natural language processing is at least one of a sentiment analysis process for analyzing the sentiment of a training performer who is a performer in the training video advertisement, a response detection process for detecting a response by a training performer who is a performer in the training video advertisement, an explanation detection process for detecting an explanation regarding a product or service introduced in the training video advertisement, and a promotion detection process for detecting a promotion related to the training video advertisement. Learning device.
2. The promotion detection process is a keyword detection process that detects at least one of a keyword related to a price of a product or service introduced in the training video advertisement, a keyword related to sales of the product or service introduced in the training video advertisement, and a keyword related to a description associated with the training video advertisement. The learning device according to claim 1 .
3. The training video information acquisition unit acquires the training video information related to facial expressions of training performers who are performers of the training video advertisement, the training video information being acquired by analyzing a video of the training video advertisement; The learning unit learns the machine learning model that estimates the advertising effectiveness from the estimated video information related to facial expressions of the estimated performers who are performers of the estimated video advertisement, the facial expressions being acquired by analyzing the video of the estimated video advertisement. The learning device according to claim 1 or 2.
4. The training purchase information acquisition unit acquires the training purchase information related to information regarding sales of a product or service introduced in the training video advertisement, The learning unit learns the machine learning model that estimates information regarding sales of a product or service introduced in the estimated video advertisement as the advertising effect. The learning device according to claim 1 or 2.
5. The training video information acquisition unit acquires a plurality of pieces of training video information, The learning unit learns the machine learning model further based on the plurality of pieces of training video information, The learning device further includes a contribution degree calculation unit that calculates a contribution degree of each of the plurality of training video information to the advertising effect. The learning device according to claim 1 or 2.
6. A training video information acquisition unit that acquires training video information acquired by analyzing a training video advertisement, which is a video advertisement for training; a training purchase information acquisition unit that acquires training purchase information regarding purchases of training viewers who are viewers of the training video advertisement; a training viewer information acquisition unit that acquires the training viewer information regarding the training comments by executing at least one of a sentiment analysis process that analyzes the sentiment of the training viewer, a comment classification process that classifies the training comments, and a keyword detection process that detects keywords from the training comments, for the training comments input by the training viewer in response to the training video advertisement; a learning unit that performs learning of a machine learning model that estimates the advertising effectiveness of the estimated video advertisement using estimated viewer information regarding the estimated comments acquired by performing at least one of a sentiment analysis process that analyzes the sentiment of the estimated viewer, a comment classification process that classifies the estimated comment, and a keyword detection process that detects a keyword from the estimated comment, for estimated comments input by an estimated viewer who is a viewer of the estimated video advertisement, which is a video advertisement for estimation, based on the training video information, the training purchase information, and the training viewer information; A learning device comprising:
7. The training video information acquisition unit acquires a plurality of pieces of training video information, the training viewer information acquisition unit acquires a plurality of pieces of training viewer information, The learning unit learns the machine learning model that estimates the advertising effectiveness from the plurality of pieces of estimated video information and the plurality of pieces of estimated viewer information based on the plurality of pieces of training video information and the plurality of pieces of training viewer information; The learning device further includes a contribution degree calculation unit that calculates a contribution degree of each of the plurality of training video information and each of the plurality of training viewer information to the advertising effect. The learning device according to claim 6.
8. the learning unit selects at least one of the plurality of pieces of training video information and the plurality of pieces of training viewer information based on the contribution degree of each of the plurality of pieces of training video information and the contribution degree of each of the plurality of pieces of training viewer information, and performs learning of a new machine learning model based on the selected at least one piece of training video information and the plurality of pieces of training viewer information; The learning device according to claim 7.
9. an estimated video information acquisition unit that acquires estimated video information by performing a process similar to a predetermined natural language process on an estimated text acquired by analyzing a sound of the estimated video advertisement, which is a video advertisement for estimation; a model storage unit that stores a machine learning model that has been trained based on training video information acquired by executing the natural language processing on a training text acquired by analyzing the audio of a training video advertisement, which is a video advertisement for training, and training purchase information regarding purchases by a training viewer who is a viewer of the training video advertisement; an estimation unit that estimates an advertising effect of the estimated video advertisement based on the estimated video information and the machine learning model; Including, The natural language processing is at least one of a sentiment analysis process for analyzing the sentiment of a training performer who is a performer in the training video advertisement, a response detection process for detecting a response by a training performer who is a performer in the training video advertisement, an explanation detection process for detecting an explanation regarding a product or service introduced in the training video advertisement, and a promotion detection process for detecting a promotion related to the training video advertisement. Estimation device.
10. An estimated video information acquisition unit that acquires estimated video information acquired by analyzing an estimated video advertisement that is a video advertisement for estimation; an estimated viewer information acquisition unit that acquires estimated viewer information regarding an estimated comment, the estimated viewer being a viewer of the estimated video advertisement, by executing at least one of a sentiment analysis process that analyzes the sentiment of the estimated viewer, a comment classification process that classifies the estimated comment, and a keyword detection process that detects a keyword from the estimated comment; a model storage unit that stores a machine learning model that estimates the advertising effectiveness of the estimated video advertisement using the estimated viewer information, the machine learning model being learned based on training video information acquired by analyzing a training video advertisement, which is a video advertisement for training, training purchase information regarding purchases of training viewers who are viewers of the training video advertisement, and training viewer information regarding the training viewers; an estimation unit that estimates an advertising effect of the estimated video advertisement based on the estimated video information, the estimated viewer information, and the machine learning model; Including, The training viewer information is information about the training comments, which is acquired by executing at least one of a sentiment analysis process for analyzing the sentiment of the training viewer, a comment classification process for classifying the training comments, and a keyword detection process for detecting keywords from the training comments, for the training comments input by the training viewer in response to the training video advertisement. Estimation device.
11. A computer comprising: A training video information acquisition step of acquiring training video information acquired by performing natural language processing on a training text acquired by analyzing a voice of a training video advertisement, which is a video advertisement for training; a training purchase information acquisition step of acquiring training purchase information regarding purchases of training viewers who are viewers of the training video advertisement; A learning step of learning a machine learning model that estimates the advertising effectiveness of the estimated video advertisement from estimated video information obtained by performing a process similar to the natural language processing on estimated text obtained by analyzing an estimated video advertisement, which is a video advertisement for estimation, based on the training video information and the training purchase information; Run The natural language processing is at least one of a sentiment analysis process for analyzing the sentiment of a training performer who is a performer in the training video advertisement, a response detection process for detecting a response by a training performer who is a performer in the training video advertisement, an explanation detection process for detecting an explanation regarding a product or service introduced in the training video advertisement, and a promotion detection process for detecting a promotion related to the training video advertisement. How to learn.
12. A computer comprising: an estimated video information acquisition step of acquiring estimated video information by performing a process similar to a predetermined natural language process on an estimated text acquired by analyzing a voice of the estimated video advertisement, which is a video advertisement for estimation; an estimation step of estimating the advertising effectiveness of the estimated video advertisement based on a machine learning model that has been trained based on training video information obtained by performing the natural language processing on a training text obtained by analyzing the audio of a training video advertisement, which is a video advertisement for training, and training purchase information regarding purchases of training viewers who are viewers of the training video advertisement, and the estimated video information; Run The natural language processing is at least one of a sentiment analysis process for analyzing the sentiment of a training performer who is a performer in the training video advertisement, a response detection process for detecting a response by a training performer who is a performer in the training video advertisement, an explanation detection process for detecting an explanation regarding a product or service introduced in the training video advertisement, and a promotion detection process for detecting a promotion related to the training video advertisement. Estimation method.
13. a training video information acquisition unit that acquires training video information acquired by performing natural language processing on a training text acquired by analyzing a voice of a training video advertisement that is a video advertisement for training; a training purchase information acquisition unit for acquiring training purchase information regarding purchases of training viewers who are viewers of the training video advertisement; a learning unit that learns a machine learning model that estimates the advertising effectiveness of the estimated video advertisement from estimated video information obtained by performing a process similar to the natural language processing on estimated text obtained by analyzing the audio of the estimated video advertisement, which is a video advertisement for estimation, based on the training video information and the training purchase information; The computer functions as The natural language processing is at least one of a sentiment analysis process for analyzing the sentiment of a training performer who is a performer in the training video advertisement, a response detection process for detecting a response by a training performer who is a performer in the training video advertisement, an explanation detection process for detecting an explanation regarding a product or service introduced in the training video advertisement, and a promotion detection process for detecting a promotion related to the training video advertisement. program.
14. an estimated video information acquisition unit that acquires estimated video information by performing a process similar to a predetermined natural language process on an estimated text acquired by analyzing the audio of the estimated video advertisement, which is a video advertisement for estimation; an estimation unit that estimates the advertising effectiveness of the estimated video advertisement based on a machine learning model that has been trained based on training video information acquired by executing the natural language processing on a training text acquired by analyzing the audio of a training video advertisement, which is a video advertisement for training, and training purchase information regarding purchases of a training viewer who is a viewer of the training video advertisement, and the estimated video information; The computer functions as The natural language processing is at least one of a sentiment analysis process for analyzing the sentiment of a training performer who is a performer in the training video advertisement, a response detection process for detecting a response by a training performer who is a performer in the training video advertisement, an explanation detection process for detecting an explanation regarding a product or service introduced in the training video advertisement, and a promotion detection process for detecting a promotion related to the training video advertisement. program.
Citation Information
Patent Citations
Viewing material evaluation method, viewing material evaluation system, and program
JP2017129923A
Ranking information providing method, program, and ranking information providing system
JP2017220235A
Program, apparatus and method capable of selecting product to be presented, and product generation program
JP2020042317A
TV program evaluation system
JP2021064945A
Deep neural networks modeling
US20190311268A1
Cited By
Behavior prediction device and behavior prediction method
JP7920479B1