Video content evaluation system, video content evaluation method, and program
Patent Information
- Application Number
- JP2025026120
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2026-09-01
AI Technical Summary
【0015】 本発明によれば、視聴者の振る舞いを計測しない場合でも、動画コンテンツが視聴者に与える感情を推定可能とする仕組みを提供することができる。 上記した以外の課題、構成及び効果は、以下の実施形態の説明により明らかにされる。
Smart Images

Figure 2026139424000001_ABST
Abstract
Description
[Technical Field]
[0001] This invention relates to a video content evaluation system, a video content evaluation method, and a program. [Background technology]
[0002] In the field of advertising creative, there is a need to evaluate advertising media. From the perspective of consumer UX, it is important to appropriately provide emotions such as surprise in response to advertising elements when watching video commercials. From the perspective of creators, being able to objectively evaluate the emotions viewers feel in prototypes and past video commercials in advance can lead to a reduction in production time and an improvement in quality.
[0003] Patent documents 1 to 7 disclose the background technology of this field.
[0004] Patent Document 1 states that the aim is to "improve the accuracy and versatility of an AI model for emotion estimation" (see abstract). Specifically, it states that "the learning method for an emotion estimation AI involves acquiring the subject's biosignals and appearance information at the same time, calculating an emotion index value based on the acquired biosignals, generating supervised learning data using the acquired appearance information as input values and the emotion index value as the correct answer, and training the AI model with the generated supervised learning data" (see abstract).
[0005] Patent Document 2 describes "more accurately grasping the physiological responses of the viewer and more accurately evaluating the ability of the object or the viewer themselves in various situations" (see abstract). Specifically, it describes "a viewer emotion determination device 100 that includes an analysis unit 102 that calculates the position of the viewpoint in the visual image a from the visual image a and the eye movement image b, and calculates the changes in the viewer's physiological response data and the acceleration that accompanies them; an accumulation unit 103 that stores information including the changes in physiological response data and the acceleration that accompanies them when the viewer is looking at an object of high interest as an emotion value K, and stores the information including the changes in physiological response data and the acceleration that accompanies them calculated by the analysis unit 102 as an analysis value A; and compares the analysis value A with the emotion value K. On the other hand, it includes a diagnosis unit 104 that analyzes the viewer's feelings towards a specific object by a face analysis method that calculates the changes in each part of the viewer's face, and diagnoses both emotion and feeling together." (see abstract).
[0006] Patent Document 3 describes "directly estimating brain activity from expressed information and evaluating emotions and sensibilities based on the estimation results" (see abstract). Specifically, it describes that "the sensibility evaluation system 100 comprises a feature extraction unit 11 that extracts feature quantities from the user's expressed information, a neural network 12 that receives feature quantities as input and estimates and outputs the user's brain activity state when the expressed information is expressed, a brain physiological index value calculation unit 13 that calculates at least one brain physiological index value related to sensibility from the brain activity state output from the neural network 12, and a sensibility evaluation value calculation unit 14 that calculates the user's sensibility evaluation value by substituting the brain physiological index value into a predetermined formula" (see abstract).
[0007] Patent Document 4 describes "a data processing device and a data processing method capable of performing data processing based on human emotions" (see abstract). Specifically, it describes that "the server device 20 includes a dictionary acquisition means 233 that generates an emotion classification dictionary in which multiple words are classified into emotion units, and a quantification means 234 that acquires the emotion classification dictionary and generates quantified data that quantifies human emotions towards content. This makes it possible to perform various data processing of content based on emotions" (see abstract).
[0008] Patent Document 5 describes "a recommendation device that estimates a user's emotions from information related to the user's viewing of content and enables recommendations appropriate to these emotions" (see abstract). Specifically, it describes "the recommendation device having means for generating a viewing history for each viewing content, associating it with previously viewed content that has already been viewed by statistical target users who viewed the content only for the viewing time within that time segment, and to generate an emotion tag for each viewing content based on the emotion information attached to the viewing content and to attach it to the viewing content, and to generate recommendation information based on previously viewed content that has been attached to an emotion tag equivalent to or similar to the emotion tag attached to the viewing content viewed by the recommended user, as associated with the time segment of the viewing content viewed by the recommended user" (see abstract).
[0009] Patent Document 6 states that its purpose is "to provide an information processing device, a video distribution method, and a video distribution program that change the display manner of a video on a mobile terminal without using the sound of the video" (see abstract). Specifically, it states that "the server performs a focus video determination process to determine which of several dance videos displayed simultaneously on the screen of the mobile terminal has the largest dancer movements, and distributes the focus video to the mobile terminal so that it is displayed larger than the other dance videos. The focus video determination process extracts images showing the dancer from the video at predetermined time intervals, and determines which of the several videos has the largest difference in the extracted subject to be the focus video" (see abstract).
[0010] Patent Document 7 describes "evaluating viewing materials objectively and quantitatively" (see abstract). Specifically, it describes "a method for evaluating viewing materials, which includes a brain activity measurement step S102 in which a brain activity measurement unit measures the brain activity of a subject who has viewed the viewing material; a first matrix generation step S103 in which a first matrix generation unit generates a first matrix that estimates the semantic content perceived by the subject based on the measurement results measured by the brain activity measurement step S102; a second matrix generation step S104 in which a second matrix generation unit performs natural language processing on textual information indicating the planning intent of the viewing material and generates a second matrix; and a similarity calculation step S105 in which a similarity calculation unit calculates the similarity between the first matrix and the second matrix" (see abstract). [Prior art documents] [Patent Documents]
[0011] [Patent Document 1] Japanese Patent Publication No. 2024-141286 [Patent Document 2] International Publication No. 2011 / 042989 [Patent Document 3] Japanese Patent Publication No. 2022-062574 [Patent Document 4] Japanese Patent Publication No. 2015-121858 [Patent Document 5] Japanese Patent Application Laid-Open No. 2015-228142 Patent Document 6 Japanese Patent Application Laid-Open No. 2023-052125 Patent Document 7 Japanese Patent Application Laid-Open No. 2017-129923 Summary of the Invention Problems to be Solved by the Invention
[0012] The emotion evaluation methods described in the above Patent Documents 1 to 7 rely on subjective evaluation and physiological indicators, and it is difficult to evaluate the emotion that unknown video content evokes in viewers.
[0013] The present invention has been made in view of such circumstances, and provides a mechanism that enables estimation of the emotion that video content evokes in viewers even when the viewer's behavior is not measured. Means for Solving the Problems
[0014] In order to solve the above problems, for example, the configurations described in the claims are adopted. The present application includes a plurality of means for solving the above problems. To list one example, the present invention is a video content evaluation system including: an acquisition unit that acquires video content; an extraction unit that inputs the acquired video content to a content understanding model and extracts time-series content representations; a first estimation unit that inputs the extracted time-series content representations to a communication content estimation model and estimates first time-series communication content; a second estimation unit that estimates the emotion of a viewer of the video content based on the estimated first time-series communication content; and an output unit that performs output based on the estimated emotion. Effects of the Invention
[0015] According to the present invention, it is possible to provide a mechanism that enables estimation of the emotion that video content evokes in viewers even when the viewer's behavior is not measured. Problems, configurations and effects other than those described above will be clarified by the following description of embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] [Figure 1] Figure 1 shows an example of the configuration of a video CM evaluation apparatus 100. [Figure 2] Figure 2 shows an example of a first time-series advertising expression. [Figure 3] Figure 3 shows an example of a second time-series advertising expression. [Figure 4] Figure 4 shows an example of a graphical model. [Figure 5] Figure 5 shows an example of a third time-series advertising expression. [Figure 6] Figure 6 shows an example of a fourth time-series advertising expression. [Figure 7] Figure 7 shows an example of emotion estimation processing 700. [Figure 8] Figure 8 shows an example of advertising expression extraction processing 800. [Figure 9] Figure 9 shows an example of transmission content estimation processing 900. [Figure 10] Figure 10 shows an example of emotion determination processing 1000. [Figure 11] Figure 11 shows an example of arousal level determination processing 1100. [Figure 12] Figure 12 shows an example of learning degree calculation processing 1200. MODE FOR CARRYING OUT THE INVENTION
[0017] 1. Example Hereinafter, examples of the present invention will be described with reference to the drawings. 1-1. Configuration Figure 1 is a diagram showing an example of the configuration of a video CM evaluation apparatus 100. The video CM evaluation apparatus 100 shown in the figure estimates emotion values of a viewer of a video CM based on changes in time-series advertising expressions included in the video CM and changes in estimated time-series transmission content.
[0018] The video commercial evaluation device 100 is composed of, for example, one or more servers located on the cloud. The video commercial evaluation device 100 is an example of a video content evaluation system according to the present invention.
[0019] This video commercial evaluation device 100 includes a main memory 101 such as RAM (Random Access Memory), an auxiliary storage device 102 such as an IC card, hard disk drive, SSD (Solid State Drive), or flash memory, a processor 103 that runs an operating system, applications, programs, etc., an input device 104 such as a touch panel, keyboard, mouse, voice input, or motion detection input from a camera, an output device 105 such as a monitor or display, and a communication control unit 106 such as a network card, wireless communication module, or mobile communication module. Note that the output device 105 may be a device or terminal that transmits information for output to an external monitor, display, printer, or other device.
[0020] Of these, the main memory 101 stores various programs and applications (modules), and the processor 103 executes these programs and applications to realize each functional element of the video commercial evaluation device 100. Note that each module may be implemented in hardware by integration or other means. Furthermore, each module may be an independent program or application, or it may be implemented as a subprogram or function within a single integrated program or application.
[0021] In this specification, each module is described as the entity (subject) that performs the processing; however, in reality, the processor that processes various programs and applications (modules) executes the processing.
[0022] The auxiliary storage device 102 stores various databases (DBs). A "database" is a functional element (storage unit) that stores a data set so that it can handle any data operations (e.g., extraction, addition, deletion, overwriting, etc.) from the processor or an external computer. The implementation method of a database is not limited; for example, it may be a database management system, spreadsheet software, or text files such as XML or JSON.
[0023] The main memory 101 specifically stores programs for the extraction module 110, the first estimation module 111, the second estimation module 112, and the output module 113. The processor 103 executes these programs to realize each functional element of the video commercial evaluation device 100. Each module will be described below.
[0024] The extraction module 110 acquires video content, inputs the acquired video content into a content understanding model, and extracts the time-series content representation. In this context, "video content" specifically refers to video commercials (or, in other words, video advertisements). A content understanding model is a pre-trained model that takes video content as input and outputs a time-series representation of that content. This content understanding model consists of an audio understanding model M1 and an image understanding model M2. The speech understanding model M1 is a pre-trained model that takes audio segments from a video as input and outputs a content representation of the first time series. On the other hand, the image understanding model M2 is a pre-trained model that takes scenes from a video (in other words, still images in time series) as input and outputs a content representation of the second time series.
[0025] Here, the first time-series content representation specifically refers to the first time-series advertising representation, which is information indicating the presence or absence of audio related to a product or trademark in each scene. On the other hand, the second time-series content representation specifically refers to the second time-series advertising representation, which is information indicating the presence or absence of images related to a product, person, company, or trademark in each scene.
[0026] In this context, advertising expression refers to elements expressed using image or audio information. For example, image-based advertising expression is binary information, indicating the presence or absence of a product image, characters, product description, logo, close-up product image, close-up characters, or a promotional message. Alternatively, image-based advertising expression is category information, indicating the type of action of the characters, the attributes of the characters, the emotional expression of the characters, or the impression of the product.
[0027] On the other hand, advertising expressions related to audio are binary information, indicating product description, brand name, sales promotion message, sound effects, or the presence or absence of background music.
[0028] Figure 2 shows an example of a first time-series advertising expression. The first time-series advertising expression 200 shown in the figure is in table format and has columns for audio section 201 and advertising expression 202. Of these, the column for audio section 201 consists of columns for number 203, start (time) 204, and end (time) 205. On the other hand, the column for advertising expression 202 consists of columns for sound effect 206, background music 207, brand name 208, product description 209, and purchase promotion message 210. Note that the column for advertising expression 202 stores a variable indicating whether or not an advertising expression is present.
[0029] Figure 3 shows an example of a second time-series advertising expression. The second time-series advertising expression 300 shown in the figure is in table format and has columns for Scene 301 and Advertising Expression 302. Of these, the Scene 301 column consists of Number 303, Start (Time) 304, and End (Time) 305. On the other hand, the Advertising Expression 302 column consists of Product UP (Close-up) 306, Character UP (Close-up) 307, Company Logo 308, Product Image 309, Product Description 310, Characters 311, Sales Promotion Message 312, and Character Actions 313. Note that the Advertising Expression 302 column (except for Character Actions 313) stores a variable indicating whether or not an advertising expression exists.
[0030] Next, we will describe the first estimated module 111. The first estimation module 111 inputs the time-series content representation extracted by the extraction module 110 into the content estimation model M3 to estimate the time-series content. Here, the content estimation model M3 is an advertising expression generation model that has been trained on data containing multiple time-series, scene-based advertising expressions. The content estimation model M3 can take time-series advertising expressions as input and estimate the next advertising expression, as well as the content of the preceding and succeeding expressions. For example, when using a hidden Markov model as the content estimation model M3, the graphical model is as shown in Figure 4, and the probability of generating an advertising expression is as shown in the following equation.
[0031]
number
[0032] Here, state s represents a combination of certain transmitted content. Observation o represents a combination of advertising expressions. The initial state vector D represents the probability distribution of the initial message transmitted in the scene. The state transition matrix B represents the probability distribution of the transitions in the transmitted content between time points. Observation matrix A represents the probability distribution of advertising expressions based on the content of the message.
[0033] The time-series information transmitted by the first estimation module 111 is information that indicates the attitude or behavior that the video content encourages viewers to exhibit, scene by scene. Specifically, it is information that indicates, for example, viewer interest (A), brand awareness (B), empathy (C), purchasing behavior (D), etc.
[0034] Figure 5 shows an example of a time-series advertising expression input to the content estimation model M3 by the first estimation module 111. The time-series advertising expression 500 shown in the figure (hereinafter referred to as the "third time-series advertising expression") is information generated by combining the first time-series advertising expression and the second time-series advertising expression described above.
[0035] This third time-series advertising expression 500 is in table format and has columns for video ID 501, scene 502, advertising expression element 503, and message content 504. Of these, the scene 502 column consists of columns for number 505, start (time) 506, and end (time) 507. The advertising expression element 503 column consists of columns such as product UP (close-up) 508 and character UP (close-up) 509. Note that the advertising expression element 503 column (excluding category information) stores a variable indicating the presence or absence of an advertising expression. In addition, the message content 504 column stores a variable representing the attitude and behavior encouraged in the viewer. In Figure 4, "A" represents viewer interest, "B" represents brand awareness, and "C" represents empathy.
[0036] The variables stored in the "Communication Content 504" column (in other words, the communication content) are assigned according to rules that associate advertising expressions with communication content. The table below shows an example of the correspondence between advertising expressions and communication content. [Table 1]
[0037] According to this table, scenes that include a close-up of a product or a character, and scenes that include sound effects or background music, are assigned the message content "A". Additionally, scenes containing product images, characters, product descriptions, or logos (all images), and scenes containing product descriptions or brand names (all audio) will be assigned the message content "B". Furthermore, scenes where the type of action, attribute, or emotional expression of a character, or the impression of a product (all represented by images), meet certain conditions will be assigned the message content "C". Additionally, scenes containing promotional messages (images or audio) are assigned the message content designation "D".
[0038] Figure 6 shows an example of a fourth time-series advertising expression generated by the first estimation module 111. The fourth time-series advertising expression 600 shown in the figure is in a table format and has columns for video ID 601, scene 602, advertising expression element 603, prior probability distribution of the message content 604, posterior probability distribution of the message content 605, and probability of the advertising expression 606. Of these, the scene 602 column consists of columns for number 607, start (time) 608, and end (time) 609. The advertising expression element 603 column consists of columns such as product UP (close-up) 610 and character UP (close-up) 611. Note that the advertising expression element 603 column (excluding category information) stores a variable indicating the presence or absence of an advertising expression. In addition, "A", "B", and "C" stored in the prior probability distribution 604 and posterior probability distribution 605 columns represent viewer interest, brand awareness, and empathy, respectively.
[0039] Next, we will explain the second estimated module 112. The second estimation module 112 estimates the emotions of viewers of video content based on the time-series transmission content estimated by the first estimation module 111. The estimated emotions are expressed as emotional values that viewers feel from the video commercial.
[0040] This second estimation module 112 consists of an arousal level calculation module 112A and a learning level calculation module 112B.
[0041] Of these, the arousal level calculation module 112A calculates a surprise value for each scene based on the time-series content representation extracted by the extraction module 110, the time-series content transmission estimated by the first estimation module 111, and the content transmission estimation model M3. Based on the calculated surprise value, it calculates the viewer's emotion (specifically, the arousal level). The arousal level calculation module 112A then calculates the viewer's emotion (specifically, the arousal level) based on the number of scenes in which the calculated surprise value is less than or equal to a first predetermined value, the number of scenes in which the calculated surprise value exceeds a second predetermined value, and the maximum, minimum, average, average of the first derivative, average of the second derivative, or a combination thereof of the calculated first surprise value. Note that the first predetermined value and the second predetermined value may be the same value.
[0042] Here, the level of arousal is a value that indicates the level of surprise a viewer feels after watching all the scenes of a video. The surprise value is a value that indicates the viewer's surprise at an advertisement at a given time. The first predetermined value is specifically the reference value (non-awakening level) for a non-awakened state. The second predetermined value is, specifically, the reference value (level of arousal) that represents the optimal state of arousal. The ratio is a value that indicates the relative number of awakening scenes compared to non-awakening scenes.
[0043] Meanwhile, the learning level calculation module 112B calculates the viewer's learning level by performing the following process. (1) Update the message content estimation model M3 to minimize the surprise value for each scene. (2) The time-series content representation extracted by the extraction module 110 is input into the updated transmission content estimation model M3 to estimate the time-series transmission content. (3) Calculate the surprise value for each scene based on the estimated time-series transmission content and the updated transmission content estimation model M3. (4) The viewer's learning level is calculated based on the change between the calculated surprise value and the surprise value calculated by the arousal level calculation module 112A.
[0044] The learning rate referred to here is a value that indicates the degree to which the sense of surprise diminishes when the same advertising expression is viewed repeatedly.
[0045] Here, we will explain how the Awakening Level Calculation Module 112A and the Learning Level Calculation Module 112B calculate the surprise value for each scene. When viewing an advertisement, the viewer is given a surprise value, which is the negative logarithm of the probability of the advertisement being displayed, -lnP(o τ ) occurs.
[0046] The learning progress calculation module 112B updates the message content estimation model M3 to minimize this surprise. Furthermore, according to Reference 1, surprise minimization can be replaced with minimizing free energy F, and its components can also be considered part of the surprise. [Reference 1] Friston, K: The free-energy principle: a unified brain theory? Nature Reviews, Neuroscience, 11(2), 127-138, 2010.
number
[0047] Contents of communication at a certain time τ The prior predicted value P(s) τ ) and advertising expression o τ Given the posterior predicted value Q(s τ ) is calculated, the difference is taken as the amount of information gained, and the remaining term -E_ Q(sτ) InP(o τ |s τ Let ) be the uncertainty. Therefore, the surprise value, which indicates the viewer's level of surprise for each scene, can be expressed by free energy, information gain, uncertainty, negative log-likelihood, or a combination thereof. According to reference 2, free energy is said to be related to a person's emotion of surprise. [Reference 2] Joffily, M., & Coricelli, G: Emotional valence and the free energy principle, PLoS Computational Biology, 9(6), e1003094, 2013.
[0048] Here, "surprise" refers to the negative logarithmic probability value of an advertising expression at a given time. Free energy is a substitute indicator of surprise, expressed in terms of the amount of information gained and uncertainty. Information acquisition is the difference between the message predicted at a given time and the message predicted after observing the advertising elements at that time. Uncertainty is the negative expected value of the probability value of an advertising expression element conditioned on the posterior predicted value of the message content at a given time.
[0049] Next, we will describe the output module 113. The output module 113 outputs based on the emotion estimated by the second estimation module 112. Specifically, the output module 113 outputs based on the emotion (specifically, the arousal level) calculated by the arousal level calculation module 112A. The output module 113 also outputs based on the learning level calculated by the learning level calculation module 112B. The output referred to here includes output to the output device 105 and transmission to other devices using the communication control unit 106.
[0050] 1-2.Operation Next, we will describe the emotion estimation process 700 performed by the video commercial evaluation device 100. Figure 7 is a flowchart showing an example of the emotion estimation process 700.
[0051] First, the extraction module 110 acquires video commercials (step 701). Then, the extraction module 110 extracts time-series advertising expressions from the acquired video commercials (step 702). At this time, the module executes the advertising expression extraction process 800. Figure 8 is a flowchart showing an example of the advertising expression extraction process 800.
[0052] First, the extraction module 110 detects audio segments within the video commercial (step 801). Then, the extraction module 110 extracts the first time-series advertising expression by analyzing the detected audio segments (step 802). In doing so, the module inputs each audio segment into the audio understanding model M1 to extract the first time-series advertising expression. An example of the extracted first time-series advertising expression is shown in Figure 2.
[0053] Next, the extraction module 110 divides the video commercial into scenes (step 803). Then, the extraction module 110 extracts a second time-series advertising expression by analyzing the images within each divided scene (step 804). At this time, the module inputs each scene into the image understanding model M2 to extract the second time-series advertising expression. An example of the extracted second time-series advertising expression is shown in Figure 3.
[0054] Finally, the extraction module 110 combines the extracted first time-series advertising expressions and the second time-series advertising expressions to generate a third time-series advertising expression, which is a scene-based advertising expression (step 805). An example of the generated third time-series advertising expression is shown in Figure 4. The above is an explanation of the advertising expression extraction process 800.
[0055] Once the third time-series advertising expression is generated, the first estimation module 111 then estimates the content of the time series based on the generated third time-series advertising expression (step 703 in Figure 7). At this time, the module executes the content estimation process 900. Figure 9 is a flowchart showing an example of the content estimation process 900.
[0056] First, the first estimation module 111 reads the communication content estimation model M3 (step 901). Next, the first estimation module 111 acquires the advertising expression of the t-th scene (with an initial value of "1") from among the advertising expressions of the third time series (step 902). Then, the first estimation module 111 calculates the prior probability distribution of the communication content of the t-th scene (step 903). Next, the first estimation module 111 calculates the posterior probability distribution of the communication content based on the advertising expression of the t-th scene (step 904). Next, the first estimation module 111 calculates the probability value of the advertising expression of the t-th scene (step 905). In the calculations of steps 903 to 905, the communication content estimation model M3 is used.
[0057] Here, the prior probability distribution, posterior probability distribution, and probability value will be described with reference to FIG. 4. S, which is a vector of probability values for each state at a given time τ τ prior probability distribution P(S τ ) is calculated by the following formula based on the estimated state at the previous time τ-1 and the transition matrix.
Mathematical Expression
[0058] Using the marginal message passing method, Q(S), which represents the posterior probability distribution of S, a vector of probability values for each state at a given time τ posterior probability distribution Q(S τ ) is based on the preceding and following estimated states S τ-1 and S τ+1 , transition matrix B τ B τ+1 and the current observation o, which is a 1-of-K vector τ is updated by the following formula. Note that σ is the softmax function.
Mathematical Expression
[0059] Observation probability value P(o τ ) and its negative log-likelihood are obtained from the current observation o τ the observation matrix A, and the prior probability distribution P(Sτ Based on this, it is calculated using the following formula. Note that S represents the total number of states.
number
[0060] For more information on the peripheral message passing method described above, please refer to Reference 1 below. [Reference 3] Parr, T., Markovic, D., Kiebel, SJ, & Friston, K. J: Neuronal message passing using Mean-field, Bethe, and Marginal approximations. Scientific Reports, 9(1), 1889, 2019.
[0061] Next, the first estimation module 111 increments the variable t (step 906) and determines whether the incremented variable t exceeds the threshold T (total number of scenes) (step 907). If the result of this determination is that the variable t does not exceed the threshold T (NO in step 907), the first estimation module 111 returns to step 902 and obtains the advertising expression for the next scene. On the other hand, if the result of this determination is that the variable t exceeds the threshold T (YES in step 907), the first estimation module 111 generates the advertising expression for the fourth time series (step 908). An example of the generated fourth time series advertising expression is shown in Figure 6. The above is an explanation of the message content estimation process 900.
[0062] Once the fourth time-series advertising expression is generated, the second estimation module 112 then estimates the viewer's emotions based on the generated fourth time-series advertising expression (step 704 in Figure 7). At this time, the module executes the emotion determination process 1000. Figure 10 is a flowchart showing an example of the emotion determination process 1000.
[0063] First, the arousal level calculation module 112A determines the viewer's arousal level (step 1001). At this time, the module executes the arousal level determination process 1100. Figure 11 is a flowchart showing an example of the arousal level determination process 1100.
[0064] First, the arousal level calculation module 112A calculates the surprise value for each scene (step 1101). Next, the arousal level calculation module 112A calculates the number of arousal scenes where the surprise value is equal to or greater than the arousal level (step 1102). Next, the arousal level calculation module 112A calculates the number of non-arousal scenes where the surprise value is less than the non-arousal level (step 1103). Next, the arousal level calculation module 112A calculates the ratio of arousal scenes to non-arousal scenes (in other words, the arousal ratio or arousal level) (step 1104). Finally, the arousal level calculation module 112A determines whether the number of arousal scenes is greater than "0" (step 1105).
[0065] If the result of this determination is "0" for the number of awakening scenes (NO in step 1105), the awakening level calculation module 112A determines that the awakening level is low (step 1107). On the other hand, if the result of this determination is that the number of awakening scenes is greater than "0" (YES in step 1105), the awakening level calculation module 112A then determines whether the awakening rate exceeds the rate level (step 1106).
[0066] If, as a result of this determination, the awakening rate does not exceed the rate level (NO in step 1106), the awakening level calculation module 112A determines that the awakening level is moderate (step 1108). On the other hand, if, as a result of this determination, the awakening rate exceeds the rate level (YES in step 1006), the awakening level calculation module 112A determines that the awakening level is high (step 1109). The above is an explanation of the awakening level determination process 1100.
[0067] Once the alertness level determination is complete, the alertness level calculation module 112A then determines whether the alertness level calculated in step 1104 exceeds a predetermined value (step 1002 in Figure 10). If the result of this determination is that the alertness level does not exceed the predetermined value (NO in step 1002), this process ends. On the other hand, if the result of this determination is that the alertness level exceeds the predetermined value (YES in step 1002), the learning level calculation module 112B then calculates the viewer's learning level (step 1003). At this time, the module executes the learning level calculation process 1200. Figure 12 is a flowchart showing an example of the learning level calculation process 1200.
[0068] First, the learning progress calculation module 112B determines whether the variable n, which represents the number of repetitions, is less than the threshold N (step 1201). If the result of this determination is that the variable n is less than the threshold N (YES in step 1201), the learning progress calculation module 112B then calculates the surprise value for each scene (step 1202). Next, the learning progress calculation module 112B updates the content estimation model M3 to minimize the surprise value calculated in step 1102 (step 1203). Next, the learning progress calculation module 112B increments the variable n (step 1204) and determines whether the incremented variable n is less than the threshold N (step 1201).
[0069] If, as a result of this determination, the variable n is not less than the threshold N (NO in step 1201), the learning rate calculation module 112B calculates the surprise reduction amount Delta_t, which is the difference between surprise S_N and S_1 for all scenes (step 1205). Here, surprise S_N is the surprise value for each scene calculated when the variable n is "N", and surprise S_1 is the surprise value for each scene calculated when the variable n is "1".
[0070] Next, the learning progress calculation module 112B calculates the sum of Delta_t (step 1206). The calculated sum is the viewer's learning progress. The above is an explanation of the learning progress calculation process 1200.
[0071] Once the learning level is calculated, the learning level calculation module 112B then determines whether the calculated learning level is greater than a predetermined value (step 1004 in Figure 10). If the result of this determination is that the calculated learning level is greater than the predetermined value (YES in step 1004), the learning level calculation module 112B determines that there was a momentary surprise (step 1005). On the other hand, if the result of this determination is that the calculated learning level is less than or equal to the predetermined value (NO in step 1004), the learning level calculation module 112B determines that there was a sustained surprise (step 1006). The above is an explanation of emotion determination process 1000.
[0072] Once the emotion estimation is complete, the output module 113 finally outputs the result of the arousal level determination process 1100 (step 705 in Figure 7). At that time, if the learning level calculation process 1200 has been executed, the output module 113 also outputs the result of step 1005 or 1006. The above is an explanation of emotion estimation processing 700.
[0073] According to the emotion estimation process 700 described above, it is possible to estimate the emotions that a video commercial evokes in viewers without measuring their behavior.
[0074] In addition, the emotion estimation process 700 described above outputs the results of the determination of arousal level and learning level (see step 705), but in addition to or instead of these determination results, the calculated arousal level and learning level themselves may also be output.
[0075] 2. Variations The above embodiment may be modified as follows. The following modifications may be combined with each other. (1) Subject to evaluation In the above embodiment, video commercials are used as the subject of evaluation, but other video content (e.g., movies) may also be used as the subject of evaluation. In that case, the content expressions extracted from the video content and the message conveyed inferred from the extracted content expressions may be appropriately changed depending on the video content being evaluated.
[0076] (2) Components of the system The video commercial evaluation device 100 may be a mobile device such as a smartphone, tablet, mobile phone, or personal digital assistant (PDA), or it may be a wearable device such as glasses, a wristwatch, or clothing. The device may also be a stationary or portable computer, or a server located on the cloud or a network. Functionally, the device may be a VR (Virtual Reality) terminal, an AR (Augmented Reality) terminal, or an MR (Mixed Reality) terminal. Alternatively, it may be a combination of multiple such terminals. For example, a combination of one smartphone and one wearable device can logically function as a single terminal. Other information processing terminals may also be used.
[0077] (3) Others It should be noted that the present invention is not limited to the embodiments described above, and various modifications are included. For example, the embodiments described above are described in detail to make the present invention easier to understand, and are not necessarily limited to those having all the configurations described. Furthermore, it is possible to replace parts of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add configurations from other embodiments to the configuration of one embodiment. In addition, it is possible to add, delete, or replace parts of the configuration of each embodiment with other configurations.
[0078] Furthermore, each of the above configurations, functions, processing units, and processing means may be implemented in hardware, either partially or entirely, by designing them as integrated circuits, for example. Alternatively, each of the above configurations and functions may be implemented in software by having the processor interpret and execute programs that implement each function. Information such as programs, tables, and files that implement each function can be stored in memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.
[0079] Furthermore, the control lines and information lines shown are those deemed necessary for explanatory purposes, and not all control lines and information lines are necessarily shown in the actual product. In reality, it can be assumed that almost all components are interconnected. Furthermore, the above-described embodiment discloses at least the configuration described in the claims. [Explanation of Symbols]
[0080] 100...Video commercial evaluation device, 110...Extraction module, 111...First estimation module, 112...Second estimation module, 112A...Arousal level calculation module, 112B...Learning level calculation module, 113...Output module,
Claims
1. The acquisition unit acquires video content, The acquired video content is input into a content understanding model, and an extraction unit extracts the time-series content representation. A first estimation unit inputs the extracted time-series content representation into a transmission content estimation model to estimate the transmission content of the first time series, A second estimation unit estimates the emotions of viewers of the video content based on the estimated first time series of transmitted content, An output unit that outputs based on the estimated emotion, A video content evaluation system equipped with the following features.
2. The second estimation unit calculates a first surprise value for each scene based on the extracted time-series content representation, the estimated first time-series communication content, and the communication content estimation model, and calculates the viewer's emotions based on the calculated first surprise value. The output unit outputs based on the calculated emotion. The video content evaluation system according to claim 1.
3. The video content evaluation system according to claim 2, wherein the second estimation unit calculates the viewer's emotions based on the number of scenes in which the calculated first surprise value is less than or equal to a first predetermined value, the number of scenes in which the calculated first surprise value exceeds a second predetermined value, and the maximum value, minimum value, average value, average of first derivatives, average of second derivatives, or a combination thereof of the calculated first surprise value.
4. The second estimation unit further includes a learning level calculation unit, The aforementioned learning level calculation unit, The aforementioned communication content estimation model is updated to minimize the surprise value for each scene. The extracted time-series content representation is input into the updated transmission content estimation model to estimate the transmission content of the second time-series. Based on the estimated second time-series transmission content and the updated transmission content estimation model, a second surprise value is calculated for each scene, and the viewer's learning level is calculated based on the amount of change between the first surprise value and the second surprise value. The output unit further outputs based on the calculated learning level. The video content evaluation system according to claim 2.
5. The video content evaluation system according to claim 1, wherein the aforementioned emotion includes the level of arousal.
6. The video content evaluation system according to claim 1, wherein the aforementioned time-series content representation is information indicating the presence or absence of images relating to products, people, companies, or trademarks in each scene, and the presence or absence of audio relating to products or trademarks in each scene.
7. The video content evaluation system according to claim 1, wherein the first time-series transmission content is information indicating an attitude or action to encourage the viewer for each scene.
8. A computer-based method for evaluating video content, Acquisition steps for obtaining video content, The acquired video content is input into a content understanding model, and an extraction step is performed to extract the time-series content representation. The extracted time-series content representation is input into a content estimation model to estimate the content of the first time-series; this is the first estimation step. A second estimation step is to estimate the emotions of the viewers of the video content based on the estimated first time series of transmitted content, An output step that outputs based on the estimated emotion, A video content evaluation method having the following characteristics.
9. In the second estimation step, a first surprise value is calculated for each scene based on the extracted time-series content representation, the estimated first time-series communication content, and the communication content estimation model, and the viewer's emotions are calculated based on the calculated first surprise value. In the output step, output is produced based on the calculated emotion. The video content evaluation method according to claim 8.
10. The video content evaluation method according to claim 9, wherein the second estimation step calculates the viewer's emotions based on the number of scenes in which the calculated first surprise value is less than or equal to a first predetermined value, the number of scenes in which the calculated first surprise value exceeds a second predetermined value, and the maximum value, minimum value, average value, average of first derivatives, average of second derivatives, or a combination thereof of the calculated first surprise value.
11. The second estimation step further includes a learning rate calculation step, In the aforementioned learning level calculation step, The aforementioned communication content estimation model is updated to minimize the surprise value for each scene. The extracted time-series content representation is input into the updated transmission content estimation model to estimate the transmission content of the second time-series. Based on the estimated second time-series transmission content and the updated transmission content estimation model, a second surprise value is calculated for each scene, and the viewer's learning level is calculated based on the amount of change between the first surprise value and the second surprise value. In the output step, output is further generated based on the calculated learning level. The video content evaluation method according to claim 8.
12. The video content evaluation method according to claim 8, wherein the aforementioned emotion includes the level of arousal.
13. The video content evaluation method according to claim 8, wherein the aforementioned time-series content representation is information indicating the presence or absence of images relating to products, people, companies, or trademarks in each scene, and the presence or absence of audio relating to products or trademarks in each scene.
14. The video content evaluation method according to claim 8, wherein the first time-series transmission content is information that indicates an attitude or action to encourage the viewer for each scene.
15. A program for causing a computer to function as any one of the parts described in claims 1 to 7.
Citation Information
Patent Citations
Data processing device and data processing method
JP2015121858A
Device for recommending content based on feeling of user, program and method
JP2015228142A
Viewing material evaluation method, viewing material evaluation system, and program
JP2017129923A
Estimation of state of brain activity from presented information of person
JP2022062574A
Information processing device, video distribution method, and video distribution program
JP2023052125A