Video content analysis system according to scene reaction of video content

The system addresses the challenge of real-time emotional response analysis by using biometric data and facial expressions to analyze viewer reactions, offering precise feedback for content improvement.

WO2026079569A1PCT designated stage Publication Date: 2026-04-16HISTRANGER INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/006213
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-10
Filing Date
2025-05-09
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

Existing film evaluation methods fail to accurately capture real-time emotional responses and immersion levels of viewers during video content consumption, relying on subjective surveys and group averages that do not account for individual differences.

Method used

A video content analysis system that collects real-time biometric data, including pulse and brainwave measurements, and facial expressions to analyze emotions, immersion, concentration, boredom, and empathy of viewers, using a combination of wearable sensors and machine learning algorithms to identify individual viewer reactions.

Benefits of technology

Enables precise analysis of viewer reactions to specific scenes, providing practical data for content creators to modify and improve their work by accurately verifying how scenes are received by audiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025006213_16042026_PF_FP_ABST
    Figure KR2025006213_16042026_PF_FP_ABST
Patent Text Reader

Abstract

A video content analysis system according to a scene reaction of video content according to one embodiment of the present invention comprises: a video processing unit which divides video content into a plurality of scenes and classifies each of the divided scenes for each preset reaction; an information collection unit for collecting biometric data of a viewer reacting while the video content is reproduced to the viewer; a reaction analysis unit for analyzing a reaction for each scene from the biometric data collected by the information collection unit; and a scene analysis unit for comparing and analyzing the reaction classified by the video processing unit and the reaction with respect to the biometric data analyzed by the reaction analysis unit for each scene.
Need to check novelty before this filing date? Find Prior Art

Description

Video content analysis system based on scene response of video content

[0001] The present invention relates to a video content analysis system based on scene response of video content.

[0002] Generally, film evaluation methods have relied primarily on quantitative indicators such as viewership ratings and audience numbers, or on qualitative analysis reflecting the subjective evaluations of critics.

[0003] Furthermore, the process of collecting viewer reactions has relied heavily on surveys. While such survey-based evaluations can gather subjective opinions after watching a movie, they had limitations in that they could not accurately capture the emotional responses or levels of immersion felt by viewers in real time while the movie was playing.

[0004] Recently, there have been attempts to analyze emotional responses by measuring brain activity or biosignals generated while viewers watch movies or content through technologies such as neurocinematics.

[0005] However, as with surveys, due to significant individual differences among viewers, approximate trends could only be identified through group averages. Furthermore, there were still limitations in the precise real-time sentiment analysis of individual viewers.

[0006] The technical problem that the present invention aims to solve is to provide a video content analysis system based on scene reactions of video content, which collects real-time biometric data and facial expression data of viewers and accurately analyzes the emotions, immersion, concentration, boredom, excitement, and empathy felt by viewers in each scene, thereby enabling content creators to verify how a specific scene is received by viewers and secure practical data that can be used for future content modification and improvement.

[0007] A video content analysis system based on scene reactions of video content according to one embodiment of the present invention comprises: a video processing unit that divides video content into a plurality of scenes and classifies each divided scene according to a preset reaction; an information collection unit that collects biometric data of a viewer that reacts while the video content is being played to a viewer; a reaction analysis unit that analyzes reactions for each scene from the biometric data collected by the information collection unit; and a scene analysis unit that compares and analyzes the reactions classified by the video processing unit and the reactions to the biometric data analyzed by the reaction analysis unit for each scene.

[0008] In addition, the information collection unit further collects viewer prior information regarding viewers, and the viewer post-information is information collected through a survey in advance before video content is provided to viewers, and the scene analysis unit is characterized by classifying and analyzing scenes according to viewer classifications classified from the viewer prior information.

[0009] In addition, the information collection unit further collects viewer post-information collected through a survey after the viewer watches video content, and the viewer post-information includes reaction information surveyed on reactions by scene, and the scene analysis unit compares and analyzes the reaction classified by the video processing unit for each scene with the reaction information included in the viewer post-information.

[0010] In addition, the information collection unit further collects facial expression data of viewers reacting while video content is being played to viewers, the reaction analysis unit analyzes facial expression changes from a pre-trained facial expression model based on the facial expression data according to each scene collected by the information collection unit and extracts one of the reactions, and the scene analysis unit compares and analyzes the reaction classified by the video processing unit, the reaction information included in the viewer post-information, and the reaction extracted by the reaction analysis unit for each scene.

[0011] In addition, it further includes a visualization unit that visualizes and displays each reaction to the scene through analysis data analyzed by the scene analysis unit.

[0012] In addition, the biometric data includes pulse data measured from a sensor worn by the viewer, and the response analysis unit is characterized by analyzing changes in pulse and abnormal peaks in each scene from the pulse data to analyze the viewer's response.

[0013] In addition, the response analysis unit extracts heart rate data indicating heart rate variability and intervals between heart rates from pulse data, and the response analysis unit analyzes the heart rate data to analyze the viewer's response.

[0014] In addition, the biometric data further includes brainwave data measured from a brainwave measuring device worn by the viewer, and the response analysis unit is characterized by analyzing changes in brainwaves for each scene from the brainwave data to analyze the viewer's response.

[0015] A video content analysis system based on scene reactions of video content according to one embodiment of the present invention collects real-time biometric data and facial expression data of viewers and accurately analyzes the emotions, immersion, concentration, boredom, excitement, and empathy felt by viewers in each scene, thereby enabling content creators to verify how a specific scene is received by viewers and secure practical data that can be used for future content modification and improvement.

[0016] FIG. 1 is a configuration diagram showing a video content analysis system based on scene response of video content according to one embodiment of the present invention.

[0017] FIG. 2 is an exemplary diagram showing an image processing unit of a video content analysis system based on a scene response of video content according to an embodiment of the present invention.

[0018] FIG. 3 is an exemplary diagram illustrating the process of extracting biometric data by removing noise from raw data of a video content analysis system based on scene response of video content according to one embodiment of the present invention.

[0019] FIG. 4 is an example diagram showing pulse data of a video content analysis system according to a scene response of video content according to an embodiment of the present invention.

[0020] FIG. 5 is an example diagram showing a pulse extraction graph, SDNN, and RR Interval graph of a video content analysis system according to a scene response of video content according to an embodiment of the present invention.

[0021] FIG. 6 is an example diagram showing a graph capable of detecting anomalies in a video content analysis system based on scene response of video content according to an embodiment of the present invention.

[0022] FIG. 7 is an example diagram showing the state of collecting facial expression data of a video content analysis system according to a scene response of video content according to an embodiment of the present invention.

[0023] FIG. 8 is an exemplary diagram showing a scene analysis unit of a video content analysis system based on a scene response of video content according to an embodiment of the present invention.

[0024] Hereinafter, various embodiments of the present invention will be described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present invention. The present invention may be embodied in various different forms and is not limited to the embodiments described herein.

[0025] To clearly explain the present invention, parts unrelated to the explanation have been omitted, and the same reference numerals are assigned to identical or similar components throughout the specification. Accordingly, the reference numerals described above may also be used in other drawings.

[0026] Furthermore, the size and thickness of each component shown in the drawings are depicted arbitrarily for the convenience of explanation, and thus the present invention is not necessarily limited to what is illustrated. Thickness may be exaggerated in the drawings to clearly represent various layers and regions.

[0027] Furthermore, the expression "identical" in the explanation may mean "substantially identical." In other words, it may be an identicality to the extent that a person with ordinary knowledge would accept it as such. Other expressions may also be those in which "substantially" has been omitted.

[0028] Furthermore, when a part in the description is described as 'including' a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. As used in this specification, 'part' refers to a unit that processes at least one function or operation, and may, for example, mean software, FPGA, or hardware components. The function provided by the 'part' may be performed separately by multiple components or integrated with other additional components. The 'part' in this specification is not necessarily limited to software or hardware, and may be configured to reside in an addressable storage medium or configured to run one or more processors. Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0029]

[0030] Referring to FIG. 1, a video content analysis system based on scene response of video content according to one embodiment of the present invention may include a video processing unit (10), an information collection unit (20), a response analysis unit (30), and a scene analysis unit (40).

[0031] The image processing unit (10) can divide the image content into multiple scenes and classify each divided scene according to a pre-set response.

[0032] Video content may include, but is not limited to, movies, etc., prior to release.

[0033] Here, the response may include, but is not limited to, emotions, moods, concentration, etc. In particular, regarding the pre-set response, for example, it may be an emotion such as joy, anger, sadness, or fear, but is not limited to these.

[0034] Specifically, referring to FIG. 2, the image processing unit (10) can divide the image content into scenes. The image processing unit (10) can organize the plot of the scenario in chronological order and classify the important events and emotional changes of each scene by listing them in detail. This can be used to analyze emotions, immersion, atmosphere, etc.

[0035] For each scene divided in this way, the image processing unit (10) can organize and classify what kind of scene each scene is (action, emotional scene, etc.), what emotional state it represents (joy, sadness, anger, etc.), and what role it plays in the progression of the story (introduction, climax, conclusion). This can be called scene-specific characteristic labeling, but is not limited thereto.

[0036] In addition, emotional states may utilize Russell's two-dimensional emotion model, but are not limited to it.

[0037] In addition, information classified by scene can be constructed as a content dataset. The content dataset may be constructed by classifying what type of scene each scene is (e.g., action scene, emotional scene, etc.), classifying the major emotions (joy, sadness, anger, etc.) appearing in each scene, and classifying what role the scene plays in the flow of the story within the scenario progression structure (e.g., introduction, conflict, climax, conclusion), but is not limited thereto.

[0038] Here, the emotions classified by the image processing unit (10) can be compared and analyzed with the biometric data (22) described later. Through this, by analyzing the level of immersion, interest, or concentration described later as a change in emotion, it is also possible to analyze how much the audience was immersed, interested, or focused on each scene.

[0039] The information collection unit (20) collects biometric data (22) of viewers in real time while video content is being played to viewers.

[0040] The biometric data (22) is based on a PPG (Photoplethysmogram, photoplethysmogram sensor), and various physiological responses such as the viewer's pulse, blood pressure, and oxygen saturation can be measured.

[0041] For example, biometric data (22) can be measured using a wearable device equipped with a PPG, and various devices such as smartwatch type, band type, earplug type, and finger-worn type can be used, but are not limited thereto.

[0042] And, the information collection unit (20) can collect information after having a viewer wear a wearable device in a movie theater where video content can be played.

[0043] Here, the information collection unit (20) can further collect brainwave data measured using a brainwave measuring device. That is, the bio-data may further include brainwave data measured from a brainwave measuring device worn by a viewer.

[0044] And the reaction analysis unit can analyze the viewer's reaction by analyzing changes in brainwaves for each scene from the brainwave data.

[0045] EEG (Electroencephalography) can measure electrical signals generated in the brain through electrodes attached to the scalp. These signals are classified into alpha waves, beta waves, and gamma waves, and each type of brainwave can indicate a viewer's cognitive state and emotional response.

[0046] 2-channel EEG can collect brainwave data using two electrodes. The electrodes are mainly attached to locations such as the forehead or temples, and can record changes in brain activity in real time while watching video content.

[0047] Alpha waves appear when the viewer is in a comfortable or relaxed state. This indicates that the viewer is not immersed or is in a relaxed state during a specific scene.

[0048] In other words, alpha waves primarily appear in a relaxed or comfortable state. They appear when a viewer comfortably watches video content without stress or tension, and can frequently occur, for example, when eyes are closed or when not concentrating.

[0049] Beta waves appear strongly when viewers are focused or excited. If a viewer's beta waves increase during a tense scene, it indicates that the viewer is focusing on the scene or is feeling tense.

[0050] In other words, beta waves appear when concentration or cognitive activity is high. In particular, they appear strongly when viewers are immersed in a specific scene, solving a problem, or concentrating. These beta waves can also appear when tension or excitement is induced.

[0051] Gamma waves occur during complex cognitive activities or states of high concentration. These brainwaves appear strongly when the viewer is completely immersed.

[0052] Here, the levels of alpha, beta, and gamma waves are determined based on a reference point, which is the viewer's normally measured brainwaves.

[0053] For example, if a viewer's beta waves increase sharply during a specific action scene, this indicates that the viewer is tense, allowing for the measurement of tension. Conversely, if alpha waves decrease, it can be analyzed that the viewer feels tension during that scene.

[0054] An increase in gamma waves during viewing indicates that the viewer is deeply immersed in the video content. For example, if gamma waves are high during thrilling scenes or scenes where a complex story unfolds, it can be seen that the viewer is strongly focused on those scenes.

[0055] 2-channel EEG data can be fused with PPG (pulse measurement), facial expression data, and viewer post-event information. For example, if beta waves increase in EEG data and the viewer's pulse rises at the same time, it can be comprehensively analyzed that the viewer is immersed in or tense during the scene.

[0056] In addition, if the response "I felt nervous" in the viewer's post-event information matches the increase in beta waves in the EEG analysis, the viewer's emotional response can be accurately verified.

[0057] The reaction analysis unit (30) analyzes the reaction for each scene from the biometric data (22) collected by the information collection unit (20). At this time, the reaction may be emotions, mood, concentration, etc.

[0058] Here, the biometric data (22) may include pulse data (23) measured from a sensor worn by the viewer.

[0059] That is, the reaction analysis unit (30) can analyze the viewer's reaction by analyzing changes in pulse and abnormal peaks for each scene from the pulse data (23). In other words, if an abnormal peak is detected in the pulse data (23), it can be seen that a special emotion was induced in the viewer in that scene.

[0060] Additionally, the response analysis unit (30) can extract heart rate data indicating heart rate variability and intervals between heart rates from the pulse data (23).

[0061] And, the reaction analysis unit (30) can analyze the viewer's reaction by analyzing heart rate data.

[0062] Meanwhile, the biometric data (22) measured through PPG is explained in detail.

[0063]

[0064] Referring to FIG. 3, the raw data (21) of the PPG is a recording of the viewer's bio-signals, such as heart rate and blood flow, over time. The raw data (21) is very complex and may contain noise. Here, if the noise is removed, changes in the dramatic reaction to the video content can be analyzed more clearly. Accordingly, the bio-data (22) is the data from which noise has been removed from the raw data (21).

[0065] For example, if unnecessary vibrations or noise are removed from raw data (21) regarding the pulse, a stable pattern can be observed. At this time, the pulse data (23) with reduced noise signals increases the possibility of interpreting the bio-signal. That is, by clearly revealing a specific pattern (regular cycle of the pulse), the rhythm of the bio-signal and regular pulse changes can be observed, and subsequently, through analysis, the pulse rate, immersion, concentration, etc. can be measured.

[0066] These pulse data (23) may show the normal peaks (Heart Rate Peaks, green dots in FIG. 4) when the pulse signal peaks appear and when it is normal, and the anomalous peaks (Anomalous Peaks, red dots in FIG. 4) when there is a sudden change that deviates from the normal pulse signal.

[0067] The abnormal peak (red dot in Fig. 4) occurs when the viewer shows an abnormally strong reaction, such as suddenly becoming tense or surprised, in a specific scene. In other words, it can refer to a moment when emotional response or immersion changes dramatically.

[0068] Meanwhile, through pulse data (23), it is possible to extract a pulse that shows changes in pulse over time, an SDNN (heart rate variability) that calculates the standard deviation of the interval by analyzing the interval between each peak, and an RR Interval graph that measures the interval between heart rates.

[0069]

[0070] First, referring to Figure 5, the pulse extraction graph visualizes the change in a viewer's pulse over time. The pulse extraction graph allows one to observe the variability of the pulse while the viewer is watching video content. The peaks appearing in the pulse extraction graph indicate rapid changes in the pulse, which may be moments when the viewer is tense or surprised in a specific scene.

[0071] Additionally, the abnormal peak in the pulse data (23) may represent a section where the viewer's emotional response is strong, and this section may be displayed on the pulse extraction graph.

[0072] Next is SDNN (Standard Deviation of NN intervals). SDNN is an indicator used to evaluate Heart Rate Variability (HRV), representing the standard deviation of heart rate intervals. This value is an important metric for assessing whether a viewer is tense or stressed. Since the graph varies for each viewer, a baseline is established, which can be defined as the viewer's normal heart rate pattern. In other words, if the SDNN value is lower than the baseline, it can be interpreted as a state of tension or immersion, while if the value is higher than the baseline, it can be interpreted as a stable or relaxed state.

[0073] In addition, a sharp decrease in SDNN value indicates the possibility that the viewer experienced strong stress or emotional changes in a specific scene.

[0074] Finally, the RR Interval is an indicator that measures the variability of heart rate intervals (NN intervals) and is another important factor in evaluating heart rate variability. Since the RR Interval graph varies for each viewer, a reference interval is set, which can be based on the viewer's normal heart rate fluctuation pattern. Here, the heart rate fluctuation pattern refers to the intervals that occur when intervals between heart rates occur repeatedly.

[0075] If the RR Interval becomes shorter than the reference interval, it indicates that the viewer is in a tense or excited state, and if it becomes longer than the reference interval, it indicates that the viewer is in a stable or relaxed state. Additionally, the RR Interval graph allows viewers to check how the heart rate interval changes over a certain period of time and to analyze scenes where the heart rate changes rapidly at specific moments.

[0076] In other words, the pulse extraction graph visually displays changes in the viewer's pulse over time, allowing identification of when specific emotional responses occurred.

[0077] Furthermore, SDNN indicates heart rate variability and plays an important role in assessing stress or tension states.

[0078] In addition, RR Interval shows the variability of heart rate intervals and can assess how tense or relaxed the viewer is.

[0079] Through these three graphs, the emotional and physiological effects of each scene in video content on viewers can be identified.

[0080] In addition, based on the biometric data (22) collected and analyzable in this way, an anomaly detection algorithm (e.g., CNN, LSTM, Isolated Forest, VAE, one class svm, etc.) can be applied to analyze whether the viewer's heart rate deviates significantly from the baseline in a specific scene. That is, anomaly detection can be performed using a graph extracted through the algorithm.

[0081] CNN (Convolutional Neural Network) is an effective algorithm for analyzing image or signal patterns and can detect specific patterns in PPG signals.

[0082] LSTM (Long Short-Term Memory) is an algorithm that detects abnormal patterns in continuous data over time (time-series data). It can detect abnormal responses by analyzing changes in heart rate over time.

[0083] Isolated Forest is an algorithm specialized in detecting abnormal values ​​in data, capable of detecting abnormal fluctuations in a viewer's biosignals.

[0084] One-class SVM (Support Vector Machine) learns only normal data and can detect outlier data that falls outside the learned range.

[0085] Anomaly detection can be performed through these machine learning algorithms. To enable anomaly detection, the algorithm can extract graphs.

[0086]

[0087] Referring to FIG. 6, a graph capable of detecting anomalies can first be extracted, which can detect Point Anomaly (24, local anomaly). This graph represents a local abnormal reaction (Contextual Anomaly (25)) that occurred at a specific moment. At this time, the peak occurring in the normal pattern may indicate a point where the viewer is likely to be emotionally surprised or show a strong reaction. That is, it can be seen that while a normal pulse pattern continues, an abnormal peak (the red dot in the Point Anomaly (24) graph of FIG. 6) occurs at a specific moment.

[0088] Additionally, a graph capable of detecting contextual anomalies (25) can be extracted. Here, this graph can analyze contextual abnormal reactions that occurred over a longer period of time. Accordingly, in this graph, it can be seen that an abnormal pattern persists, and it can be seen that the viewer may have been abnormally tense or immersed during that period. Referring to the Contextual Anomaly (25) graph in Fig. 6, the red dots represent abnormal heart rate fluctuations, the green dots represent normal fluctuations, and the yellow shaded area represents the section where an abnormal pattern occurred during a specific time period.

[0089] Finally, a graph detecting Collective Anomaly (26) can be extracted. This graph may indicate a continuous period of collective abnormal reaction. This is when abnormal pulse changes occur over several moments. This period may be when viewers repeatedly showed strong emotional reactions in a specific scene, or when tension became very high, causing significant changes in heart rate or pulse variability. Referring to the Collective Anomaly (26) graph in Fig. 6, the red shaded area indicates an area where collective abnormal heart rate fluctuations occurred, which means that during a specific period, the audience continuously reacted strongly with emotions of tension or immersion.

[0090] Therefore, Point Anomaly (24) can detect momentary abnormal reactions to identify dramatic reactions (surprise, fear, etc.) to video content, and Contextual Anomaly (25) indicates that abnormal emotional reactions have occurred continuously in specific sections, so it is possible to detect changes in mood due to scene transitions. Also, Collective Anomaly (26) shows a situation where an abnormal pattern occurs collectively over a long period of time, so it is possible to identify the possibility of decreased concentration on video content, such as restlessness, or to identify high levels of concentration, such as tension or immersion.

[0091] Accordingly, the scene analysis unit (40) can compare and analyze the reaction classified by the image processing unit (10) and the reaction to the bio-data (22) analyzed by the reaction analysis unit (30) for each scene.

[0092] For example, the image processing unit (10) emotionally analyzes the action scene of a movie and predicts that the scene will cause tension. This analysis classifies that the scene will cause tension or surprise in the viewer based on the scenario, composition of the scene, music, direction, etc. And if tension or surprise is actually analyzed through the abnormal peak reaction in the biometric data (22), it can be determined that the prediction of the analysis by the image processing unit (10) is successful.

[0093] That is, the scene analysis unit (30) can verify whether the video content has accurately induced the intended emotional effect through a comparative analysis of the classification of the video processing unit (10) and the actual reaction of the viewer (22), and through this, the content creator can specifically analyze the emotional reaction of the viewer and use it as an important factor for future content improvement.

[0094]

[0095] Referring to FIG. 7, the information collection unit (20) collects the viewer's pulse data (23) and facial expression data in real time while the viewer is watching an action scene.

[0096] The reaction analysis unit (30) found that the viewer's heart rate increased rapidly as a result of analyzing the pulse data (23) from the biometric data (22), and that a surprised expression was detected on the viewer's face as a result of analyzing the facial expression data.

[0097] In other words, an abnormal peak in the pulse is detected, indicating that the viewer has fallen into a state of surprise or tension at a specific moment.

[0098] The scene analysis unit (40) can compare and analyze that the scene analysis unit (40) has classified the scene as one that induces tension as the image processing unit (10), and that the reaction analysis unit (30) has detected a biological reaction in which the viewer actually felt tension through a rise in heart rate and a surprised expression. That is, the analysis of the scene analysis unit (40) is consistent, so it can be confirmed that this scene induces strong tension in the viewer as intended.

[0099] Conversely, if there were no abnormal reactions in the viewer's biometric data (22), the producer could conclude that the scene was not as effective as expected.

[0100] Accordingly, the scene analysis unit (40) can compare the emotion prediction (tension) of the image processing unit (10) with the bio-data (22) and facial expression data analysis (heart rate increase, surprise expression) of the reaction analysis unit (30) to analyze whether the viewer's actual emotional response matches the intended effect of the video content.

[0101] Meanwhile, the information collection unit (20) can collect more viewer prior information about the viewer.

[0102] In this case, viewer pre-information refers to information collected through a survey in advance before video content is provided to viewers.

[0103] The scene analysis unit (40) can classify and analyze scenes according to the viewer classification classified in the viewer dictionary information.

[0104] Viewer prior information can include various factors such as age, gender, movie genre preferences, and basic tendencies regarding emotional responses. For example, questions related to prior knowledge shown in images can play a crucial role in identifying viewers' prior expectations regarding video content. This allows for the classification of viewer groups based on specific characteristics. For instance, it is possible to distinguish between groups that prefer a particular genre and those that do not, or between viewers who are likely to exhibit specific emotional responses.

[0105] At this time, the scene analysis unit (40) can divide viewers into specific classification groups based on viewer prior information collected from a preliminary survey and analyze the reactions of viewers belonging to those groups. For example, it can analyze whether viewers of a specific age group or those who prefer a specific genre showed a specific emotional reaction in a certain scene or how immersed they were. Through this, accurate analysis of each scene can be performed and the emotional effect of the content on a specific group can be evaluated.

[0106] Additionally, the information collection unit (20) can collect more facial expression data of viewers reacting while video content is being played to viewers.

[0107] At this time, facial expression data can be collected using a infrared camera (NIR). Infrared cameras can accurately collect facial expression data of viewers even in dark environments (under 1 lux). This is useful for detecting the facial expressions of audiences in dark places such as movie theaters.

[0108] The reaction analysis unit (30) can extract one reaction after analyzing the facial expression change from a pre-trained facial expression model for each scene of the facial expression data collected by the information collection unit (20).

[0109] Specifically, the viewer's face is recognized in the collected facial expression data, and only the face portion is cropped and analyzed. Through this, data capable of analyzing the key features of facial expressions can be obtained. By utilizing a pre-trained vision model, 45 key landmarks are extracted from the viewer's face, and the extracted landmarks allow for the analysis of changes in facial expressions based on the movements of the eyes, mouth, nose, etc.

[0110] That is, the reaction analysis unit (30) analyzes the facial expression data for each scene collected by the information collection unit (20), and can detect changes in the viewer's facial expression based on changes in landmarks to extract emotional reactions (e.g., joy, sadness, surprise, anger). This process can precisely analyze changes in facial expression using a pre-trained facial expression model.

[0111] Additionally, the scene analysis unit (40) can compare and analyze the reactions classified by the video processing unit (10) for each scene, the reaction information included in the viewer post-information, and the real-time facial expression data extracted from the reaction analysis unit (30). Through this, the degree of agreement between the facial expression reaction shown by the viewer in a specific scene and the reaction information can be analyzed, and the viewer's reaction to the video content can be comprehensively evaluated.

[0112] For example, referring to FIG. 8, regarding a specific scene, the image processing unit (10) predicted that it would induce tension (Fig. 8A), and in a post-viewer survey, the viewer evaluated the scene as a very tense scene (Fig. 8B). As a result of analyzing facial expression data, the viewer's eyes widened and a surprised expression was detected, confirming that the viewer was actually tense and surprised (Fig. 8C). When the scene analysis unit (40) compared these three pieces of information, it was clearly proven that the viewer felt tension in this scene as intended. Here, if tension or surprise is actually analyzed through abnormal peak reactions in the biometric data (22) (Fig. 8D), the scene analysis unit (40) can comprehensively analyze that the viewer felt tension in this scene as intended.

[0113] Through this, content creators can successfully evaluate the emotional impact of specific scenes and utilize it as an important factor in future content creation and improvement.

[0114] In addition, the information collection unit (20) can collect additional viewer post-information collected through a survey after the viewer watches the video content.

[0115] At this time, the viewer post-information may include response information surveyed on reactions by scene, and the response information may be emotions, mood, and concentration.

[0116] In other words, viewer post-event information refers to information collected from viewers through surveys after video content has been provided. This refers to data that specifically identifies how viewers reacted to the video content. This may include questions regarding impressions by scene, genre, or empathy. Through this, it is possible to determine what emotional reactions viewers showed to specific scenes, how much they empathized with content of a particular genre, and so on.

[0117] Specifically, the survey is based on Russell's two-dimensional emotion archetype model and can classify the emotions felt by viewers in a specific scene. Emotions are primarily divided into Valence (positive-negative axis) and Arousal (axe level axis), and based on this, the seven major emotional states (Neutral, Sad, Happy, Angry, Disgust, Fear, Surprise) felt by viewers can be specifically recorded.

[0118] In addition, viewers can record in detail the emotions they felt while watching specific scenes in the survey. For example, they can assign specific scores for emotional reactions (interest, positive / negative emotions, boredom, etc.) to each scene. Furthermore, the survey may include questions to record not only the viewer's mood and emotional reactions but also their level of empathy for the genre. For instance, it may include a question evaluating how much empathy the viewer felt with a movie of a specific genre.

[0119] And, the scene analysis unit (40) can compare and analyze the reactions classified by the image processing unit (10) for each scene with the reaction information included in the viewer post-information. In addition, it can comprehensively compare and analyze with facial expression data.

[0120] In other words, viewer post-event data, which records the reactions viewers showed to specific scenes, is used to analyze viewers' emotional responses, concentration, and interest levels in a multidimensional manner. Through this, it is possible to determine whether a specific scene evoked positive or negative emotions in viewers, or whether it triggered a strong emotional response. For example, a scene rated as "interesting" in a survey may exhibit high levels of arousal and positive emotions on the axes of Valence and Arousal.

[0121] In addition, the aforementioned scene analysis unit (40) can verify whether the emotions and immersion felt by the viewer in a specific scene actually match the biometric data (22) or whether the reaction was stronger than expected. Through this, the content creator can use it to determine which scene the viewer's emotional immersion was high in, or to modify the editing direction so that the audience can empathize or immerse themselves more in a specific scene.

[0122] The video content analysis system based on the scene response of video content according to the present invention further includes a visualization unit (50).

[0123] The visualization unit (50) can visualize and display each reaction to the scene through the analysis data analyzed by the scene analysis unit (40).

[0124] For example, it may be graphics and charts to allow users to intuitively understand the viewer's reaction at a glance, but is not limited to this.

[0125] Specifically, real-time emotional reactions felt by viewers in each scene are visualized through various forms of graphs and charts. For example, emotional reactions such as joy, sadness, surprise, and anger felt by viewers in a specific scene can be represented as bar charts or line graphs to track changes in each emotion over time.

[0126] In addition, indicators such as interest and immersion are visually represented scene by scene, allowing one to identify the extent to which a specific scene elicited an emotional response from the viewer. By highlighting scenes with high levels of emotional engagement, it demonstrates which scenes the audience focused on more.

[0127] Additionally, the visualization unit (50) classifies changes in facial expressions based on the viewer's facial expression data and visualizes them as emotion labels. Changes in facial expressions shown by the viewer in each scene are visualized graphically, and an emotional state (e.g., joy, sadness, surprise, fear) is derived based on the results of the facial expression analysis. Based on the Russell 2D emotion prototype model, whether the viewer's emotion is positive or negative is visualized through Valence (horizontal axis) and Arousal (vertical axis), allowing the intensity and type of emotion to be identified at a glance.

[0128] Additionally, the visualization unit (50) analyzes the viewer's pulse data (23) to detect changes in heart rate in specific scenes and links this to changes in emotion. Through this, physiological reactions when the viewer is nervous or surprised can be detected, and the results of the emotion analysis based on pulse fluctuations are visualized as a graph. By detecting abnormal peaks, the point in time when the viewer reacted strongly emotionally at a specific moment is highlighted, and this is indicated as a red area (example) so that the user can understand it intuitively.

[0129] Accordingly, the visualization unit (50) visualizes the real-time reactions of viewers to each scene in graphics and charts, allowing the user to visually understand the changes in the viewer's emotions and easily grasp the reactions to each scene of the video. Through this, the content creator can analyze how a specific scene affected the viewer and play an important role in improving the quality of the content based on interest, immersion, and emotional reactions.

[0130]

[0131] The drawings and detailed description of the invention referenced so far are merely exemplary of the invention and are used only for the purpose of explaining the invention, not to limit the meaning or the scope of the invention as defined in the claims. Therefore, those skilled in the art will understand that various modifications and equivalent alternative embodiments are possible therefrom. Accordingly, the true technical scope of protection of the invention should be determined by the technical spirit of the appended claims.

[0132] The embodiments described above may be implemented as hardware components, software components, and / or combinations of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an Arithmetic Logic Unit (ALU), a Digital Signal Processor (DSP), a microcomputer, a Field Programmable Gate Array (FPGA), a Programmable Logic Unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions.

[0133] The processing unit may execute an operating system and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For convenience of understanding, the processing unit may be described as being used as a single unit, but a person of ordinary skill in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements.

[0134] For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as a parallel processor, are also possible. Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure the processing unit to operate as desired or command the processing unit independently or collectively.

[0135] Software and / or data may be embodied in any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by a processing device or to provide instructions or data to a processing device. Software may be distributed over networked computer systems and stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0136] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software.

[0137] Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiments, and vice versa.

[0138] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents. Therefore, other implementations, other embodiments, and equivalents to the claims are also included within the scope of the claims set forth below.

Claims

1. An image processing unit that divides video content into multiple scenes and classifies each divided scene according to a pre-set response; An information collection unit that collects biometric data of the viewer reacting while the above video content is being played to the viewer; A reaction analysis unit that analyzes reactions for each scene from the biometric data collected by the above information collection unit; and A video content analysis system based on scene reactions of video content, comprising: a scene analysis unit that, for each scene, compares and analyzes the reaction classified by the image processing unit and the reaction to the bio-data analyzed by the reaction analysis unit.

2. In Paragraph 1, The above information collection unit is, Collect more viewer prior information regarding the aforementioned viewer, and The above viewer prior information is information collected through a survey in advance before the above video content is provided to viewers, and A video content analysis system based on scene response of video content, characterized in that the scene analysis unit classifies and analyzes scenes according to viewer classifications classified in the viewer dictionary information.

3. In Paragraph 1, The above information collection unit is, After the aforementioned viewer watches the video content, further collect viewer follow-up information collected through surveys, and The above viewer post-information includes response information surveyed on reactions for each scene, and A video content analysis system based on scene reactions of video content, characterized in that the scene analysis unit compares and analyzes, for each scene, the reaction classified by the image processing unit and the reaction information included in the viewer post-information.

4. In Paragraph 3, The above information collection unit is, Further collect facial expression data of the viewer reacting while the above video content is being played to the viewer, and The above-mentioned response analysis unit analyzes facial expression changes using a pre-trained facial expression model based on the facial expression data according to each scene collected by the above-mentioned information collection unit, and extracts one of the responses. A video content analysis system based on scene reactions of video content, characterized in that the scene analysis unit compares and analyzes, for each scene, the reaction classified by the image processing unit, the reaction information included in the viewer post-information, and the reaction extracted by the reaction analysis unit.

5. In any one of paragraphs 1 through 4, A video content analysis system based on scene reactions of video content, further comprising: a visualization unit that visualizes and displays each reaction to the scene through analysis data analyzed by the scene analysis unit.

6. In Paragraph 1, The above biometric data includes pulse data measured from a sensor worn by the viewer, and A video content analysis system based on scene reactions of video content, characterized in that the above reaction analysis unit analyzes changes in pulse and abnormal peaks for each scene from the pulse data to analyze the viewer's reaction.

7. In Paragraph 6, The above response analysis unit extracts heart rate data representing heart rate variability and intervals between heart rates from the above pulse data, and A video content analysis system based on scene reactions of video content, characterized in that the above reaction analysis unit analyzes the heart rate data to analyze the viewer's reaction.

8. In Paragraph 6, The above biometric data further includes brainwave data measured from a brainwave measuring device worn by the viewer, and A video content analysis system based on scene reactions of video content, characterized in that the above reaction analysis unit analyzes changes in brainwaves for each scene from the brainwave data to analyze the viewer's reaction.

Citation Information

Patent Citations

  • Audience emotion recognition method, device and system

    CN111401198A

  • Information processing device and program

    JP2021039541A

  • Method of providing customized learning contents based on brainwave information

    KR1020120113573A

  • Apparatus and method for analyzing viewers' responses to video content

    KR102512468B1

  • Video indexing based on viewers' behavior and emotion feedback

    US20030118974A1