Audience emotion analysis method, device and equipment based on multi-modal data and medium thereof

By analyzing audience emotions through multimodal data, generating an emotional baseline curve, and identifying weaknesses, this approach solves the problem of inaccurate positioning in traditional film evaluation and provides precise suggestions for film optimization.

CN121504516APending Publication Date: 2026-02-10LIAONING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511723884.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-22
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional film evaluation methods rely on audience recollections and subjective expressions, which cannot accurately pinpoint the emotional low points within a film, making it difficult for creators to optimize narrative pacing and audiovisual design in a targeted manner.

Method used

Based on multimodal analysis of audience skin conductance data and electrocardiogram data, a baseline curve of film emotion is generated and compared with the emotional curve of the film to be tested to identify weak links in emotional arousal. Optimization suggestions are then generated by combining film metadata.

Benefits of technology

It enables precise identification of weak points in evoking emotions in films, provides scientific and accurate optimization guidance, and improves the practicality and accuracy of film emotion assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504516A_ABST
    Figure CN121504516A_ABST
Patent Text Reader

Abstract

The invention relates to an audience emotion analysis method and device based on multi-modal data, equipment and a medium thereof. The method comprises the following steps: screening a plurality of historical reference films of a target type from a historical film library, and fusing audience skin electrical response and electrocardio data to establish a type film emotion reference curve; collecting same-mode physiological data of a to-be-tested film to generate a to-be-tested film emotion curve, comparing the continuous time periods when the recognition values of the two curves are continuously lower than the reference, and positioning the continuous time period when emotion arousing is weak. By adopting the method, the emotional valley in the movie can be accurately positioned by using objective physiological data, and the defect that traditional subjective evaluation cannot quantify time-period play is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of emotion analysis, and particularly relates to a viewer emotion analysis method, device and equipment based on multi-modal data and a medium thereof. BACKGROUND

[0002] With the continuous improvement of the evaluation system of the film and television industry, there are film evaluation technologies based on questionnaires, box office or word-of-mouth scores in the market, which are characterized by low data collection cost and simple implementation, thereby forming the traditional evaluation method of post-subjective investigation + macro-indicator statistics which is currently widely used.

[0003] In the traditional technology, the film producer usually obtains viewer feedback through questionnaire investigation, social media public opinion monitoring or simple heart rate collection after the film is released, and then makes an empirical judgment on the overall appeal of the film according to the scoring results or text sentiment polarity statistics, and makes subjective adjustments in subsequent projects accordingly.

[0004] However, the current traditional evaluation method cannot accurately locate the emotional trough of a specific period due to the dependence on viewer recall and subjective expression, which makes it difficult for the creator to know the weak link of the film that really triggers the viewer to leave the film, so that the narrative rhythm and sound design cannot be quantitatively and targetedly optimized. SUMMARY

[0005] Therefore, it is necessary to provide a viewer emotion analysis method, device and equipment based on multi-modal data which can accurately locate and quantify the weak link of film emotion arousal based on objective physiological data in view of the above technical problems.

[0006] In a first aspect, the application provides a viewer emotion analysis method based on multi-modal data, comprising:

[0007] screening a historical film library according to a preset screening condition to obtain a plurality of benchmark films of a target type;

[0008] performing first segmented statistical processing on viewer galvanic skin response data and electrocardiogram data corresponding to the plurality of benchmark films to generate a type film emotion benchmark curve;

[0009] performing second segmented statistical processing on viewer galvanic skin response data and electrocardiogram data corresponding to the to-be-tested film to generate a to-be-tested film emotion curve;

[0010] comparing the to-be-tested film emotion curve with the type film emotion benchmark curve to identify a continuous time period in which the value is continuously lower than the type film emotion benchmark curve;

[0011] generating a corresponding viewer emotion analysis result based on the continuous time period; the viewer emotion analysis result is used to represent a weak link of the to-be-tested film emotion arousal.

[0012] In one of the embodiments, the skin conductance response data of the audience corresponding to the plurality of reference films and the electrocardiogram data are subjected to first segment statistical processing to generate a type film emotion reference curve, including:

[0013] The skin conductance response data of the audience is subjected to skin conductance response event detection to generate an SCR event sequence;

[0014] The electrocardiogram data is subjected to heart rate variability analysis to generate an HRV feature sequence;

[0015] The SCR event sequence and the HRV feature sequence are subjected to time alignment processing according to a film timeline to generate synchronized physiological feature data;

[0016] The synchronized physiological feature data is subjected to segment average processing to generate a type film emotion reference curve.

[0017] In one of the embodiments, the skin conductance response data of the audience is subjected to skin conductance response event detection to generate an SCR event sequence, including:

[0018] The skin conductance response data of the audience is subjected to band-pass filtering processing to generate a filtered signal;

[0019] The filtered signal is subjected to baseline correction processing to generate a clean skin conductance signal;

[0020] Based on the clean skin conductance signal, peak points greater than preset amplitude threshold and rising slope threshold are screened to generate a candidate SCR event set;

[0021] The candidate SCR event set is subjected to motion artifact filtering processing to generate an SCR event sequence.

[0022] In one of the embodiments, the to-be-tested film emotion curve is compared with the type film emotion reference curve to identify a continuous time period in which the numerical value is continuously lower than the type film emotion reference curve, including:

[0023] The type film emotion reference curve and the to-be-tested film emotion curve are subjected to point-by-point difference operation to generate an emotion difference curve;

[0024] The emotion difference curve is subjected to sliding window average filtering processing to generate a smoothed emotion difference curve;

[0025] A continuous data point sequence in which the numerical value of the smoothed emotion difference curve is lower than a preset negative threshold is marked as a candidate low arousal section;

[0026] From the candidate low arousal section, a section with a duration exceeding a preset duration threshold is screened to generate a continuous time period.

[0027] In one of the embodiments, after generating the corresponding audience emotion analysis result based on the continuous time period, further comprising:

[0028] Extracting corresponding timestamp information from the metadata of the to-be-tested film based on the continuous time period;

[0029] According to the timestamp information, performing video content feature analysis on the to-be-tested film to generate a plot type label, a picture dynamic index, and an audio intensity index;

[0030] Based on the plot type label, the picture dynamic index, and the audio intensity index, performing modification suggestion reasoning to generate a structured optimization suggestion report.

[0031] In one of the embodiments, according to the preset screening condition, the historical film library is screened to obtain a plurality of benchmark films of a target type, comprising:

[0032] Obtaining film box office data and word-of-mouth score data of a plurality of professional film review platforms;

[0033] Preliminarily screening films in the historical film library that are greater than a preset box office ranking threshold and a word-of-mouth score threshold to generate a candidate film set;

[0034] Performing multi-platform ranking consistency verification on the candidate film set, and generating a plurality of benchmark films from the candidate films that pass the verification.

[0035] In one of the embodiments, before generating the to-be-tested film emotion curve based on the audience galvanic skin response data and electrocardiogram data corresponding to the to-be-tested film, further comprising:

[0036] Receiving pre-recorded galvanic skin response data and electrocardiogram data matched with the target audience portrait;

[0037] Performing signal quality assessment on the pre-recorded galvanic skin response data and electrocardiogram data to generate a quality assessment result;

[0038] Based on the quality assessment result, performing rejection processing on unqualified data to generate effective multi-modal physiological data; the effective multi-modal physiological data is used for the second segmented statistical processing.

[0039] In a second aspect, the application further provides an audience emotion analysis device based on multi-modal data, comprising:

[0040] A historical film screening module for screening a historical film library according to a preset screening condition to obtain a plurality of benchmark films of a target type;

[0041] A benchmark curve generation module for performing first segmented statistical processing on audience galvanic skin response data and electrocardiogram data corresponding to the plurality of benchmark films to generate a type film emotion benchmark curve.

[0042] The test curve generation module is used to perform second-segment statistical processing based on the audience's skin conductance response data and electrocardiogram data corresponding to the test film to generate the emotional curve of the test film.

[0043] The emotion comparison module is used to compare the emotion curve of the film under test with the emotion benchmark curve of the genre film and identify the continuous time period when the value is consistently lower than the emotion benchmark curve of the genre film.

[0044] The results output module is used to generate corresponding audience sentiment analysis results based on continuous time periods; the audience sentiment analysis results are used to characterize the weak links in the emotional arousal of the film under test.

[0045] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the audience sentiment analysis method based on multimodal data as described in the first aspect.

[0046] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audience sentiment analysis method based on multimodal data as described in the first aspect.

[0047] The aforementioned audience sentiment analysis method, device, equipment, and media based on multimodal data, by selecting benchmark films of target genres according to preset conditions, can identify representative films from a historical film database, providing a reliable sample basis. By generating a genre-specific sentiment benchmark curve through segmented statistical analysis of electrodermal and electrocardiogram data, the representational effect of these two physiological data on emotions can be combined, making the benchmark curve more closely reflect the audience's actual emotional state. The same logic is used to generate sentiment curves for the films under test, ensuring consistency and comparability with the benchmark curve. By identifying continuous time periods where values ​​are consistently lower than the benchmark through curve comparison, specific segments in the film under test with insufficient emotional arousal can be accurately located. The analysis results pointing to weak points in emotional arousal can directly provide clear directions for film optimization. This method can scientifically and accurately complete audience sentiment analysis for genre films, providing reliable guidance for the emotional optimization of films under test and improving the practicality and accuracy of genre film sentiment evaluation. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1A flowchart illustrating an audience sentiment analysis method based on multimodal data provided by this invention;

[0050] Figure 2 A flowchart illustrating a method for generating a structured optimization suggestion report in one optional embodiment of the present invention;

[0051] Figure 3 This is a schematic diagram of the structure of an audience sentiment analysis device based on multimodal data provided by the present invention. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0053] In one embodiment, such as Figure 1 As shown, a method for audience sentiment analysis based on multimodal data is provided. This embodiment illustrates the application of this method to a sentiment analysis terminal. It is understood that this method can also be applied to a server, or to a system including both a sentiment analysis terminal and a server, and is implemented through the interaction between the sentiment analysis terminal and the server. In this embodiment, the method includes the following steps:

[0054] S101. Filter the historical film library according to the preset filtering conditions to obtain multiple benchmark films of the target type.

[0055] Optionally, film data from multiple professional film review platforms, such as Douban, Maoyan Professional Edition, and Lighthouse Professional Edition, can be obtained through compliant data collection methods. The preset screening criteria include both box office and word-of-mouth indicators. First, films with high box office rankings and word-of-mouth scores on a single platform are initially screened from the historical film database. Then, the consistency of the rankings of the corresponding films across multiple platforms is verified by comparing their positions on different platforms. Films with high rankings in both dimensions on each platform are retained. Finally, multiple benchmark films of the target type are obtained. The above process is based on the psychological triangular verification theory to ensure that the benchmark films have high commercial value and broad representativeness.

[0056] S102. Based on the audience's skin conductance response data and electrocardiogram data corresponding to multiple benchmark films, perform the first segment statistical processing to generate the genre film emotion benchmark curve.

[0057] Optionally, audience skin conductance response (SCRR) data and electrocardiogram (ECG) data are acquired when viewers watch multiple benchmark films. SCRR data reflects the intensity of emotional arousal, while ECG data reflects autonomic nervous activity and emotional stability through heart rate variability (HRV). The first segmented statistical processing divides the film into several consecutive time periods based on its length, taking into account the film's plot rhythm; for example, the segment length can be appropriately shortened for action films. Then, statistics are calculated for both types of data within each time period, such as the event frequency of SCRR and the mean of ECG data. These two types of statistics are then fused to obtain the emotional representation values ​​for each time period. All representation values ​​are then concatenated chronologically to generate a genre-specific emotional baseline curve.

[0058] S103. Based on the audience's skin conductance response data and electrocardiogram data corresponding to the film to be tested, perform the second segmented statistical processing to generate the emotional curve of the film to be tested.

[0059] Optionally, obtain the skin conductance response data and electrocardiogram data of viewers matching the target audience while watching the film to be tested. Using the same segmentation method and statistical approach as the first segmentation, i.e., the same time period division rules and the same statistical calculation and fusion logic, process the corresponding data to obtain the emotional representation values ​​for each time period of the film to be tested. Then, concatenate the corresponding representation values ​​in chronological order to generate the emotional curve of the film to be tested. It is essential to ensure that the time dimension and data statistical standards of the two curves are consistent.

[0060] S104. Compare the emotional curve of the film to be tested with the emotional baseline curve of the genre film to identify the continuous time period in which the value is consistently lower than the emotional baseline curve of the genre film.

[0061] Optionally, the sentiment baseline curve of the genre film is aligned with the sentiment curve of the film under test at the same time node. The difference between the data of the curve under test and the data of the baseline curve is calculated at each time node. Data points with negative differences are identified. It is further determined whether these negative difference data points are continuous, forming a continuous time period sequence. By setting reasonable continuity judgment rules, such as continuous coverage of at least one complete segment, segments with practical significance are selected from the sequence, and finally at least one continuous time period with a value continuously lower than the baseline curve is obtained.

[0062] S105. Based on continuous time periods, generate corresponding audience sentiment analysis results; the audience sentiment analysis results are used to characterize the weak links in the emotional arousal of the film under test.

[0063] Optionally, the obtained continuous time periods are precisely matched with the timeline of the film under test to determine the specific start and end positions of the corresponding time periods within the film. Combined with the degree of difference in the emotional curve values, audience emotional analysis results are generated. These results clearly identify segments in the film where the level of emotional arousal is consistently lower than the benchmark of similar high-value films, directly characterizing the weak points in emotional arousal within the film under test and providing a clear basis for film optimization.

[0064] The aforementioned audience sentiment analysis method based on multimodal data utilizes box office and word-of-mouth data from multiple professional film review platforms, combined with the psychological triangulation theory to select benchmark films for the target genre. Then, it performs segmented statistical processing on audience skin conductance and electrocardiogram data for both the benchmark and the test films, generating corresponding sentiment curves. Subsequently, by comparing these sentiment curves, it identifies consecutive time periods where the values ​​are consistently lower than the benchmark, ultimately generating analytical results representing weak points in emotional arousal. This method addresses the shortcomings of traditional methods, such as insufficient representativeness of benchmark film selection from a single platform, one-sided quantification of emotions using single physiological data, and vague identification of weak points. It effectively improves the scientific rigor and relevance of audience sentiment analysis for genre films, providing precise direction for optimizing film sentiment.

[0065] In one embodiment, a first segmented statistical processing is performed based on audience skin conductance data and electrocardiogram data corresponding to multiple benchmark films to generate a genre film sentiment benchmark curve, including:

[0066] S201. Detect skin conductance response events from audience skin conductance response data and generate SCR event sequences.

[0067] Optionally, the acquired raw skin conductance response (SCR) data is preprocessed. First, a Butterworth bandpass filter algorithm is used to remove high-frequency noise and low-frequency drift. High-frequency noise includes electrode contact interference, and low-frequency drift includes slow skin changes, retaining effective frequency band signals related to emotions. Then, the resting period, i.e., the mean of the data in the calm state before the viewer watches, is calculated. The filtered data is subtracted from this mean to complete baseline correction, resulting in clean SCR signals. An amplitude threshold and a rise slope threshold are set to select peak points that meet the conditions as candidate SCR events. Combined with synchronously recorded behavioral data, invalid events caused by motion artifacts are eliminated to generate an SCR event sequence.

[0068] S202. Perform heart rate variability analysis on electrocardiogram data to generate HRV feature sequences.

[0069] Optionally, the acquired raw ECG data is preprocessed to remove power line interference and motion artifacts, and then the RR interval sequence is extracted. The RR interval is the time interval between two adjacent R wave peaks. HRV (Heart Rate Variability) features are calculated based on this sequence, exemplarily including time-domain and frequency-domain features. Time-domain features include the standard deviation of the RR interval, and frequency-domain features include high-frequency and low-frequency components. The time-domain features reflect the overall heart rate fluctuations, while the frequency-domain features reflect the balance of the autonomic nervous system. The calculated HRV features are arranged in chronological order to generate an HRV feature sequence, which can help characterize the emotional stability of the audience.

[0070] S203. Based on the film timeline, perform time alignment processing on the SCR event sequence and HRV feature sequence to generate synchronized physiological feature data.

[0071] Optionally, timeline information is extracted from the metadata of the benchmark film. This timeline includes the film's start time, timestamps for each frame, and key plot points. The occurrence timestamps of each event in the SCR event sequence and the calculation timestamps of each feature in the HRV feature sequence are extracted separately. These two types of timestamps are then precisely matched with the film's timeline to ensure a one-to-one correspondence between SCR events and HRV features at the same time point. A time alignment algorithm is used to eliminate potential time discrepancies during data acquisition, generating synchronized physiological characteristic data.

[0072] S204. Perform segmented averaging on the synchronous physiological characteristic data to generate a type film emotion baseline curve.

[0073] Optionally, based on the duration and emotional change patterns of the benchmark film, the film's timeline is divided into several equal or unequal time segments. When dividing, natural plot segments should be prioritized to avoid interrupting the complete storyline. For the synchronous physiological characteristic data within each time segment, statistics on SCR events (such as event frequency and average amplitude) and HRV characteristics (such as mean and standard deviation) are calculated. A weighted average method is used to fuse the two types of statistics to obtain the emotional baseline value for that time segment. The weights are set according to the contribution of the two types of data to emotional representation. The emotional baseline values ​​of all time segments are concatenated in chronological order to generate a genre film emotional baseline curve.

[0074] In the above embodiment, the audience's electrodermal response (EDR) data is first filtered, baseline corrected, thresholded, and artifact removed to generate an SCR event sequence. Simultaneously, heart rate variability analysis is performed on the electrocardiogram (ECG) data to generate an HRV feature sequence. Then, the two sequences are time-aligned according to the film's timeline to obtain synchronized physiological characteristic data. Finally, a segmented averaging process is used to generate a genre film emotion baseline curve. This embodiment achieves the synergistic fusion of EDR and ECG multimodal data, avoiding the limitations of a single physiological indicator in comprehensively reflecting emotions, arousal intensity, and stability. This makes the generated baseline curve more closely match the audience's actual emotional state, significantly improving the accuracy and reliability of the baseline curve.

[0075] In one embodiment, skin conductance response event detection is performed on the audience's skin conductance response data to generate a skin conductance response (SCR) event sequence, including:

[0076] S301. Bandpass filter the audience's skin conductance response data to generate a filtered signal.

[0077] Optionally, the Butterworth bandpass filter algorithm is used to process the audience's electrodermal response data. First, the passband cutoff frequency is set according to the frequency range of effective emotional information in the electrodermal signal. Then, the preset filter order and signal sampling rate are determined. The filter coefficients are calculated using the standard design formula for Butterworth filtering, which is:

[0078]

[0079] in, This is the complex frequency domain transfer function of the Butterworth filter; For complex frequencies, that is, complex variables in the complex frequency domain, the mathematical form is: , For the actual part, For imaginary units, Angular frequency, Its function is to characterize the frequency and attenuation characteristics of a signal, providing a mathematical basis for filters to select signals in specific frequency bands; This is the preset filter order; The passband cutoff angular frequency, By passband cutoff frequency With sampling rate The derivation shows that the relationship is: Using this formula, with the filter order, passband cutoff frequency, and sampling rate as input parameters, the numerator coefficient *b* and denominator coefficient *a* of the filter are calculated. The original data is then input into the filter model constructed based on coefficients *b* and *a*. Through convolution operations, the effective frequency band signal is retained, while high-frequency noise above the upper cutoff frequency limit and low-frequency drift below the lower cutoff frequency limit are filtered out, ultimately generating the filtered signal.

[0080] S302. Perform baseline correction processing on the filtered signal to generate a clean skin electrical signal.

[0081] Optionally, first determine the rest period, which is the time period during which the audience does not watch the film and is in a calm state. Extract all data points of the filtered signal within this period, calculate their arithmetic mean, and record it as the rest period mean. The formula used is: The data at each time point t in the filtered signal are corrected, where This represents the original value of the filtered signal at time t. This is the corrected value. This baseline correction operation eliminates the influence of differences in baseline skin conductance levels among individual viewers on the emotional signal, generating a clean skin conductance signal that more accurately reflects emotional changes.

[0082] S303. Based on the clean skin electrical signals, peak points that are greater than the preset amplitude threshold and rise slope threshold are selected to generate a candidate SCR event set.

[0083] Optionally, an amplitude threshold and a rise slope threshold are preset. The amplitude threshold is used to exclude minor signal fluctuations caused by random noise, while the rise slope threshold is used to ensure that the selected signal changes have obvious emotional arousal characteristics. Each data point of the clean skin electrodermatology signal is traversed, and the amplitude of each data point relative to the resting period baseline is calculated, i.e., the difference between the data point value and the resting period mean. Simultaneously, the difference between the data point and the previous adjacent data point is calculated to obtain the rise slope of the signal. Data points with amplitudes greater than the amplitude threshold and rise slopes greater than the rise slope threshold are selected as peak points. The timestamp, amplitude, and other information of each peak point are recorded and integrated to form a candidate SCR event set.

[0084] S304. Perform motion artifact filtering on the candidate SCR event set to generate an SCR event sequence.

[0085] Optionally, acquire video recordings of viewer behavior synchronized with viewer electrodermal response (EDR) data, and use video analytics to identify periods in the video where viewer exhibits physical movements, such as raising an arm, adjusting posture, or significant head rotation, marking the start and end times of these periods. Extract the timestamp of each event from the candidate SCR event set and determine if the timestamp falls within the marked physical movement period. If the timestamp falls within the physical movement period, the event is considered invalid due to motion artifact interference and is removed from the set; if the timestamp does not fall within the physical movement period, the event is retained as a valid SCR event. Arrange all valid events in chronological order to generate an SCR event sequence.

[0086] In the above embodiment, Butterworth bandpass filtering is used to remove high-frequency noise and low-frequency drift from the audience's electrodermal response (EDR) data. Baseline correction is then performed by subtracting the resting period mean to eliminate individual differences in baseline EDR. Subsequently, effective peak points are selected based on preset amplitude and rise slope thresholds to form a candidate SCR event set. Finally, invalid events interfering with motion artifacts are removed by combining synchronized behavioral video to generate an SCR event sequence. This embodiment eliminates noise and interference factors layer by layer, ensuring the effectiveness and purity of the SCR event sequence.

[0087] In one embodiment, the sentiment curve of the film under test is compared with a genre film sentiment benchmark curve to identify continuous time periods in which the value is consistently lower than the genre film sentiment benchmark curve, including:

[0088] S401. Perform point-by-point difference calculation on the emotional baseline curve of the genre film and the emotional curve of the film to be tested to generate an emotional difference curve.

[0089] Optionally, the sentiment baseline curve for the genre film and the sentiment curve for the film under test are sampled at the same temporal resolution to ensure that the two curves have the same number of data points, and that each data point corresponds one-to-one with the same time point in the film. The formula used is: Perform point-by-point difference calculations on the two curves, where This represents the value of the emotion difference curve at time t. The value of the emotional curve of the film under test at time t. This represents the baseline emotional curve for the genre film at time t. All calculated values... Arranged chronologically, an emotion difference curve is generated, which visually reflects the emotional differences between the test film and the benchmark film at each time point.

[0090] S402. Perform sliding window averaging filtering on the emotion difference curve to generate a smooth emotion difference curve.

[0091] Optionally, the width and step size of the sliding window are preset. The window width is determined based on the required smoothness of the emotional changes, such as covering multiple consecutive data points to reduce the impact of random fluctuations. The step size is set according to the film's temporal resolution to ensure complete coverage of the curve without repetition or redundancy. The emotional difference curve is traversed in units of the sliding window. The arithmetic mean of all data points within each window is calculated, and this mean is used as the filtered value corresponding to the center time node of the window. The window is moved according to the set step size, and the above calculation process is repeated until the window has traversed the entire emotional difference curve, generating a smooth emotional difference curve. This curve can more clearly present the overall trend of the emotional changes of the film under test relative to the baseline curve.

[0092] S403. Mark the sequence of continuous data points in the smoothed emotion difference curve whose values ​​are lower than the preset negative threshold as candidate low arousal segments.

[0093] Optionally, a preset negative threshold is established. This threshold is determined based on the fluctuation range of the genre film's emotional baseline curve and the minimum requirement for emotional arousal. It is used to define the critical value at which the emotional level of the film under test is lower than the baseline curve. The smoothed emotional difference curve is traversed, and each data point is checked to see if its value is lower than the preset negative threshold. If it is lower, the data point is marked as a negative difference point. The time period corresponding to multiple consecutive negative difference points is extracted. During this time period, the emotional level of the film under test is consistently lower than the baseline curve. This time period is marked as a candidate low-arousal segment, ensuring that each segment corresponds to a potential part of the film where the emotion is consistently weak.

[0094] S404. From the candidate low wake-up segments, select segments whose duration exceeds a preset duration threshold and generate a continuous time period.

[0095] Optionally, a preset duration threshold is set, determined based on the minimum effective duration of a video segment. This avoids misjudging extremely short emotional fluctuations as weak points requiring optimization, ensuring that the selected segments have actual significance for emotional optimization. The start and end times of each candidate low-arousal segment are extracted, and the actual duration of each segment is obtained by calculating the difference between the end and start times. The duration of each segment is compared with the preset duration threshold, and candidate low-arousal segments whose duration exceeds the threshold are retained, ultimately generating a continuous time period. This time period represents the parts of the video under test where emotional arousal is insufficient and requires focused optimization.

[0096] In the above embodiment, a point-by-point difference calculation is performed between the emotion baseline curve of the genre film and the emotion curve of the film under test to generate an emotion difference curve. Then, a sliding window averaging filter is used to smooth the curve to reduce random fluctuations. Subsequently, continuous data points with values ​​below a preset negative threshold are marked as candidate low-arousal segments. Finally, segments with a duration exceeding a preset threshold are selected as valid continuous time periods. This embodiment accurately locates specific segments in the film under test where emotional arousal is insufficient, avoiding the problem that traditional overall numerical comparison cannot pinpoint weak points in detail, thus improving the accuracy and practical value of locating emotionally weak points.

[0097] In one embodiment, after generating the corresponding audience sentiment analysis results based on continuous time periods, the method further includes:

[0098] S501. Based on a continuous time period, extract the corresponding timestamp information from the metadata of the film to be tested.

[0099] Optionally, complete timeline information is extracted from the metadata of the film under test. This metadata includes precise timestamps for each frame, time markers for scene transitions, chapter divisions, and other information. The start and end times of the identified continuous time segments are matched with the time information in the metadata timeline to determine the specific position of each time segment on the film's timeline. Detailed information such as the corresponding frame timestamp range, scene number, and chapter name is extracted to form a timestamp information set. This set ensures that the analysis of the film content can accurately pinpoint specific segments with weak emotional impact, avoiding positioning bias.

[0100] S502. Based on the timestamp information, perform video content feature analysis on the film to be tested, and generate plot type tags, dynamic performance indicators and audio intensity indicators.

[0101] Optionally, based on the extracted timestamp information, video clips corresponding to periods of emotional weakness in the film to be tested are extracted. These video clips undergo multi-dimensional content feature analysis: plot analysis tools are used to identify plot types within the clips, such as dialogue scenes, action scenes, and transition scenes, generating plot type tags; image analysis algorithms are used to calculate parameters such as motion speed, shot switching frequency, and color saturation within the clips, quantifying the dynamics of the image and generating a dynamic index; audio analysis tools are used to extract audio features such as volume, frequency distribution, and rhythm changes, generating an audio intensity index, comprehensively capturing the content features and emotional transmission factors of the clips.

[0102] S503. Based on plot type tags, visual dynamism indicators, and audio intensity indicators, modification suggestions are derived, and a structured optimization suggestion report is generated.

[0103] Optionally, a rule base for associating content features with emotional arousal effects is first established. This rule base is constructed based on the general rules of emotional transmission in genre films. For example, low visual dynamism in action scenes can easily lead to insufficient emotional arousal, while abnormal audio intensity (too high or too low) in dialogue scenes can affect the audience's attention span. The parsed plot type tags, visual dynamism indicators, and audio intensity indicators are matched with the rules in the rule base to identify the specific content factors that cause weak emotional arousal in the corresponding time periods. Based on these factors, targeted modification suggestions are generated, such as increasing the frequency of shot transitions in a certain action scene or adjusting the audio volume of a certain dialogue segment. The suggestions are organized by content dimension to generate a structured optimization suggestion report.

[0104] In the above embodiments, based on the identified continuous low-arousal time periods, the corresponding precise timestamp information is extracted from the metadata of the film under test. Then, video segments are extracted based on the timestamps, and their plot type tags, visual dynamism indicators, and audio intensity indicators are analyzed. Finally, a structured optimization suggestion report is generated by combining the emotional transmission rules of genre films. This embodiment deeply correlates the emotional weakness reflected by physiological data with the specific content characteristics of the film, breaking through the limitations of traditional methods that rely solely on physiological data or subjective questionnaires to provide vague suggestions. This makes the generated optimization suggestions more targeted and operable, providing clear technical guidance for film editing and adjustment.

[0105] In one embodiment, the historical film library is filtered according to preset filtering conditions to obtain multiple benchmark films of the target type, including:

[0106] S601: Obtain film box office data and word-of-mouth rating data from multiple professional film review platforms.

[0107] Optionally, relevant data for the target type of film can be obtained from platforms such as Douban, Maoyan Professional Edition, and Lighthouse Professional Edition through open data interfaces or compliant data acquisition channels of various professional film review platforms. The acquired data includes box office data (such as total box office, opening week box office, daily average box office, and box office ranking) and reputation data (such as platform user ratings, number of ratings, professional film critic ratings, and reputation ranking). The acquired data is then linked and integrated by film name to ensure that the data for the same film on different platforms corresponds one-to-one, forming a raw dataset containing basic film information, box office data from multiple platforms, and reputation data from multiple platforms.

[0108] S602. Perform preliminary screening on films in the historical film database that exceed the preset box office ranking threshold and word-of-mouth rating threshold to generate a candidate film set.

[0109] Optionally, reasonable ranking thresholds are set for the box office data and audience rating data in the original dataset. The box office ranking threshold is determined by referencing the overall box office distribution of target-type films across various platforms, selecting films that rank among the top performers. The audience rating threshold is determined by referencing the general level of film ratings within a platform, selecting films with audience ratings higher than a certain score or ranking among the top performers. For each target-type film in the historical film database, it is determined whether it simultaneously meets the criteria of having a box office ranking higher than the box office ranking threshold and an audience rating higher than the audience rating threshold on any platform. Films that simultaneously meet these criteria are included in the candidate film set, initially selecting films that combine commercial performance and audience acceptance.

[0110] S603. Perform multi-platform ranking consistency verification on the candidate film set, and generate multiple benchmark films from the verified candidate films.

[0111] Optionally, for each film in the candidate film set, its box office ranking and audience rating ranking across all professional film review platforms that have acquired data are extracted. Consistency analysis methods, such as the Kendall coefficient of harmony test, are used to calculate the film's ranking consistency across different platforms, determining whether its ranking remains high across all platforms without significant differences. If a film ranks high in both box office and audience ratings on most platforms, and the ranking fluctuations are within a reasonable range, then the film is deemed to have passed the multi-platform ranking consistency verification. All verified films are then integrated to generate multiple benchmark films for the target genre, ensuring that the representativeness of the benchmark films is not affected by data bias from a single platform.

[0112] In the above embodiment, box office and audience rating data from multiple professional film review platforms are obtained through compliant channels. Films in the historical film database are then preliminarily screened according to preset box office rankings and audience rating thresholds to generate a candidate film set. Finally, the candidate films undergo multi-platform ranking consistency verification to determine multiple benchmark films. This embodiment, based on the psychological triangular verification theory, avoids the problem of insufficient representativeness of benchmark films due to single-platform or single-dimensional screening through multi-platform dual-dimensional screening and consistency verification, ensuring that the benchmark films possess high commercial value and broad audience acceptance.

[0113] In one embodiment, before generating the emotional curve of the film under test by performing a second segmented statistical processing based on the audience's skin conductance response data and electrocardiogram data corresponding to the film under test, the method further includes:

[0114] S701: Receives pre-recorded skin conductance data and electrocardiogram data that match the target audience profile.

[0115] Optionally, based on the target audience positioning of the film to be tested, a target audience profile is first determined. This profile clearly defines the age range, gender ratio, viewing preferences (e.g., young people who prefer action films), viewing habits, and other characteristics of the target audience. Pre-acquired skin conductance response data and electrocardiogram data are received. This data must come from a subject group that matches the target audience profile and must include the subject's basic information, such as age, gender, and viewing preferences. The subject information is matched and verified against the target audience profile to ensure that the received data accurately reflects the emotional responses of the target audience and avoids data distortion due to audience mismatch.

[0116] S702. Perform signal quality assessment on the pre-recorded skin conductance response data and electrocardiogram data, and generate quality assessment results.

[0117] Optionally, signal quality assessment indicators are used to perform quality checks on the pre-recorded skin conductance response data and electrocardiogram (ECG) data. For skin conductance response data, the assessment indicators include signal-to-noise ratio (the ratio of effective signal to noise), baseline stability (the range of signal fluctuation during the resting period), and signal integrity (the degree of absence of missing or interrupted data). For ECG data, the assessment indicators include RR interval integrity (the proportion of RR intervals without abnormal intervals), signal-to-noise level (the intensity of power line interference and motion artifacts), and the continuity of heart rate data. The specific values ​​of the above indicators are calculated and compared with preset signal quality standards to generate quality assessment results, clarifying the quality level of each data segment, indicating whether it is qualified or unqualified.

[0118] S703. Based on the quality assessment results, non-compliant data are removed to generate valid multimodal physiological data; the valid multimodal physiological data is used for the second segment statistical processing.

[0119] Optionally, based on the signal quality assessment results, qualified audience skin conductance response data and electrocardiogram (ECG) data are selected. Data with unqualified quality, such as excessively low signal-to-noise ratios in skin conductance data or numerous abnormal RR intervals in ECG data, are directly discarded to avoid poor-quality data affecting the accuracy of emotion curve generation. The selected qualified audience skin conductance response data and ECG data are then linked and integrated according to participant number, viewing time period, and timestamp to ensure that the two types of data from the same participant within the same time period can correspond and match, generating effective multimodal physiological data and providing high-quality data input for the accurate generation of the emotion curve for the film under test.

[0120] In the above embodiment, pre-recorded skin conductance response data and electrocardiogram data matched with the target audience profile of the film under test are received. Signal quality assessment is then performed on both types of data to generate a quality assessment result. Finally, based on the quality assessment result, unqualified data is eliminated to generate valid multimodal physiological data. This embodiment ensures the correlation between the data and the target audience's emotional response through audience profile matching, eliminates interference from inferior data through quality screening, and avoids distortion of the emotional curve due to mismatched data sources or quality issues. This effectively improves the accuracy of generating the emotional curve for the film under test and provides reliable data input for comparative analysis of emotional curves.

[0121] The aforementioned audience sentiment analysis methods, devices, equipment, and media based on multimodal data ensure high representativeness of benchmark films through multi-platform dual-dimensional screening and consistency verification. Synchronous processing and segmented statistics of electrodermal and electrocardiogram data enhance the comprehensiveness and accuracy of sentiment quantification. Curve difference calculation, smoothing filtering, and duration screening accurately locate continuous time periods of weak sentiment. Then, by associating video content features, structured optimization suggestions are generated. This effectively solves the problems of insufficient representativeness of benchmark samples, one-sided sentiment quantification, and ambiguous location of weak points in existing technologies. It significantly improves the scientific rigor and practicality of audience sentiment analysis for genre films, providing reliable technical support for film sentiment optimization.

[0122] To further illustrate the solutions of the embodiments of this application, a specific example is provided below.

[0123] This embodiment takes action films as the object and presents the complete process of sample selection, experiment execution, data processing, benchmark construction and application of the film to be tested, based on the core technical logic of the action film emotion assessment benchmark curve construction method.

[0124] 1. Screening of action movie sample films

[0125] Action film data was acquired from three platforms—Douban, Maoyan Professional Edition, and Lighthouse Professional Edition—using compliant data interfaces, encompassing both box office and critical reception. Based on the psychological triangular verification theory, screening rules were established. First, a top-down comparison was performed across the action film databases of each platform to initially select films with high box office rankings and high critical reception scores on each platform. Then, the initial screening results underwent multi-platform consistency verification, comparing the dual-dimensional rankings of the same film across all three platforms. Films that ranked highly in both dimensions across all three platforms were retained until a final 10 films meeting the requirements were selected, forming a high-commercial-value sample library of action films. This sample library is dynamically updated annually based on platform data for newly released action films, providing a representative sample foundation for constructing an emotional baseline curve.

[0126] 2. Experimental preparation and environmental control

[0127] The study population consisted of Chinese youth aged 18-24, representing the core audience for action films. The male-to-female ratio was set at 1:1, with 21 participants per film. This number was determined using G.power software to meet statistical requirements. Experimental equipment included bipolar silver / silver chloride electrodes, a biosignal amplifier with a sampling rate of at least 100Hz, physiological signal analysis software such as LabChart or AcqKnowledge, and a video synchronization trigger. The experimental environment was controlled at room temperature of 22±1℃ and humidity of 40-60%. A soundproof room or controlled laboratory was used to minimize external noise and light interference, avoiding environmental factors that could affect the quality of physiological signal acquisition.

[0128] 3. Experimental Procedure for Measuring Emotion Curves in Sample Films

[0129] The experimental procedure consisted of five stages, each strictly adhering to standardized procedures. The first stage, lasting 15 minutes, involved preparing the subjects. First, the skin at the electrode placement sites was cleaned. The index and middle fingers of the non-dominant hand were wiped with 75% alcohol wipes, and then 0.5% NaCl conductive paste was applied to ensure the electrode impedance was below 10kΩ. After starting the device, the subjects were allowed to stand still for 2 minutes to record the baseline noise level, which was required to be below 0.02μS. Simultaneously, instructions were given to the subjects to remain still to avoid interference with the signal from body movements and to maintain natural breathing while watching the video.

[0130] The second stage was a resting baseline recording lasting 1 minute, during which a dark video was played without sound output. During this process, the sensor was activated to record the skin conductance signal. Finally, the data from the last 3 minutes of this stage was selected as the baseline value to exclude initial adaptation interference.

[0131] The third stage is neutral stimulus calibration, lasting 1 minute, which involves playing a neutral video. The video content is a scene from the daily life of the participants. Taking college students as an example, the video content is set to be a daily campus life without any special plot. During this process, the stability of the SCR response is verified, and the response is required to have no significant fluctuations.

[0132] The fourth stage involved watching a 120-minute movie, using a video synchronization trigger to mark the start and end timestamps of the movie, and continuously recording the raw skin conductance signals.

[0133] The fifth stage is data export, which saves the collected raw data in Excel format. The file contains timestamps and synchronization markers to ensure that the data is traceable and meets the needs of subsequent processing.

[0134] 4. Skin conductance response data processing and analysis

[0135] The exported raw EEG data underwent multi-step processing. The first step was data preprocessing, starting with noise reduction using a 4th-order Butterworth bandpass filter algorithm with a passband frequency set to 0.05-2Hz. The algorithm was implemented by calling the `butter` and `filtfilt` functions from the `scipy.signal` library. First, the filter coefficients `b` and `a` were calculated using the formula: `b,a=butter(4,[0.05,2],btype='band',fs=100)`. Then, zero-phase filtering was applied to the raw data to obtain `filtered_data`, using the formula:

[0136] filtered_data = filtfilt(b, a, raw_data); then baseline correction is performed using the formula... ,in Let be the filtered signal value at time t. This represents the average value during the resting period.

[0137] The second step is SCR event detection. A threshold method is used to adapt to the long-term characteristics of the film. The amplitude threshold is set to ΔSCR≥0.05μS to avoid interference from small fluctuations, and the rise slope threshold is the rise rate>0.01μS / s. From the signal that meets the threshold requirements, four parameters are extracted: peak amplitude, latency, rise time, and recovery time. Peak amplitude is the maximum change value relative to the baseline, latency is the time from the start of stimulation to the start of SCR rise, rise time is the time from 10% to 90% of the peak value, and recovery time is the time from the peak value to 50%, forming an SCR event sequence.

[0138] The third step is segmented statistical analysis, which divides the 120-minute film into 24 five-minute segments. During the division process, the segment length is finely adjusted according to the rhythm characteristics of action films to avoid interrupting the plot. For each segment, the SCR event frequency, average amplitude, and cumulative response area are calculated. The SCR event frequency is measured in times per minute, and the cumulative response area is calculated using the integral formula ∫SCR(t)dt to reflect the sustained arousal level within the segment.

[0139] The fourth step is dynamic trend analysis, which uses the sliding window method with a window width of 3 minutes and a step size of 30 seconds. The moving average and standard deviation of the mean within the window are calculated, and the standard deviation is used to reflect the intensity of emotional fluctuations. The t-test is used to analyze the significant difference between the signal within the window and the signal in the non-event period, with the significance level set at p<0.05.

[0140] 5. Construction of the emotional baseline curve for action films

[0141] The results of the Galvanic Skin Response (GSR) processing of 10 action films in the sample library were statistically analyzed. The average cumulative response area of ​​each 5-minute segment of each film was calculated. The mean of the corresponding segments of the 10 films was then taken and concatenated in chronological order to form the emotional baseline curve of the action films. At the same time, the overall mean GSR of the 10 films was calculated. This mean was used as the baseline threshold for the emotional arousal of the action films. The baseline threshold was adjusted synchronously with the dynamic updates of the sample library to ensure that the baseline curve always conforms to the current emotional characteristics and market demand of the action films.

[0142] 6. Emotional Analysis and Optimization Application of Action Films Under Test

[0143] Taking an unreleased action film (Part A) as an example, the above experimental procedure and data processing steps were repeated. First, skin conductance data were collected from 21 target subjects who watched the film. An emotional curve for the film was generated using the same segmented statistical and trend analysis methods. The emotional curve was compared with the baseline emotional curve for action films. If the average cumulative response area for a certain continuous time period was found to be lower than the baseline threshold, that time period was identified as a weak segment for emotional arousal. A subjective questionnaire was designed for this weak segment, focusing on the subjects' evaluation of the segment's plot tightness and visual appeal. The reasons for the weakness were analyzed by combining physiological data and questionnaire results. Based on the analysis results, the segment of the film was edited and adjusted. After adjustment, the experiment and data processing were repeated to verify whether the emotional arousal level improved until the baseline curve requirements were met.

[0144] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0145] Based on the same inventive concept, this application also provides an apparatus for implementing the audience sentiment analysis method based on multimodal data as described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or more embodiments of the audience sentiment analysis apparatus based on multimodal data provided below can be found in the limitations of the audience sentiment analysis method based on multimodal data described above, and will not be repeated here.

[0146] In one exemplary embodiment, such as Figure 3 As shown, a multimodal data-based audience sentiment analysis device 10 is provided, comprising:

[0147] The historical film filtering module 11 is used to filter the historical film library according to preset filtering conditions to obtain multiple benchmark films of the target type.

[0148] The baseline curve generation module 12 is used to perform first segmented statistical processing on audience skin conductance data and electrocardiogram data corresponding to multiple baseline films to generate a genre film emotion baseline curve.

[0149] The test curve generation module 13 is used to perform second segmented statistical processing based on the audience's skin conductance response data and electrocardiogram data corresponding to the test film to generate the test film's emotion curve.

[0150] The emotion comparison module 14 is used to compare the emotion curve of the film under test with the emotion benchmark curve of the genre film and identify the continuous time period when the value is consistently lower than the emotion benchmark curve of the genre film.

[0151] The results output module 15 is used to generate corresponding audience sentiment analysis results based on continuous time periods; the audience sentiment analysis results are used to characterize the weak links in the emotional arousal of the film under test.

[0152] In one embodiment, the baseline curve generation module includes:

[0153] The SCR event detection unit is used to detect skin conductance response events in audience skin conductance response data and generate SCR event sequences.

[0154] The HRV feature extraction unit is used to perform heart rate variability analysis on electrocardiogram data and generate HRV feature sequences.

[0155] The time alignment unit is used to perform time alignment processing on the SCR event sequence and HRV feature sequence according to the film timeline to generate synchronized physiological feature data;

[0156] The segmented averaging unit is used to perform segmented averaging on synchronous physiological characteristic data to generate a genre film emotion baseline curve.

[0157] In one embodiment, the SCR event detection unit includes:

[0158] The bandpass filter subunit is used to perform bandpass filtering on the audience's skin conductance response data to generate a filtered signal.

[0159] The baseline correction subunit is used to perform baseline correction processing on the filtered signal to generate a clean skin electrical signal.

[0160] The peak filtering subunit is used to filter out peak points that are greater than preset amplitude thresholds and rise slope thresholds based on clean skin electrical signals, and generate a candidate SCR event set.

[0161] The artifact filtering subunit is used to perform motion artifact filtering on the candidate SCR event set and generate an SCR event sequence.

[0162] In one embodiment, the emotion comparison module includes:

[0163] The difference calculation unit is used to perform point-by-point difference calculation between the emotional baseline curve of the genre film and the emotional curve of the film to be tested, and generate an emotional difference curve.

[0164] The smoothing filter unit is used to perform sliding window averaging filtering on the sentiment difference curve to generate a smooth sentiment difference curve.

[0165] The low-arousal marking unit is used to mark a sequence of continuous data points in the smoothed emotion difference curve whose values ​​are lower than a preset negative threshold as candidate low-arousal segments.

[0166] The duration filtering unit is used to filter out segments from candidate low-wake segments whose duration exceeds a preset duration threshold, and generate continuous time periods.

[0167] In one embodiment, the result output module further includes:

[0168] The timestamp extraction unit is used to extract the corresponding timestamp information from the metadata of the video under test based on a continuous time period.

[0169] The content feature analysis unit is used to analyze the video content features of the film under test based on timestamp information, and generate plot type tags, dynamic performance indicators and audio intensity indicators.

[0170] The suggestion reasoning unit is used to reason about modifications based on plot type tags, visual dynamism indicators, and audio intensity indicators, and generate a structured optimization suggestion report.

[0171] In one embodiment, the historical video filtering module includes:

[0172] The box office and word-of-mouth data acquisition unit is used to obtain film box office data and word-of-mouth rating data from multiple professional film review platforms;

[0173] The threshold filtering unit is used to perform preliminary filtering of films in the historical film library that exceed the preset box office ranking threshold and word-of-mouth rating threshold, and generate a set of candidate films.

[0174] The consistency verification unit is used to perform multi-platform ranking consistency verification on the candidate film set, and generates multiple benchmark films from the candidate films that pass the verification.

[0175] In one embodiment, the curve generation module further includes:

[0176] The pre-recorded data receiving unit is used to receive pre-recorded skin conductance response data and electrocardiogram data that are matched with the target audience profile;

[0177] The signal quality assessment unit is used to assess the signal quality of pre-recorded skin conductance response data and electrocardiogram data, and generate quality assessment results.

[0178] The data removal unit is used to remove unqualified data based on the quality assessment results and generate valid multimodal physiological data; the valid multimodal physiological data is used for the second segment statistical processing.

[0179] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the audience sentiment analysis method based on multimodal data as described above.

[0180] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the audience sentiment analysis method based on multimodal data as described above.

[0181] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0182] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A method for audience sentiment analysis based on multimodal data, characterized in that, The method includes: The historical film library is filtered according to preset filtering criteria to obtain multiple benchmark films of the target type; Based on the audience's skin conductance response data and electrocardiogram data corresponding to the multiple benchmark films, a first segment statistical processing is performed to generate a genre film emotion benchmark curve. The second segment statistical processing is performed based on the audience's skin conductance response data and electrocardiogram data corresponding to the film under test to generate the emotional curve of the film under test. The emotional curve of the film under test is compared with the emotional baseline curve of the genre film to identify continuous time periods in which the value is consistently lower than the emotional baseline curve of the genre film. Based on the continuous time period, corresponding audience sentiment analysis results are generated; the audience sentiment analysis results are used to characterize the weak links in the emotional arousal of the film under test.

2. The method according to claim 1, characterized in that, The first segmentation statistical processing, based on audience skin conductance data and electrocardiogram data corresponding to the multiple benchmark films, generates a genre film emotional benchmark curve, including: The audience's electroskin response data were analyzed to detect electroskin conductance response events and generate SCR event sequences. Heart rate variability analysis was performed on the electrocardiogram data to generate HRV feature sequences; Based on the film timeline, the SCR event sequence and the HRV feature sequence are time-aligned to generate synchronized physiological feature data; The synchronous physiological characteristic data are processed by segmented averaging to generate the emotional baseline curve of the type of film.

3. The method according to claim 2, characterized in that, The step of detecting skin conductance response events in the audience's skin conductance response data and generating SCR event sequences includes: The audience's skin conductance response data were bandpass filtered to generate a filtered signal. The filtered signal is subjected to baseline correction processing to generate a clean skin electrical signal; Based on the clean skin electrical signals, peak points that are greater than preset amplitude thresholds and rise slope thresholds are selected to generate a candidate SCR event set. Motion artifact filtering is performed on the candidate SCR event set to generate the SCR event sequence.

4. The method according to claim 1, characterized in that, The step of comparing the emotional curve of the film under test with the emotional baseline curve of the genre film to identify continuous time periods in which the value is consistently lower than the emotional baseline curve of the genre film includes: A point-by-point difference calculation is performed between the emotional baseline curve of the aforementioned type of film and the emotional curve of the film to be tested to generate an emotional difference curve. The emotion difference curve is subjected to sliding window averaging filtering to generate a smooth emotion difference curve; The sequence of consecutive data points in the smoothed emotion difference curve whose values ​​are lower than a preset negative threshold is marked as a candidate low arousal segment. From the candidate low wake-up segments, segments whose duration exceeds a preset duration threshold are selected to generate the continuous time period.

5. The method according to claim 1, characterized in that, After generating the corresponding audience sentiment analysis results based on the continuous time period, the process further includes: Based on the continuous time period, the corresponding timestamp information is extracted from the metadata of the film to be tested; Based on the timestamp information, the video content features of the film to be tested are analyzed to generate plot type tags, dynamic performance indicators and audio intensity indicators. Based on the aforementioned plot type tags, visual dynamism indicators, and audio intensity indicators, modification suggestions are derived, and a structured optimization suggestion report is generated.

6. The method according to claim 1, characterized in that, The process of filtering the historical film library according to preset filtering conditions to obtain multiple benchmark films of the target type includes: Obtain film box office data and audience ratings from multiple professional film review platforms; Films in the historical film library that exceed preset box office ranking thresholds and word-of-mouth rating thresholds are initially screened to generate a candidate film set; The candidate film set is subjected to multi-platform ranking consistency verification, and the candidate films that pass the verification are used to generate the multiple benchmark films.

7. The method according to claim 1, characterized in that, Before generating the emotional curve of the film under test by performing the second segmented statistical processing based on the audience's skin conductance response data and electrocardiogram data corresponding to the film under test, the process also includes: Receive pre-recorded skin conductance data and electrocardiogram data that match the profile of the target movie-going audience; Signal quality assessment is performed on the pre-recorded skin conductance data and electrocardiogram data to generate quality assessment results; Based on the quality assessment results, non-compliant data are removed to generate valid multimodal physiological data; the valid multimodal physiological data is used for the second segmented statistical processing.

8. A device for analyzing audience sentiment based on multimodal data, characterized in that, The device includes: The historical film filtering module is used to filter the historical film library according to preset filtering conditions to obtain multiple benchmark films of the target type; The baseline curve generation module is used to perform first segmented statistical processing on audience skin conductance data and electrocardiogram data corresponding to the multiple baseline films to generate a genre film emotion baseline curve. The test curve generation module is used to perform second-segment statistical processing based on the audience's skin conductance response data and electrocardiogram data corresponding to the test film to generate the test film's emotion curve. The emotion comparison module is used to compare the emotion curve of the film under test with the emotion benchmark curve of the genre film, and identify the continuous time period when the value is consistently lower than the emotion benchmark curve of the genre film. The results output module is used to generate corresponding audience sentiment analysis results based on the continuous time period; the audience sentiment analysis results are used to characterize the weak links in the emotional arousal of the film under test.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.