Audio-visual language audience response evaluation system and method based on multi-view synchronous perception and data fusion

The audiovisual language audience response evaluation system, which integrates multi-perspective synchronous perception and data fusion, solves the problem of systematic measurement and optimization of audience responses from multiple angles. It realizes unified analysis of multimodal data and content optimization suggestions, thereby improving the accuracy of evaluation results and practical application capabilities.

CN122002086APending Publication Date: 2026-05-08LIAONING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING UNIVERSITY
Filing Date
2026-02-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies lack systematic multi-angle comparisons, have a single dimension of perceptual data collection, have fragmented data collection and analysis processes, and lack content-level feedback mechanisms, making it difficult to truly reflect audience reactions in diverse viewing environments. In particular, they lack quantitative research and optimization suggestions from non-standard perspectives.

Method used

An audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion is constructed, including an adjustable viewing platform, a synchronous perception acquisition module, a multimodal data fusion and analysis module, and an adaptive audiovisual language optimization module, to achieve multi-angle data acquisition, unified timeline analysis, and content optimization suggestions.

Benefits of technology

It enables system comparison at horizontal and tilt angles, integrates eye movement, physiological signals and subjective feedback, improves the accuracy of assessment results and practical application capabilities, and supports optimization suggestions in fields such as film and television production and advertising design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122002086A_ABST
    Figure CN122002086A_ABST
Patent Text Reader

Abstract

The invention relates to an audio-visual language audience response evaluation system and method based on multi-view synchronous perception and data fusion. The system comprises an adjustable viewing angle watching platform, an audio-visual content playing module, a synchronous perception acquisition module, a subjective feedback acquisition module, a multi-modal data fusion analysis module and an evaluation result output module, eye movement data, physiological signals and behavior responses of audiences can be synchronously acquired in real time at different watching angles, and subjective feedback information is combined, so that the audiences can be evaluated in real time. And the perception intensity, the attention concentration degree, the cognitive load and the emotion titer of the audience are comprehensively evaluated. The system can be widely applied to the fields of audio-visual language content optimization, immersive medium design and audience experience research, and has the advantages of being flexible in view angle switching, high in data fusion degree, high in analysis precision and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of audiovisual communication, human-computer interaction, experimental psychology, physiological signal detection, and multimodal data fusion analysis. In particular, it relates to an audiovisual language audience response assessment system and method based on multi-view synchronous perception and data fusion. This system is widely applicable to audience research on audiovisual content, user experience optimization, media content creation feedback, virtual reality environment assessment, and intelligent display device adaptation, belonging to a comprehensive application system of multi-view media perception and effect measurement technology. Background Technology

[0002] With the diversification of media environments and the variety of terminal devices, the way audiovisual content is viewed is undergoing significant changes. Traditional frontal, horizontal viewing scenarios, primarily on televisions or computers, are gradually being replaced by tilted or non-standard viewing angles, such as in-vehicle entertainment systems, subway advertising screens, and reclining viewing on mobile devices. Against this backdrop, understanding the impact of different viewing angles on the cognitive psychological mechanisms of audience attention, perception, comprehension, and emotional responses has become a research hotspot in fields such as communication studies, psychology, and interaction design.

[0003] Current research on audience response largely focuses on a single perspective or ideal viewing environment, exploring the degree to which audiences accept certain information, such as social media marketing messages, accessibility information, and interactive films. Methods typically employed include subjective questionnaires, interviews, and basic behavioral observation. While these methods can reflect audience subjective feelings to some extent, they have the following limitations:

[0004] (1) Lack of systematic multi-angle comparison: Existing studies lack systematic measurement and comparison of audience reactions under different physical perspectives such as horizontal and tilted, making it difficult to truly reproduce diverse viewing environments.

[0005] (2) Single dimension of perception data collection: It is mainly based on subjective evaluation and lacks the support of objective data such as eye tracking, physiological signal monitoring, and facial expression recognition. The evaluation results are highly subjective and lack credibility.

[0006] (3) Fragmented data acquisition and analysis process: Even if some studies introduce multiple sensor devices, they often lack multimodal data fusion analysis under a unified time axis, resulting in fragmented analysis results and difficulty in forming a complete audience response model.

[0007] (4) Lack of content-level feedback mechanism: Existing technology makes it difficult to directly feed audience reaction data back to audiovisual content creators, thereby guiding the optimization of elements such as composition, editing, and sound, and lacks a practical application path.

[0008] Furthermore, there is currently no mature system to support quantitative research and optimization suggestions for "non-standard perspective viewing experience" in new media forms such as virtual reality and immersive interactive video.

[0009] Therefore, there is an urgent need to build a systematic solution that can cover multiple viewing angles, simultaneously collect multimodal physiological and behavioral data, integrate subjective and objective evaluations, and effectively output evaluation results, so as to meet the dual needs of in-depth understanding of audience response and content optimization support in the new media environment. Summary of the Invention

[0010] To address the aforementioned technical problems, this invention provides a system and method for evaluating the audience response to audiovisual language based on multi-view synchronous perception and data fusion.

[0011] The present invention is specifically achieved through the following technical solution: an audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion, including an information input part, an information collection part, an analysis and evaluation part, and an optimization suggestion part;

[0012] The information input section includes an adjustable viewing platform and an audio-visual content playback module;

[0013] The information acquisition section includes a synchronous sensing acquisition module and a subjective feedback acquisition module. The synchronous sensing acquisition module consists of an eye-tracking submodule, a synchronous marker recorder, and a physiological signal acquisition submodule.

[0014] The analysis and evaluation section includes a multimodal data fusion and analysis module and an evaluation result output module;

[0015] The optimization suggestions section includes an adaptive audiovisual language optimization module.

[0016] The adjustable viewing platform includes an adjustable viewing angle bracket, a tilt sensor module, and a tilt screen mounting interface. The viewing angle can be switched by electric drive or manual control to simulate various viewing angles and environments in real-world scenarios.

[0017] The audiovisual content playback module includes a content management unit, a synchronization control unit, and a playback control interface, used to display standardized or experimental audiovisual language materials; the content includes controllable screen composition, editing structure, and sound design parameters, and supports the import, editing, annotation, and timed playback of video clips.

[0018] The aforementioned synchronous sensing and acquisition module is used to collect physiological, behavioral, and cognitive data of the audience from different viewing angles in real time.

[0019] The eye-tracking submodule is based on an infrared or video eye tracker and records the gaze point, gaze duration, saccade path, and pupil diameter in real time. The eye-tracking submodule includes an infrared light source component, a high-speed camera, and a coordinate calculation unit.

[0020] The physiological signal acquisition submodule is used to monitor the physiological changes of the audience during viewing: including heart rate sensor, skin conductance sensor, and muscle signal acquisition interface;

[0021] The subjective feedback collection module is used to acquire the audience's subjective evaluation information on audiovisual content, including a questionnaire evaluation interface, an input interface, and a data preprocessing unit; the evaluation dimensions include visual clarity, plot comprehension, rhythm comfort, immersive experience, and emotional touch indicators, and automatically records timestamps.

[0022] The multimodal data fusion and analysis module synchronizes and standardizes the collected multi-source data, constructs a unified timeline, extracts core feature indicators, and performs statistical modeling and correlation analysis on the data; it supports cross-analysis of eye-tracking data, physiological signals, behavioral features, and subjective ratings, and establishes a mapping relationship between different viewing perspectives and audience perception responses; the multimodal data fusion and analysis module includes a multimodal data integration engine, a feature extraction unit, a model analysis unit, and a visualization module.

[0023] The evaluation result output module, based on a preset evaluation index system, transforms the fusion analysis results into visual charts and quantitative reports. The output content includes attention heatmaps, cognitive load curves, emotional state distribution maps, and comprehension score maps, and can be exported as PDF, CSV, or JSON formats. The evaluation result output unit includes a content comprehension score unit, an emotional response evaluation unit, and an output interface.

[0024] The adaptive audiovisual language optimization module, based on the evaluation results and according to the differences in audience performance from different angles, proposes adjustment suggestions for picture composition, editing density, rhythm design, and sound layout, and supports automatic content adjustment and recommendation. The adaptive audiovisual language optimization module includes a composition optimization suggestion unit, an editing rhythm suggestion unit, and an audio-visual synchronization adjustment module.

[0025] The evaluation method for the audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion includes the following steps:

[0026] Step 1) Set the experimental conditions, configure the required viewing angle, and adjust the viewing platform to the corresponding angle;

[0027] Step 2) Load the audiovisual content to be evaluated into the audiovisual content playback module and start the playback process, while activating the synchronous perception acquisition module.

[0028] Step 3) During the playback of audiovisual content, the eye-tracking submodule records the viewer's gaze behavior in real time, and the physiological acquisition submodule continuously collects heart rate, skin conductance, and electroencephalogram (EEG) indicators.

[0029] Step 4) After the content playback ends, the audience completes their subjective evaluation through the subjective feedback collection module. The system numbers, aligns, and preprocesses all the data.

[0030] Step 5) The multimodal data fusion and analysis module performs feature extraction, time series analysis and statistical modeling on various types of data, and outputs key indicators, including average fixation rate, peak maximum cognitive load, heart rate variation range, average skin conductance, α / β wave frequency distribution, subjective understanding score and emotional valence score.

[0031] Step 6) Analyze the differences in the physiological and cognitive responses of the audience under different perspectives, and construct a perspective-perception relationship model;

[0032] Step 7) The evaluation results output module outputs a complete audience response evaluation report, including visualization charts, feature curves and decision support suggestions, and pushes it to researchers or content production systems through the platform.

[0033] The evaluation methods described include Support Vector Machine (SVM), Convolutional Neural Network (CNN), Time Series Clustering, Principal Component Analysis (PCA), and Random Forest machine learning methods, which are used to classify, predict, or model the correlation of audience response data under perspective variables.

[0034] The beneficial effects of this invention are as follows:

[0035] 1. By incorporating physical viewing perspective as a system variable into audience research, a systematic comparison can be achieved under horizontal and tilt angles, breaking the singularity of traditional audience research environment settings.

[0036] 2. By integrating four data channels—eye movement, physiological signals, behavioral recognition, and subjective feedback—an assessment mechanism that combines subjective and objective perspectives and provides multimodal verification is constructed, overcoming the problem of single data dimensions in previous assessments.

[0037] 3. A unified time axis aligned data processing strategy and a scalable evaluation model framework are proposed, enabling data from different modalities to be analyzed and predicted on a unified scale, thereby improving the accuracy of the evaluation results.

[0038] 4. It possesses the application capability for practical creation, and the system evaluation results can be directly transformed into audiovisual language adjustment suggestions, providing decision-making basis for film and television production, advertising design, immersive content development, etc., significantly improving creative efficiency and user experience satisfaction; 5. It is applicable to various real-world environments such as in-vehicle scenarios, mobile viewing, public media terminals, VR / AR scenarios, etc., and has broad adaptability and good industrialization prospects.

[0039] Therefore, this invention constructs a scientific, systematic, and scalable audience response evaluation system for the future audiovisual content dissemination ecosystem from multiple levels, including technical structure, acquisition process, data fusion, and application feedback. It has outstanding theoretical innovation value and practical application potential. Attached Figure Description

[0040] Figure 1 : A schematic diagram of the overall structure of the system of the present invention.

[0041] Figure 2 : Flowchart of multi-view audiovisual audience response data acquisition and fusion processing of the present invention. Detailed Implementation

[0042] To better understand the present invention, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. The present invention proposes an audiovisual language audience response evaluation system and method applicable to multi-viewpoint conditions. It can achieve multimodal synchronous acquisition of user attention, physiological responses, and subjective feedback from horizontal to tilted viewpoints, and perform data fusion analysis, thereby systematically evaluating the differences in the effects of audiovisual language elements under different viewing angles. It has broad application value and technical feasibility.

[0043] like Figure 1 As shown, an audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion includes an information input part, an information collection part, an analysis and evaluation part, and an optimization suggestion part.

[0044] The information input section includes an adjustable viewing platform and an audio-visual content playback module.

[0045] The adjustable viewing platform includes an adjustable viewing angle bracket, a tilt sensor module, and a tilt screen mounting interface. Viewing angle switching is achieved through electric drive or manual control. This platform simulates the physical environment of a user viewing audiovisual content from different angles, including horizontal viewing angles (such as looking directly at a desktop monitor) and tilted viewing angles (such as looking up at a ceiling advertising screen or lying on one's side to view a mobile phone screen). The platform is equipped with an angle adjustment device and sensors, enabling precise setting and feedback of the current viewing angle, which is then input as a system variable for subsequent analysis.

[0046] The audiovisual content playback module includes a content management unit, a synchronization control unit, and a playback control interface. It is used to display standardized or experimental audiovisual language materials and is responsible for uniformly controlling the presentation time, sequence, and rhythm of the audiovisual content. The content includes controllable frame composition, editing structure, and sound design parameters. It supports the import, editing, annotation, and timed playback of video clips, maintaining time synchronization. This module supports loading video content in various formats and has timecode synchronization functionality to ensure accurate alignment with the eye-tracking and physiological data acquisition system on the timeline, facilitating subsequent event-driven response analysis.

[0047] II. The information acquisition section includes a synchronous sensing acquisition module and a subjective feedback acquisition module. The synchronous sensing acquisition module consists of an eye-tracking sub-module, a synchronous marker recorder, and a physiological signal acquisition sub-module.

[0048] Synchronous perception and acquisition module: This module records the user's eye movement data (including fixation point, fixation time, saccade path, etc.) and physiological data (such as skin conductance, heart rate, electromyography, electroencephalography, etc.) in real time to capture changes in attention and emotional arousal. This module supports multiple sensor inputs and features a unified timestamp protocol to ensure the consistency and integrity of multi-source data.

[0049] The eye-tracking submodule is based on an infrared or video eye tracker and records the gaze point, gaze duration, saccade path, and pupil diameter in real time. The eye-tracking submodule includes an infrared light source component, a high-speed camera, and a coordinate calculation unit.

[0050] The physiological signal acquisition submodule is used to monitor the physiological changes of the audience during viewing: including heart rate sensor, skin conductance sensor, and muscle signal acquisition interface;

[0051] The subjective feedback collection module is used to acquire audience subjective evaluation information of audiovisual content, including a questionnaire evaluation interface, an input interface, and a data preprocessing unit. Evaluation dimensions include visual clarity, plot comprehension, pacing comfort, immersive experience, and emotional engagement indicators, and it automatically records timestamps. In use, after completing a viewing task, users are guided to complete a subjective feedback task, including cognitive load perception, self-emotional evaluation, and information acceptance rating. Data is collected through electronic questionnaires or touch interaction. This module can also insert feedback on content viewing interruptions (such as post-keyframe evaluation) to capture more granular psychological reactions.

[0052] III. The analysis and evaluation section includes a multimodal data fusion and analysis module and an evaluation result output module;

[0053] The multimodal data fusion and analysis module synchronizes and standardizes the collected multi-source data, constructs a unified timeline, extracts core feature indicators, and performs statistical modeling and correlation analysis on the data. It supports cross-analysis of eye-tracking data, physiological signals, behavioral characteristics, and subjective ratings, establishing a mapping relationship between different viewing perspectives and audience perceptual responses. This module includes a multimodal data integration engine, a feature extraction unit, a model analysis unit, and a visualization module. It can use a combination of traditional statistical methods (such as analysis of variance and linear regression) and machine learning methods (such as support vector machines and random forests) to model and extract key differential features from different perspectives.

[0054] The evaluation result output module, based on a preset evaluation index system, transforms the fusion analysis results into visual charts and quantitative reports. Output content includes attention heatmaps, cognitive load curves, emotional state distribution maps, and comprehension score maps, exported in PDF, CSV, or JSON format. The evaluation result output unit includes a content comprehension score unit, an emotional response evaluation unit, and an output interface. For example, it can output conclusions such as user attention shift, increased emotional arousal, or decreased comprehension at a certain tilt angle for a specific composition.

[0055] IV. The optimization suggestions section includes an adaptive audiovisual language optimization module.

[0056] The adaptive audiovisual language optimization module, based on the evaluation results and considering the differences in audience performance from different angles, proposes adjustments to screen composition, editing density, pacing, and sound layout. It supports automatic content adjustment and recommendation. The module includes a composition optimization suggestion unit, an editing pacing suggestion unit, and an audio-visual synchronization adjustment module. Based on the evaluation results, the system can generate corresponding audiovisual language adjustment suggestions, such as adjusting the screen focus position, extending the duration of keyframes, optimizing editing pacing, or sound design, to enhance the viewing experience from multiple angles.

[0057] Step 1) Construct a standardized experimental environment, set the target viewing angle (including but not limited to 0°, 30°, 60°, etc.) through an adjustable viewing angle platform, calibrate the viewing angle recognition and recording device, initialize each data acquisition module, complete system time synchronization and benchmark parameter calibration, and ensure that the system can accurately identify and record viewing angle parameters.

[0058] Step 2) Load the audiovisual content to be evaluated into the audiovisual content playback module, establish a unified playback timecode, and insert stimulus marker frames or specific analysis segments at the required time nodes for the corresponding analysis of subsequent key data segments.

[0059] Step 3) During the playback of audiovisual content, the multimodal synchronous perception and acquisition module is activated to collect audience eye movement data, physiological signals and behavioral feature data in real time, and simultaneously record playback timecode and viewing angle parameters to achieve unified time stamping of multi-source data.

[0060] Step 4) After the playback ends, guide the audience to complete the structured evaluation questionnaire through the subjective feedback collection module, collect subjective feedback data including emotional valence score, comprehension score, cognitive load self-assessment and preference selection, and assign a unified number to the objective data collected.

[0061] Step 5) Perform noise reduction, outlier removal, data standardization, and cross-modal time alignment on the collected eye-tracking data, physiological data, behavioral data, and subjective data to form a unified multimodal analysis dataset;

[0062] Step 6) Perform feature extraction and time series analysis on the unified dataset, construct a statistical model or machine learning model, extract key features including average fixation concentration rate, peak maximum cognitive load, heart rate variation range, mean skin conductance, EEG α / β wave frequency distribution and subjective rating indicators, and compare the differences of various indicators under different perspective conditions.

[0063] Step 7) Based on the analysis results under different perspective conditions, construct a correlation model between perspective parameters, audiovisual language element features and audience physiological cognitive response, extract audiovisual language elements such as picture composition, editing rhythm and sound effect layout that have a significant impact, and form a perspective-perception relationship model.

[0064] Step 8) Generate a complete evaluation report through the evaluation results output module, which includes visualization charts, gaze hot zone distribution, physiological curve changes and three-dimensional relationship model analysis results. Based on historical data and rule models, provide targeted optimization suggestions for content paragraphs that perform poorly from a specific perspective, thus achieving a closed loop of content improvement.

[0065] Example 1:

[0066] Three viewing conditions were set: horizontal viewing angle (0°), slightly tilted viewing angle (30°), and moderately tilted viewing angle (60°). Video content of the same length (such as advertising clips or plot clips) was played for the subjects. Eye movement indicators such as fixation duration and fixation distribution density, physiological indicators such as skin conductance and heart rate variability were collected, and subjective cognitive load scores and emotional responses were recorded.

[0067] Based on multimodal data, this system determines that there is a coupling effect between editing frequency and the center of gravity of the frame on attention shift, especially at tilted viewpoints. Therefore, the system can suggest reducing editing frequency or adjusting the placement of key information to the lower center of the frame to adapt to different visual habits from various perspectives.

[0068] The system structure of this invention is highly modular, with all modules capable of independent deployment or integration into a complete platform. Sensor types and audiovisual content formats can be flexibly configured according to experimental objectives, making it widely applicable to the following typical scenarios:

[0069] (1) In advertising and communication research, the differences in advertising effects under different devices (such as outdoor angled screens and vehicle screens) were tested;

[0070] (2) Pre-production content testing for film and television production, used to evaluate the audience response to the editing scheme under multiple viewing postures;

[0071] (3) User experience evaluation of virtual reality or augmented reality devices, simulating multi-angle interactive scenarios;

[0072] (4) Test the cognitive effect of educational or medical promotional videos to ensure the accuracy of information delivery.

[0073] In summary, by introducing multi-perspective variables and integrating eye-tracking, physiological, and subjective feedback data, this invention constructs a well-structured, fully functional, and highly scalable audience perception assessment system. It solves the technical problem that existing technologies cannot accurately reflect the impact of changes in viewing posture on audience reactions, and has significant innovation and practical value.

Claims

1. An audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion, characterized in that: It includes an information input section, an information collection section, an analysis and evaluation section, and an optimization suggestion section; The information input section includes an adjustable viewing platform and an audio-visual content playback module; The information acquisition section includes a synchronous sensing acquisition module and a subjective feedback acquisition module. The synchronous sensing acquisition module consists of an eye-tracking submodule, a synchronous marker recorder, and a physiological signal acquisition submodule. The analysis and evaluation section includes a multimodal data fusion and analysis module and an evaluation result output module; The optimization suggestions section includes an adaptive audiovisual language optimization module.

2. The audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion according to claim 1, characterized in that: The adjustable viewing platform includes an adjustable viewing angle bracket, a tilt sensor module, and a tilt screen mounting interface. The viewing angle can be switched by electric drive or manual control to simulate various viewing angles and environments in real-world scenarios. The audiovisual content playback module includes a content management unit, a synchronization control unit, and a playback control interface, used to display standardized or experimental audiovisual language materials; the content includes controllable screen composition, editing structure, and sound design parameters, and supports the import, editing, annotation, and timed playback of video clips.

3. The audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion according to claim 1, characterized in that: The aforementioned synchronous sensing and acquisition module is used to collect physiological, behavioral, and cognitive data of the audience from different viewing angles in real time. The eye-tracking submodule is based on an infrared or video eye tracker and records the gaze point, gaze duration, saccade path, and pupil diameter in real time. The eye-tracking submodule includes an infrared light source component, a high-speed camera, and a coordinate calculation unit. The physiological signal acquisition submodule is used to monitor the physiological changes of the audience during viewing: it includes a heart rate sensor, a skin conductance sensor, and a muscle signal acquisition interface.

4. The audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion according to claim 1, characterized in that: The subjective feedback collection module is used to acquire the audience's subjective evaluation information on audiovisual content, including a questionnaire evaluation interface, an input interface, and a data preprocessing unit; the evaluation dimensions include visual clarity, plot comprehension, rhythm comfort, immersive experience, and emotional touch indicators, and automatically records timestamps.

5. The audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion according to claim 1, characterized in that: The multimodal data fusion and analysis module synchronizes and standardizes the collected multi-source data, constructs a unified timeline, extracts core feature indicators, and performs statistical modeling and correlation analysis on the data; it supports cross-analysis of eye-tracking data, physiological signals, behavioral features, and subjective ratings, and establishes a mapping relationship between different viewing perspectives and audience perception responses; the multimodal data fusion and analysis module includes a multimodal data integration engine, a feature extraction unit, a model analysis unit, and a visualization module.

6. The audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion according to claim 1, characterized in that: The evaluation result output module, based on a preset evaluation index system, transforms the fusion analysis results into visual charts and quantitative reports. The output content includes attention heatmaps, cognitive load curves, emotional state distribution maps, and comprehension score maps, and can be exported as PDF, CSV, or JSON formats. The evaluation result output unit includes a content comprehension score unit, an emotional response evaluation unit, and an output interface.

7. The audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion according to claim 1, characterized in that: The adaptive audiovisual language optimization module, based on the evaluation results and according to the differences in audience performance from different angles, proposes adjustment suggestions for picture composition, editing density, rhythm design, and sound layout, and supports automatic content adjustment and recommendation. The adaptive audiovisual language optimization module includes a composition optimization suggestion unit, an editing rhythm suggestion unit, and an audio-visual synchronization adjustment module.

8. An evaluation method using an audiovisual language audience response evaluation system based on multi-view synchronous perception and data fusion as described in any one of claims 1-7, characterized in that, The steps are as follows: Step 1) Set the experimental conditions, configure the required viewing angle, and adjust the viewing platform to the corresponding angle; Step 2) Load the audiovisual content to be evaluated into the audiovisual content playback module and start the playback process, while activating the synchronous perception acquisition module. Step 3) During the playback of audiovisual content, the eye-tracking submodule records the viewer's gaze behavior in real time, and the physiological acquisition submodule continuously collects heart rate, skin conductance, and electroencephalogram (EEG) indicators. Step 4) After the content playback ends, the audience completes their subjective evaluation through the subjective feedback collection module. The system numbers, aligns, and preprocesses all the data. Step 5) The multimodal data fusion and analysis module performs feature extraction, time series analysis and statistical modeling on various types of data, and outputs key indicators, including average fixation rate, peak maximum cognitive load, heart rate variation range, average skin conductance, α / β wave frequency distribution, subjective understanding score and emotional valence score. Step 6) Analyze the differences in the physiological and cognitive responses of the audience under different perspectives, and construct a perspective-perception relationship model; Step 7) The evaluation results output module outputs a complete audience response evaluation report, including visualization charts, feature curves and decision support suggestions, and pushes it to researchers or content production systems through the platform.

9. The evaluation method according to claim 8, characterized in that: The evaluation methods described include Support Vector Machine (SVM), Convolutional Neural Network (CNN), Time Series Clustering, Principal Component Analysis (PCA), and Random Forest machine learning methods, which are used to classify, predict, or model the correlation of audience response data under perspective variables.