Multimedia display method and system

By acquiring data on exhibition space and audience characteristics, monitoring behavioral data in real time, and dynamically adjusting multimedia display content, the system addresses the issues of insufficient flexibility and interactivity in multimedia display systems, thereby improving the display effect and audience participation.

CN121442136APending Publication Date: 2026-01-30广州市美术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511523650.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing multimedia display systems suffer from poor information display flexibility and insufficient interactivity, failing to dynamically adjust the displayed content according to audience needs, resulting in subpar display effects.

Method used

By acquiring the initial environmental parameters of the exhibition space and the characteristic data of the audience, the initial display parameters of the multimedia content are generated. The audience behavior data is monitored in real time, the dynamic deviation between the current display effect and the expected goal is calculated, and content adaptation processing is performed to dynamically adjust the presentation form, interaction method and information density of the multimedia content.

Benefits of technology

It enables personalized adaptation and dynamic adjustment of multimedia display systems, enhances audience immersion and participation, improves display effects and information dissemination efficiency, and solves the problems of poor flexibility and insufficient interactivity in traditional display systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121442136A_ABST
    Figure CN121442136A_ABST
Patent Text Reader

Abstract

The invention discloses a multimedia display method and system, and relates to the technical field of multimedia information processing, and the method comprises the steps: obtaining an initial environment parameter of a display space and audience group feature data, and generating an initial display parameter of multimedia content; in the process of displaying the multimedia content according to the initial display parameters, monitoring audience behavior data, and calculating a dynamic deviation between a current display effect and an expected target; performing content adaptation processing on the dynamic deviation to obtain dynamic adjustment parameters of multi-modal display; according to the dynamic adjustment parameter, calculating a content dynamic adaptation amount in the current display scene, and synchronously adjusting a presentation form, an interaction mode and information density of the multimedia content according to the content dynamic adaptation amount; the method and the device have the effects of improving the display effect of multimedia display and improving audience interactivity.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multimedia information processing, and in particular to a multimedia display method and system. BACKGROUND

[0002] At present, with the rapid development of information technology, multimedia display has become an important way of scientific and technological information dissemination, which can deliver complex information to the audience in a lively and intuitive form. The existing multimedia display technology plays an important role in improving audience experience and information dissemination efficiency.

[0003] The existing traditional multimedia display system usually relies on a single display device, such as a large screen display or a projector, combined with a simple interactive device for operation. This display method has the problems of poor information display flexibility, insufficient interactivity, and inability to dynamically adjust the display content according to the audience's needs, and therefore has the defect of poor display effect of multimedia display, which needs to be improved. SUMMARY

[0004] In order to improve the display effect of multimedia display and improve the audience interactivity, the present application provides a multimedia display method and system.

[0005] In the first aspect, the application aims to achieve the following technical solutions: A multimedia display method, comprising: Obtaining initial environmental parameters of a display space and audience group feature data, generating initial display parameters of multimedia content; monitoring audience behavior data in the process of displaying the multimedia content according to the initial display parameters, and calculating the dynamic deviation of the current display effect from the expected target; Performing content adaptation processing on the dynamic deviation to obtain dynamic adjustment parameters of multi-modal display; According to the dynamic adjustment parameters, calculating the content dynamic adaptation amount in the current display scene, and synchronously adjusting the presentation form, interactive mode and information density of the multimedia content according to the content dynamic adaptation amount.

[0006] By adopting the technical scheme, the environment parameters (such as spatial size / sound field distribution / illumination intensity) and the audience group features (such as age distribution / cognitive level / cultural background) are acquired to construct a multi-dimensional initial parameter matrix, the problem of poor adaptability of a single device in the traditional scheme is solved, the initial display parameters are preliminarily matched with the physical space and the audience group, the audience behavior data (such as gaze hotspot shift rate / interaction response delay / physiological signal fluctuation) are monitored to construct a dynamic mapping model of the display system and audience cognition, the traditional static deviation detection is upgraded to real-time dynamic evaluation, the display effect is adjusted according to the audience demand, and whether the current display meets the expected information dissemination target is evaluated through behavior data analysis (such as gaze time, movement trajectory, and interaction frequency).

[0007] In a preferred example of the present application, the content dynamic adaptation amount in the current display scene is calculated according to the dynamic adjustment parameters, and the presentation form, interactive mode and information density of the multimedia content are synchronously adjusted according to the content dynamic adaptation amount, specifically including: According to the dynamic adjustment parameters, the attention heat map corresponding to the real-time behavior characteristics of the audience and the content expected value corresponding to the initial display parameters are obtained; Based on the difference between the attention heat map and the content expected value, the dynamic adaptation amount of the multi-modal content is comprehensively calculated, including information level compression rate, interaction response delay threshold and sensory channel weight; According to the dynamic adaptation amount, the rendering resolution and interaction response rate of the multimedia content are adjusted, and the space-time consistency parameters of visual, auditory and tactile feedback are controlled through a multi-channel synchronous algorithm; According to the space-time consistency parameters, the adaptive pushing logic of the multimedia content is optimized, and dynamic calibration is formed with real-time feedback of the audience.

[0008] By adopting the technical scheme, the dynamic adaptation amount is calculated based on the difference between the attention heat map and the content expected value, the difference between the attention heat map and the expected value is quantified, the dynamic distribution of the weight of the visual / auditory / tactile modal is realized, the rendering resolution is adjusted in real time (such as from 4K to 1080P) according to the information level compression rate β∈{1, 2, 4}, the GPU load is reduced under the premise of maintaining the coherence of the content, the time delay difference of the visual, auditory and tactile is controlled through the multi-channel synchronization algorithm, and the perception fragmentation problem in the traditional scheme is effectively solved.

[0009] In a preferred example of the present application, the adaptive pushing logic of the multimedia content according to the spatio-temporal consistency parameter is optimized, and dynamic calibration is formed with real-time feedback of the audience, which further comprises: Obtaining the positioning trajectory and physiological sensing data of the audience in the exhibition space, and calculating the content adaptation degree of the next stage in combination with the environmental lighting parameter; According to the content adaptation degree, the narrative rhythm of the exhibition content is dynamically adjusted, and the combination strategy of the multi-modal content is optimized through the reinforcement learning model; Based on the historical data of the audience interaction operation, an audience cognitive model is constructed and a personalized content branch path is generated.

[0010] By adopting the technical scheme, the attention change trend of the audience is predicted based on the moving trajectory and physiological state (such as heart rate and eye movement frequency) of the audience; the visibility and comfort of the exhibition scene are dynamically evaluated in combination with factors such as environmental lighting, to provide a basis for content rhythm adjustment. The reinforcement learning model can continuously optimize the combination mode of the multi-modal content according to the audience response; the exhibition rhythm is adjusted to adapt to the emotional fluctuation of the audience, to prevent attention loss caused by too fast or too slow rhythm. The constructed audience cognitive model can understand the knowledge background, interest preference and acceptance ability of the audience.

[0011] In a preferred example of the present application, the personalized content branch path is calculated by formula (1), and formula (1) is as follows: B=α×E+(1-α)×I (1) Wherein, B represents the content branch weight, E represents the audience emotion entropy value, I represents the interaction depth index, and a is a dynamic adjustment coefficient (0.6≤a≤0.9); According to the device network state of the exhibition space, the real-time rendering load parameter of each terminal node is obtained, and the synchronization compensation amount of the content dynamic adaptation amount in the distributed system is calculated; According to the synchronization compensation amount, the priority strategy of multi-terminal content distribution is adjusted.

[0012] By adopting the technical scheme, the emotional entropy value E in the formula reflects the emotional state of the audience, and the interactive depth index I measures the participation degree; the dynamic adjustment coefficient a is between 0.6 and 0.9, the emotional factor is given a higher weight, the concept of emotion-driven display is embodied, the introduction of the synchronization compensation amount can cope with the differences of each terminal node in rendering load, network delay and the like; and the risk of inconsistent display content between terminals is effectively reduced.

[0013] In a preferred example of the present application: the priority strategy of the multi-terminal content distribution according to the synchronization compensation amount further comprises: A digital twin model of the display content is established to map the content state difference between the physical space and the virtual space in real time; According to the spatial distribution characteristics of the audience group, the content push area is dynamically divided and the network bandwidth allocation is optimized; The audience behavior data of the multi-terminal nodes is aggregated through a federated learning mechanism to continuously optimize the group adaptability of the content recommendation algorithm.

[0014] By adopting the technical scheme, the digital twin model can reflect the presentation state of the display content in the physical space in real time, which is convenient for finding content deviation, device anomaly and the like; the push area is dynamically divided according to the audience group distribution to avoid resource waste; the federated learning mechanism can aggregate the data of multiple terminals on the premise of protecting privacy; and the content recommendation algorithm is continuously optimized to make it more suitable for the behavior mode of the audience group.

[0015] In a preferred example of the present application: the monitoring of the audience behavior data and the calculation of the dynamic deviation of the current display effect and the expected target specifically comprise: The multi-modal interaction data stream of the audience in the display space is obtained, the multi-modal interaction data stream is input into a preset audience cognitive load prediction model, and an audience attention attenuation coefficient and an information reception effectiveness index are output; the entropy value change rate ΔH of the display content is dynamically calculated according to the attention attenuation coefficient, and a dynamic deviation calculation formula is constructed: ΔD = λ × ΔH + (1-λ) × T f (2) Wherein, ΔD is the dynamic deviation, λ is a cognitive attenuation weight coefficient (0.7≤λ≤0.95), T f is a multi-modal data fusion delay time; When ΔD exceeds a preset threshold, a multi-channel compensation mechanism is triggered, and the heat weight distribution of the content focus area is adjusted through an eye movement trajectory prediction algorithm.

[0016] By adopting the above technical solution, multimodal interactive data (such as voice, action, gaze, touch, etc.) are input into the cognitive load prediction model, which can determine in real time whether the audience is currently in a state of focused / distracted attention, thereby achieving a quantitative assessment of the audience's cognitive state. The entropy change rate ΔH is introduced to reflect the trend of content complexity changing over time; the cognitive decay weight λ (ranging from 0.7 to 0.95) emphasizes the impact of decreased audience attention on the presentation effect; and the multimodal data fusion delay time T... f This demonstrates the impact of system response delays on user experience.

[0017] In a preferred embodiment of this application: the content adaptation processing of the dynamic deviation to obtain dynamic adjustment parameters for multimodal display includes: A dynamic Bayesian network model of the display system is constructed, and the dynamic deviation ΔD is input into the network as the observation node. The posterior probability distribution of the parameter adjustment is calculated using a variational inference algorithm, and the core adjustment parameters are extracted. These core adjustment parameters include the sensory channel coordination coefficient α. ′ Content decoupling granularity β and interactive response elasticity γ; A dynamic adjustment matrix M = (m) is generated based on the core adjustment parameters. ij )_{3×3}, where: m 11 =α ′ ×exp(-β×ΔD) m 22 =γ×(1+ΔH / σ) m 33 = 1 / (1+e^{-ΔD}) By performing a nonlinear transformation on the multimodal content stream using matrix M, adaptive remapping of display parameters is achieved.

[0018] By adopting the above technical solutions, Dynamic Bayesian Network (DBN) is a probabilistic graphical model suitable for time series modeling; by inputting the dynamic bias ΔD as the observed variable into the network, the uncertainty and dependencies between variables in the presentation process can be captured; in the face of high-dimensional and uncertain presentation environments, traditional methods are difficult to solve for the optimal parameters quickly; using variational inference algorithms can significantly improve computational efficiency while ensuring accuracy; and it enables real-time estimation and adjustment of presentation parameters (such as sensory channel synergy coefficients, content decoupling granularity, and interactive response elasticity).

[0019] Secondly, the objective of this invention is achieved through the following technical solution: A multimedia display system for performing a multimedia display method as described above, the system comprising: A data collection module is configured to acquire initial environmental parameters of a display space and audience group feature data. A content generation module is configured to generate initial display parameters of multimedia content based on the initial environmental parameters and the audience group feature data. A display control module is configured to control a display device to output corresponding content during display of the multimedia content according to the initial display parameters. A behavior detection module is configured to monitor audience behavior data in real time. A dynamic adjustment module is configured to calculate a dynamic deviation between a current display effect and an expected target based on the audience behavior data, and perform content adaptation processing on the dynamic deviation to obtain dynamic adjustment parameters of multi-modal display. The display control module is further configured to calculate a content dynamic adaptation amount in a current display scenario based on the dynamic adjustment parameters, and synchronously adjust a presentation form, an interactive mode, and an information density of the multimedia content.

[0020] In a third aspect, the application aims to achieve the object by adopting the following technical solution: A display device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above-mentioned multimedia display method when executing the computer program.

[0021] In a fourth aspect, the application aims to achieve the object by adopting the following technical solution: A computer-readable storage medium stores a computer program, and the computer program implements the steps of the above-mentioned multimedia display method when executed by a processor.

[0022] In summary, the application includes at least one of the following beneficial technical effects: 1. A closed-loop display system based on environmental and audience feature perception is constructed; audience behavior data is monitored in real time and a dynamic deviation is calculated, content adaptation processing is performed based on the dynamic deviation, and finally the presentation form, interactive mode, and information density of the multimedia content are synchronously adjusted according to the dynamic adjustment parameters; the application effectively solves the problems of poor flexibility and insufficient interactivity of traditional display systems, realizes personalized adaptation and dynamic adjustment of display content, enhances audience immersion and participation, and improves overall display effect and information dissemination efficiency; 2. Each element in the matrix M is related to the dynamic deviation DD and the entropy change rate DH, which reflects the nonlinear characteristics of content adjustment. 11 represents intensity attenuation control of the visual channel; 22 represents the response of the auditory channel to changes in information density; 33The activation threshold function of the haptic feedback is expressed by a nonlinear expression to more flexibly cope with the demand changes in different display scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is a flowchart of a multimedia display method in an embodiment of the present application; Figure 2 is a flowchart of step S4 in a multimedia display method in an embodiment of the present application. DETAILED DESCRIPTION

[0024] The present application will be further described in detail below with reference to the accompanying drawings.

[0025] In an embodiment, as shown in Figure 1 The present application discloses a multimedia display method, which specifically comprises the following steps: S1: Obtain initial environmental parameters of the display space and audience group characteristic data, and generate initial display parameters of the multimedia content.

[0026] In this embodiment, the present application is applicable to the interactive display of scientific and technological information in science and technology exhibition halls, enterprise exhibition halls, education and training institutions, and various public display spaces; applicable technical scenarios include the quantum physics interactive exhibition area of a certain science and technology museum, the display space is a 300 square meter circular exhibition hall, equipped with an LED ring-shaped curtain wall, a floor pressure sensing floor, and an infrared positioning base station. The target audience is middle school students and above.

[0027] Specifically, the initial environmental parameters include light intensity (100-500 lux adjustable), sound field frequency distribution (20 Hz-20 kHz frequency response test), and spatial coordinate data (three-dimensional laser scanning modeling), which can be collected by an IoT sensor network. The audience group characteristic data refers to the audience age distribution (72% of 15-45 years old), occupation label (63% of students / technology practitioners), and historical behavior characteristics (such as the average stay time of visitors in similar exhibition items in the past 3 months is 18 minutes) obtained through the gate system or the exhibition hall management system.

[0028] Further, the genetic algorithm is used to optimize the display parameters: such as video resolution: 4K (3840x2160), interactive response delay: ≤80ms, information density: switch one knowledge point every 30 seconds, to obtain the initial display parameters.

[0029] S2: In the process of displaying the multimedia content according to the initial display parameters, monitor the audience behavior data, and calculate the dynamic deviation of the current display effect from the expected target.

[0030] In this embodiment, the audience behavior data is multi-modal behavior data, including visual data, interaction data and physiological data, wherein the visual data is the audience visual trajectory (processed by OpenCV, sampling rate 60fps), the interaction data includes the handle operation frequency of the display device, the touch point coordinate, and the physiological data includes the heart rate variability (HRV) standard deviation ≤5ms as the focused state.

[0031] Specifically, the baseline behavior mode is set based on the expert knowledge base: such as the average gaze duration ≥12 seconds per exhibit, the interaction operation success rate ≥85%, and the positive emotional fluctuation rate ≥70%, to obtain the expected target display effect.

[0032] The multi-modal interaction data stream of the audience in the display space is obtained, the multi-modal interaction data stream is input into a preset audience cognitive load prediction model, and the audience attention decay coefficient and information reception effectiveness index are output; the entropy value change rate H of the display content is dynamically calculated according to the attention decay coefficient, and a dynamic deviation calculation formula is constructed: ΔD = λ × ΔH + (1-λ) × T f (2) Wherein, ΔD is the dynamic deviation, λ is the cognitive decay weight coefficient (0.7≤λ≤0.95), T f is the multi-modal data fusion delay time; when ΔD exceeds the preset threshold, a multi-channel compensation mechanism is triggered, and the heat weight distribution of the content focus area is adjusted through the eye movement trajectory prediction algorithm.

[0033] In this embodiment, the cognitive load prediction model is a time series prediction model based on LSTM, which inputs multi-modal data stream and outputs audience cognitive state index, the input layer includes 3 branches (visual / tactile / physiological) LSTM layer with 128 hidden units, and the time series is processed in dependence; the attention decay coefficient is used to reflect the speed of the audience attention decay with time (value range 0-1); the information reception effectiveness is the proportion of the information actually received and understood by the audience (value range 0-1).

[0034] Further, the cognitive load prediction model adopts the calculation method of loss function and weighted cross entropy, wherein the weight of the attention decay coefficient is 0.6, and the weight of the information reception effectiveness is 0.4.

[0035] Specifically, the attention decay coefficient α s is the ratio of the current attention value to the initial value (α s ∈[0,1]);the entropy value change rate (H) is the instantaneous change rate of content information entropy (bit / s); the multi-modal fusion delay (τ) is the total delay from data acquisition to model output (ms).

[0036] The calculation formula of the attention decay coefficient is: wherein, a s (t) is the attention decay coefficient (dimensionless, value range 0-1) of the current time; Current at is the real-time attention value (dimensionless, value range 0-1) of the current time; Initial at is the initial attention reference value set by the baseline experiment; set by the baseline experiment (e.g., the value is 1 when the audience first gazes at the exhibit); β s = 0.05 is the decay coefficient.

[0037] The content information entropy is calculated using the Shannon entropy formula wherein pi1(t) is the probability of the i1th knowledge point being focused on, and i1 is the knowledge point identifier. The content information entropy value change rate is:

[0038] S3: performing content adaptation processing on the dynamic deviation to obtain a dynamic adjustment parameter of the multi-modal display.

[0039] In this embodiment, a dynamic Bayesian network model of the display system is constructed, and the dynamic deviation DD is input into the network as an observation node. Specifically, the dynamic Bayesian network model is used to model the dynamic dependency relationship of the time series data, the input layer is the dynamic deviation DD, the attention decay coefficient a s , the entropy value change rate The hidden layer includes the cognitive load C, the emotional fluctuation E, and the device load L. The output layer includes the sensory synergy coefficient S, the content granularity G, and the interactive elasticity R.

[0040] The edge weight of the dynamic Bayesian network model is determined by calculating the mutual information (MI) of the historical data: wherein, ω i′j is the weight of the edge connecting the nodes X i′ and X j , indicating the dependency strength (value range: 0 to 1) between the variables X i′ and X j . The greater the weight value, the stronger the correlation between the two, and vice versa. If X i ' is the audience emotional entropy value, X j is the interactive depth index, and ω i′j is 0.8, it indicates that there is a strong correlation between the two; MI(X i′ , X j ) is the mutual information of the variables X i′ and X j , which measures the amount of information shared by the two; max(MI) is the maximum value of the mutual information of all node pairs (X i′ , X j ).

[0041] The posterior probability distribution of the parameter adjustment is calculated by variational inference algorithm, and the core adjustment parameters are extracted. The core adjustment parameters include the sensory channel synergy coefficient α′, the content decoupling granularity β, and the interaction response elasticity γ.

[0042] Generate a dynamic adjustment matrix M = (m) based on the core adjustment parameters. ij )_{3×3}, where the matrix elements are calculated as follows: m 11 =α′×exp(-β×ΔD) m 22 =γ×(1+ΔH / σ) m 33 = 1 / (1+e^{-ΔD}) By performing a nonlinear transformation on the multimodal content stream using matrix M, adaptive remapping of display parameters is achieved.

[0043] Specifically, the dynamic adjustment matrix M is a 3×3 nonlinear transformation matrix used to map multimodal content stream parameters; nonlinear transformation refers to the dynamic coupling and decoupling of parameters through matrix multiplication; adaptive remapping refers to adjusting the priority of content parameters in real time according to scenario requirements.

[0044] S4: Calculate the dynamic adaptation amount of the content in the current display scenario based on the dynamic adjustment parameters, and adjust the presentation form, interaction method and information density of the multimedia content in sync based on the dynamic adaptation amount of the content.

[0045] In this embodiment, the presentation format adjustment includes resolution adaptive adjustment and interaction rate adjustment (e.g., automatically disabling high-precision gesture recognition when network latency > 80ms); the interaction method optimization includes haptic feedback adjustment and multi-channel synchronization, where haptic feedback adjustment refers to dynamically adjusting vibration intensity according to ambient light intensity (the stronger the light, the more obvious the vibration feedback); multi-channel synchronization refers to controlling the latency difference between visual effects (elemental illumination) and sound effects (electronic sounds) within 30ms. Information density control includes knowledge unit segmentation and dynamic push logic, where knowledge unit segmentation is divided into basic, advanced, and expert levels based on the displayed knowledge content; the dynamic push logic, for example, unlocks advanced content based on the audience's dwell time (> 5 minutes).

[0046] In this embodiment, as Figure 2 As shown, step S4 includes: S41: Based on dynamically adjusted parameters, obtain the attention heatmap corresponding to the real-time behavioral characteristics of the audience and the expected content value corresponding to the initial display parameters.

[0047] In this embodiment, assuming a "Brain Science Exploration" exhibit at a science museum, visitors wearing EEG caps and holding interactive pens participate in interactive visualization of neural signals in front of a circular screen. The system needs to adjust the content in real time based on changes in visitor attention. The attention heatmap refers to the density distribution map of visitor gaze points collected by an eye tracker (sampling rate 60Hz, accuracy ±1° angle of view); the expected content value refers to the preset visitor behavior benchmark (e.g., average gaze duration ≥15 seconds) based on initial display parameters (e.g., 4K resolution, 80ms latency).

[0048] Specifically, the power spectrum ratio of theta waves (4-8Hz) to beta waves (13-30Hz) (reflecting focus) is collected using an EEG cap; the attention heatmap uses OpenCV's CamShift algorithm to process eye-tracking data, generating a real-time gaze heatmap, overlaid with an EEG theta / beta ratio layer, and marking cognitive load areas (θ / β>2 is considered high load). Expected value comparison involves calling a baseline heatmap from the initial parameter database (e.g., audience attention concentrated in the core area of ​​the exhibit ≥70%), and calculating the Jaccard similarity index (JSI) between the actual heatmap and the baseline heatmap.

[0049] S42: Based on the difference between the attention heatmap and the expected content value, the dynamic adaptation of multimodal content is calculated comprehensively, including information level compression rate, interaction response delay threshold and sensory channel weight.

[0050] In this embodiment, the information hierarchy compression rate refers to the ratio of simplifying complex content into a multi-layered structure (e.g., jumping from the atomic level, molecular level to the reactive level in the displayed content's molecular structure model). The interaction latency threshold refers to the maximum allowable interaction response time (e.g., haptic feedback latency ≤ 50ms); sensory channel weight refers to the priority allocation coefficient of visual / auditory / tactile senses (∑α). i =1, i includes visual / auditory / tactile senses).

[0051] Specifically, the pixel-level difference between the attention heatmap and the expected value is calculated using the MSE loss function, and the number of audience interaction errors is counted.

[0052] Dynamic adaptation calculation includes: Information hierarchy compression ratio: If JSI < 0.6, then hierarchical dimensionality reduction is triggered (compression ratio = 0.5, such as only retaining the core reaction formula in the display of chemical knowledge); The interaction delay threshold is adjusted based on the theta wave energy of the brain (the delay threshold is relaxed to 80ms when the theta energy is >6μV); Sensory weighting adjustment refers to reducing auditory weighting (α) when the audience frequently turns their heads (angular velocity > 30° / s). 听觉=0.3), by fusing multi-source data through a fuzzy logic controller, the compression rate × delay threshold × weight coefficient is ensured to be ≤1.2.

[0053] S43: Adjust the rendering resolution and interactive response rate of multimedia content according to the dynamic adaptation amount, and control the spatiotemporal consistency parameters of visual, auditory and tactile feedback through a multi-channel synchronization algorithm.

[0054] In this embodiment, the spatiotemporal consistency parameter refers to the time synchronization error (≤50ms) and spatial coordinate deviation (≤5cm) of visual / auditory / tactile feedback; the nonlinear transformation function refers to the mathematical model (such as an S-curve function) that maps the adaptation parameters to device control commands.

[0055] Specifically, rendering parameter adjustments include dynamically adjusting the video bitrate based on the compression ratio (compression ratio 0.5 → bitrate 12Mbps → 24Mbps) and switching the model's Level of Detail (LOD) using OpenGL ES shader programs. Interactive response optimization includes deploying edge computing nodes (Intel NUC) to process haptic feedback data (latency <30ms).

[0056] S44: Optimize the adaptive push logic of multimedia content based on spatiotemporal consistency parameters, and form dynamic calibration with real-time audience feedback.

[0057] In this embodiment, dynamic calibration refers to a closed-loop mechanism that adjusts the content push strategy based on real-time feedback.

[0058] Specifically, the push strategy optimization includes: Update content and push it to the Q-table using the Q-learning algorithm: State space S = {low load, medium load, high load}; Action space A = {accelerated explanation, repeated demonstration, branch switching}, and the reward function R is adjusted according to TD-error: R = α1 × Knowledge Acquisition Rate + β1 × Interaction Satisfaction - γ1 × Cognitive Load. Where α1, β1, and γ1 are weighting coefficients.

[0059] Furthermore, step S44 also includes: S441: Obtain the audience's location trajectory and physiological sensor data in the exhibition space, and calculate the content adaptation degree for the next stage by combining the ambient lighting parameters.

[0060] In this embodiment, the positioning trajectory refers to the spatial movement path of the audience in the exhibition hall (collected by UWB base stations); physiological sensor data refers to bioelectrical signals such as heart rate variability (HRV) and skin conductance response (EDA); and ambient light parameters refer to color temperature (2700K-6500K), illuminance (lux), and glare index (UGR).

[0061] Specifically, by aligning UWB positioning data (x, y, z coordinates) with physiological sensor data (time resolution 1 ms) using timestamps, and based on ambient lighting parameters, the Dynamic Illuminance Index (LVI) is calculated as: LVI = (illuminance / glare index) × color temperature weight, where the weight is 1.2 for color temperature > 4000K and 1.0 for color temperature ≤ 4000K. The fitness calculation model is trained using the random forest algorithm and can output a fitness score, where the fitness score formula is: A = ω1 × trajectory entropy + ω2 × heart rate standard deviation + ω3 × (LVI / interaction count), where ω1, ω2 and ω3 are weighting coefficients; the interaction count refers to the number of times the audience actively interacts within a set time period, and the statistical scope includes effective interaction events such as touching exhibits, voice commands, and gesture recognition.

[0062] S442: Dynamically adjust the narrative rhythm of the displayed content based on content suitability, and optimize the combination strategy of multimodal content through reinforcement learning models.

[0063] In this embodiment, narrative rhythm refers to the rate of change of information density per unit time (such as the number of knowledge point switching times per minute); multimodal combination strategy refers to the content matching strategy of visual, auditory and tactile senses.

[0064] Specifically, the Q-learning algorithm is used to optimize the narrative rhythm through reinforcement learning. The state space is defined as: S = {low fitness, medium fitness (0.3 ≤ A ≤ 0.7), high fitness (A > 0.7)}, and the action space is A = {accelerate the narrative (+20% rhythm), maintain the status quo, slow down the narrative (-15% rhythm)}. Based on the preset fitness score threshold and fitness score, basic rewards and penalties are set, and the ε-greedy strategy is used to explore the optimal action sequence.

[0065] S443: Based on historical data of audience interaction, construct an audience cognitive model and generate personalized content branch paths.

[0066] In this embodiment, the interactive operation history refers to the audience's full-link behavioral data within the exhibits, including tactile data (such as touchscreen click coordinates, swipe trajectories, and pressure intensity), visual data (such as gaze hotspots and gaze duration), and physiological data (such as the power spectra of EEG theta waves (4-8Hz) and beta waves (13-30Hz). The audience cognitive model is a time-series model based on an LSTM neural network to predict the audience's knowledge acquisition level. The emotion entropy value (E) is used to measure the complexity of the audience's emotional fluctuations, with a value range of [0, 1]. E = 0: completely stable emotions (such as full focus); E = 1: drastic emotional fluctuations (such as frequent switching of points of interest).

[0067] Specifically, a Softmax classifier is used to classify cognitive levels and output cognitive labels. For example, personalized branch generation includes: Basic Path: Explanation of Core Knowledge Points (5 minutes) Advanced Branch: Extended Experiment Demonstration (3 minutes) Expert Branch: Mathematical Modeling Derivation (8 minutes).

[0068] Specifically, the personalized content branch path is calculated using formula (1), which is shown below: B=α×E+(1-α)×I (1) Where B represents the content branch weight, E represents the audience emotion entropy value, I represents the interaction depth index, and α is the dynamic adjustment coefficient (0.6≤α≤0.9).

[0069] For example, determine the branch path based on the value of B and the threshold: If B is greater than 0.7, it means that the audience has a high level of understanding and interest in the content being presented, so guide them to the path of "advanced quantum knowledge"; If B is between 0.4 and 0.7, it indicates that the audience is at an intermediate level, and the course proceeds to the "intermediate experimental explanation" path. If B is less than or equal to 0.4, it means that the audience may not be very familiar with the content, so they should start learning from the "Introduction to Basic Concepts".

[0070] S444: Based on the device network status of the exhibition space, obtain the real-time rendering load parameters of each terminal node, and calculate the synchronization compensation amount of the content dynamic adaptation in the distributed system.

[0071] In this embodiment, the real-time rendering load parameters refer to the GPU utilization, CPU load, memory utilization, and network bandwidth utilization of the terminal device; the synchronization compensation amount refers to the compensation value (unit: milliseconds) for the difference in content rendering latency caused by differences in device performance.

[0072] Specifically, the adaptation compensation coefficient C for each terminal is calculated. i = Maximum rendering time of the terminal / Global average rendering time, where i is the terminal identifier. The formula for generating the synchronization compensation amount is: Latency i For network latency of terminal i, for example, when ΔT = 98ms, a 98ms buffer needs to be added during content distribution.

[0073] S445: Adjust the priority strategy for multi-terminal content distribution based on the amount of synchronization compensation.

[0074] In this embodiment, the content distribution priority strategy determines the decision rules for the order of content push to multiple terminals; edge computing nodes refer to data centers deployed locally, which are responsible for real-time data processing and caching; network bandwidth allocation refers to the dynamic allocation of uplink / downlink bandwidth resources according to terminal needs.

[0075] Specifically, the strategy for adjusting the priority of multi-terminal content distribution based on the amount of synchronization compensation also includes: S4451: Establish a digital twin model of the displayed content to map the differences in content status between physical and virtual spaces in real time.

[0076] In this embodiment, the digital twin model is a virtual mirror of the physical space, including real-time mapping of spatial layout, device status, and content version; the state difference mapping refers to the difference parameters between the physical space and the virtual space in terms of content version, rendering parameters, and interaction status.

[0077] Specifically, a 1:1 digital twin model of the exhibition hall is constructed (containing 1000+ interactive objects); the model includes the following dynamic attributes: exhibit location coordinates (x, y, z), terminal device status (running / faulty), and content version number; device status topics are subscribed to via MQTT protocol, and OpenCV is used for image recognition to compare the positional deviations between the physical exhibits and the virtual model; when an inconsistency in content version is detected, a version synchronization command is triggered on the edge nodes.

[0078] S4452: Dynamically divide content push areas and optimize network bandwidth allocation based on the spatial distribution characteristics of audience groups.

[0079] In this embodiment, the spatial distribution characteristics of the audience group are spatial clustering indicators calculated based on UWB positioning data; the content push area refers to the priority area dynamically adjusted according to the audience density (such as prioritizing core content in high-density areas); and the network bandwidth allocation is the uplink and downlink bandwidth quota dynamically adjusted through SDN (Software Defined Networking).

[0080] Specifically, the DBSCAN algorithm is used to divide the audience into groups: for example, by setting two parameters: `eps=3`: If a viewer has at least 3 other viewer members (including themselves) within a 3-meter radius, they are grouped into the same group; `min_samples=3`: Each group must have at least 3 people; groups with fewer than 3 people are not counted. This yields the clustering results for the viewer groups.

[0081] Based on the clustering results, three content push areas were divided, for example: Core area: audience density > 5 people / m² 2 (Prioritize 4K content); Transition zone: Audience density 1-5 people / m² 2 (Pushing 1080P content); Edge area: Audience density <1 person / m²2 (Push text summary)

[0082] Bandwidth optimization strategy refers to dynamically adjusting the QoS queue through the OpenFlow protocol.

[0083] S4453: Aggregates audience behavior data from multiple terminal nodes through a federated learning mechanism to continuously optimize the group adaptability of the content recommendation algorithm.

[0084] In this embodiment, federated learning is a distributed machine learning framework in which each terminal node only shares encrypted gradient parameters; group adaptability refers to the recommendation algorithm's ability to generalize to the behavioral patterns of the audience group.

[0085] Specifically, each terminal runs a lightweight LSTM model, which mainly consists of two parts: the LSTM layer (LSTM): it is a "long short-term memory" network layer that is good at processing ordered data, such as time series, speech, and action trajectories, and can extract 64-dimensional temporal features; for example, it can be used to understand the movement behavior or interaction habits of visitors in the exhibition hall over a period of time.

[0086] Fully connected layer (dense): The output layer, which has only one neuron and uses the sigmoid activation function; the output of a fully connected layer is a value between 0 and 1, which is often used to determine the probability of something happening; for example, predicting whether the audience will be interested in a certain display content, or whether they will continue to watch.

[0087] Population adaptability validation: The population accuracy of the classification model on the validation set of the population dataset. group Perform calculations when Accuracy group Model deployment is triggered when the percentage is >85%.

[0088] in Among them, TP group A true positive for the population represents the number of samples correctly predicted as positive by the model in the population dataset, indicating the actual number of viewers correctly identified by the model as having high engagement; TN group For a true negative result, this indicates the actual number of viewers correctly identified by the model as having low engagement; FP group A false positive in the population indicates the number of samples that were actually low-engagement viewers but were misclassified as high-engagement by the model; FN group A false negative indicates the number of samples that were actually highly engaged viewers but were misclassified as low-engaged by the model.

[0089] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0090] In one embodiment, a multimedia display system is provided, which corresponds to a multimedia display method in the above embodiments.

[0091] A multimedia display system includes a data acquisition module, a content generation module, a display control module, a behavior detection module, and a dynamic adjustment module. Detailed descriptions of each functional module are as follows: The data acquisition module is used to obtain the initial environmental parameters of the exhibition space and the characteristics of the audience. The content generation module is used to generate initial display parameters for multimedia content based on initial environmental parameters and audience characteristic data. The display control module is used to control the display device to output corresponding content during the process of displaying multimedia content according to the initial display parameters; The behavior detection module is used to monitor audience behavior data in real time. The dynamic adjustment module is used to calculate the dynamic deviation between the current display effect and the expected goal based on audience behavior data, and to perform content adaptation processing on the dynamic deviation to obtain the dynamic adjustment parameters for multimodal display. The display control module is also used to calculate the dynamic adaptation amount of content in the current display scenario based on the dynamically adjusted parameters, and to simultaneously adjust the presentation form, interaction method and information density of multimedia content.

[0092] For specific limitations regarding a multimedia display system, please refer to the limitations of a multimedia display method mentioned above, which will not be repeated here. Each module in the multimedia display system described above can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0093] In one embodiment, a display device is provided, which may be a server. The display device includes a processor, memory, a network interface, and a database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores initial environment parameters, audience group characteristic data, and audience behavior data, etc. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a multimedia display method.

[0094] In one embodiment, a display device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: S1: Obtain the initial environmental parameters of the exhibition space and the characteristic data of the audience group to generate the initial display parameters of the multimedia content; S2: During the process of displaying the multimedia content according to the initial display parameters, monitor the audience behavior data and calculate the dynamic deviation between the current display effect and the expected goal. S3: Perform content adaptation processing on dynamic deviations to obtain dynamic adjustment parameters for multimodal display; S4: Calculate the dynamic adaptation amount of the content in the current display scenario based on the dynamic adjustment parameters, and adjust the presentation form, interaction method and information density of the multimedia content in sync based on the dynamic adaptation amount of the content.

[0095] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: S1: Obtain the initial environmental parameters of the exhibition space and the characteristic data of the audience group to generate the initial display parameters of the multimedia content; S2: During the process of displaying the multimedia content according to the initial display parameters, monitor the audience behavior data and calculate the dynamic deviation between the current display effect and the expected goal. S3: Perform content adaptation processing on dynamic deviations to obtain dynamic adjustment parameters for multimodal display; S4: Calculate the dynamic adaptation amount of the content in the current display scenario based on the dynamic adjustment parameters, and adjust the presentation form, interaction method and information density of the multimedia content in sync based on the dynamic adaptation amount of the content.

[0096] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0097] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0098] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method of multimedia presentation, characterized by, The application comprises the following steps: Obtain initial environmental parameters of the display space and audience group characteristics data, and generate initial display parameters of the multimedia content; Monitor audience behavior data during the display of the multimedia content according to the initial display parameters, and calculate the dynamic deviation of the current display effect from the expected target; Perform content adaptation processing on the dynamic deviation to obtain dynamic adjustment parameters of multi-modal display; According to the dynamic adjustment parameters, calculate the content dynamic adaptation amount in the current display scenario, and synchronously adjust the presentation form, interactive mode and information density of the multimedia content according to the content dynamic adaptation amount.

2. The method of claim 1, wherein, According to the dynamic adjustment parameters, calculate the content dynamic adaptation amount in the current display scenario, and synchronously adjust the presentation form, interactive mode and information density of the multimedia content according to the content dynamic adaptation amount, specifically comprising: According to the dynamic adjustment parameters, obtain the attention heat map corresponding to the real-time behavior characteristics of the audience and the content expected value corresponding to the initial display parameters; Based on the difference between the attention heat map and the content expected value, the dynamic adaptation amount of the multi-modal content is comprehensively calculated, including information level compression rate, interactive response delay threshold and sensory channel weight; According to the dynamic adaptation amount, adjust the rendering resolution and interactive response rate of the multimedia content, and control the space-time consistency parameters of visual, auditory and tactile feedback through a multi-channel synchronous algorithm; According to the space-time consistency parameters, optimize the adaptive pushing logic of the multimedia content, and form dynamic calibration with real-time feedback of the audience.

3. A method of presenting multimedia according to claim 2, wherein, According to the space-time consistency parameters, optimize the adaptive pushing logic of the multimedia content, and form dynamic calibration with real-time feedback of the audience, further comprising: obtaining the positioning trajectory and physiological sensing data of the audience in the display space, and calculating the content adaptation degree of the next stage in combination with the environmental lighting parameters; According to the content adaptation degree, dynamically adjust the narrative rhythm of the display content, and optimize the combination strategy of multi-modal content through a reinforcement learning model; Based on the historical data of audience interactive operation, build an audience cognitive model and generate a personalized content branch path.

4. The method of claim 3, wherein, The personalized content branch path is calculated by formula (1), and formula (1) is as follows: B=α×E+(1-α)×I (1) Wherein, B represents the content branch weight, E represents the audience emotion entropy value, I represents the interactive depth index, and a is a dynamic adjustment coefficient (0.6≤a≤0.9); According to the device network state of the display space, obtain the real-time rendering load parameters of each terminal node, and calculate the synchronization compensation amount of the content dynamic adaptation amount in the distributed system; According to the synchronization compensation amount, adjust the priority strategy of multi-terminal content distribution.

5. The method of claim 4, wherein, According to the synchronization compensation amount, adjust the priority strategy of multi-terminal content distribution, further comprising: Establish a digital twin model of the display content to real-time map the content state difference between the physical space and the virtual space; According to the spatial distribution characteristics of the audience group, dynamically divide the content pushing area and optimize the network bandwidth allocation; Aggregate the audience behavior data of multiple terminal nodes through a federated learning mechanism to continuously optimize the group adaptability of the content recommendation algorithm.

6. A method of presenting multimedia according to claim 1 or 5, characterized in that, The monitoring audience behavior data, calculating the dynamic deviation of the current display effect and the expected target, specifically comprising: Obtaining the multi-modal interaction data stream of the audience in the display space, inputting the multi-modal interaction data stream into the preset audience cognitive load prediction model, outputting the audience attention attenuation coefficient and the information reception effective degree index; calculating the entropy value change rate ΔH of the display content according to the attention attenuation coefficient, and constructing a dynamic deviation calculation formula: AD = λ x ΔH + (1 - λ) x T f (2) Wherein, ΔD is dynamic deviation, λ is cognitive attenuation weight coefficient (0.7≤λ≤0.95), T f is the multimodal data fusion delay time; When ΔD exceeds the preset threshold, triggering a multi-channel compensation mechanism, and adjusting the heat weight distribution of the content focus area through an eye movement trajectory prediction algorithm.

7. A method of presenting multimedia according to claim 6, wherein, The content adaptation processing of the dynamic deviation is performed to obtain the dynamic adjustment parameters of multi-modal display, comprising: Building a dynamic Bayesian network model of the display system, inputting the dynamic deviation ΔD as an observation node into the network; The posterior probability distribution of the parameter adjustment is calculated using a variational inference algorithm, and the core adjustment parameters are extracted. These core adjustment parameters include the sensory channel coordination coefficient α. ′ Content decoupling granularity β and interactive response elasticity γ; A dynamic adjustment matrix M = (m ij )_{3×3} is generated according to the core adjustment parameters, wherein: m 11 = a ′ × exp(-β×ΔD) m 22 = γ x (1 + ΔH / σ) m 33 = 1 / (1 + e^(-ΔD)) Through the matrix M, the multi-modal content stream is nonlinearly transformed to realize the adaptive remapping of the display parameters.

8. A multimedia presentation system characterized by A system for performing a multimedia display method as claimed in any one of claims 1 to 7, the system comprising: A data acquisition module for acquiring initial environmental parameters of a display space and audience group characteristic data; A content generation module for generating initial display parameters of multimedia content according to the initial environmental parameters and audience group characteristic data; A display control module for controlling the display device to output corresponding content during the display of the multimedia content according to the initial display parameters; A behavior detection module for monitoring audience behavior data in real time; A dynamic adjustment module for calculating the dynamic deviation between the current display effect and the expected target according to the audience behavior data, and performing content adaptation processing on the dynamic deviation to obtain dynamic adjustment parameters of multi-modal display; The display control module is also used for calculating the content dynamic adaptation amount in the current display scene according to the dynamic adjustment parameters, and synchronously adjusting the presentation form, interactive mode and information density of the multimedia content.

9. A display device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the multimedia display method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to realize the steps of the multimedia display method as claimed in any one of claims 1 to 7.