An immersive digital-generated media system and method based on real-time interactive feedback
By integrating closed-loop collaborative technologies of multimodal interactive acquisition, real-time feedback processing, and AI adaptive generation, the problems of delayed interactive response, insufficient multimodal interaction integration, and monotonous immersion in digital media content generation devices have been solved, achieving instant feedback and a comprehensive immersive experience, thereby enhancing user immersion and adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN TEXTILE UNIV
- Filing Date
- 2026-01-20
- Publication Date
- 2026-06-23
AI Technical Summary
Existing digital media content generation devices suffer from problems such as delayed interactive response, insufficient integration of multimodal interaction, limited immersion, and low flexibility in content generation, making it impossible to achieve real-time feedback and a fully immersive experience.
It employs a multimodal interactive acquisition module, a real-time feedback processing module, an AI content generation engine, an environment adaptation module, an immersive output module, a storage and transmission module, and a human-computer interaction control module. By integrating multimodal interactive acquisition, real-time feedback processing, AI adaptive generation, and multi-channel immersive output, it achieves a closed-loop collaboration of "interaction-feedback-generation-presentation," thereby enhancing interactivity and immersion.
It achieves real-time interactive feedback with an interaction response latency of ≤150ms and a feedback response time of ≤20ms. It also features multimodal collaboration and integration, with content generation adaptability to user needs and scenarios exceeding 90%, thus expanding the scope of application scenarios.
Smart Images

Figure CN122261371A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital media technology, and specifically to an immersive digital media generation system and method based on real-time interactive feedback. Background Technology
[0002] With the rapid development of digital media technology, the media industry has placed higher demands on the interactivity and immersiveness of content. Existing digital media content generation equipment and technologies face the following key problems:
[0003] 1. Delayed interactive response: Traditional content generation devices mostly adopt a serial working mode of "acquisition-processing-generation", which results in high latency in interactive data processing (usually >200ms), making it impossible to achieve real-time feedback and ruining the user's immersive experience;
[0004] 2. Insufficient integration of multimodal interaction: Existing devices mostly collect action or voice data in a single way, lacking comprehensive consideration of physiological perception and environmental parameters, resulting in poor adaptability of content generation to user needs and scenarios;
[0005] 3. Limited immersive experience: Most systems focus only on visual or auditory output, lacking multi-channel synergy such as tactile and spatial sound effects, making it difficult to create a fully immersive experience;
[0006] 4. Low content generation flexibility: Traditional content generation relies on preset templates or manual intervention, and cannot dynamically adjust the content style and plot direction according to real-time user interaction, thus limiting its applicability to various scenarios.
[0007] Therefore, there is an urgent need to design a digital media content generation system that combines real-time interactive feedback, multimodal data fusion, adaptive content generation, and multi-channel immersive presentation to address the shortcomings of existing technologies and meet the needs of high-quality development in the media industry. Summary of the Invention
[0008] Purpose of the invention:
[0009] The purpose of this invention is to provide an immersive digital media content generation system and method based on real-time interactive feedback. By integrating multimodal interactive acquisition, real-time feedback processing, AI adaptive generation, and multi-channel immersive output modules, it achieves a closed-loop collaboration of "interaction-feedback-generation-presentation", improves the interactivity, immersion, and adaptability of digital media content, and meets the diversified needs of professional media creation, digital entertainment, virtual live streaming, and other scenarios.
[0010] An immersive digital media content generation system based on real-time interactive feedback includes a multimodal interactive acquisition module, a real-time feedback processing module, an AI content generation engine, an environment adaptation module, an immersive output module, a storage and transmission module, a power supply module, and a human-computer interaction control module.
[0011] The multimodal interaction acquisition module is used to simultaneously acquire the user's action data, voice data, and physiological perception multimodal data, and convert the multimodal data into standardized digital signals and send them to the real-time feedback processing module.
[0012] The real-time feedback processing module is electrically connected to the multimodal interaction acquisition module and includes a data preprocessing unit, an interaction intent recognition unit, and a feedback signal generation unit. The real-time feedback processing module is used to perform real-time noise reduction, feature extraction, and intent parsing on the acquired multimodal data to generate multimodal instantaneous feedback signals.
[0013] The AI content generation engine is electrically connected to the real-time feedback processing module and the environment adaptation module, and includes a multimodal data fusion unit, an adaptive content generation unit, and a content optimization unit. The AI content generation engine fuses multimodal real-time feedback signals and dynamically generates and optimizes digital media content based on the interaction intent and the environmental parameters collected by the environment adaptation module. The optimized digital media content is then output through the immersive output module.
[0014] The environment adaptation module includes an ambient light sensor, a spatial positioning sensor, and an acoustic sensor. The environment adaptation module is used to collect ambient light intensity, system spatial position, and environmental noise data and send them to the AI content generation engine to provide an environment adaptation basis for content generation.
[0015] The immersive output module is electrically connected to the AI content generation engine and includes a visual output unit, an auditory output unit, and a tactile feedback unit, used to present the generated digital media content synchronously through multiple channels.
[0016] The storage and transmission module is electrically connected to the AI content generation engine and is used to store the original interactive data, generated content data and configuration parameters, and to achieve high-speed data transmission through the 5G / Wi-Fi 6 / USB4 interface;
[0017] The power supply module is used to provide a stable voltage and supports two modes: continuous power supply and external power supply.
[0018] The human-computer interaction control module is used to electrically connect with the AI content generation engine and the immersive output module, and to perform parameter configuration, content preview and interaction mode switching.
[0019] Furthermore, the multimodal interaction acquisition module includes an action acquisition unit and a voice acquisition unit;
[0020] The motion acquisition unit uses a 1080P RGB-D depth camera and a six-axis MEMS motion sensor, which can simultaneously capture the user's limb movements, posture changes and motion trajectories;
[0021] The voice acquisition unit uses an array microphone, which supports long-distance voice acquisition and noise suppression;
[0022] The physiological sensing unit employs a photoelectric heart rate sensor and a skin conductance sensor to capture changes in the user's physiological state.
[0023] Furthermore, the real-time feedback processing module includes a noise reduction unit, a feature extraction unit, an interaction intent recognition unit, and a feedback signal generation unit. The noise reduction unit is used to reduce noise in multimodal data. The feature extraction unit is used to extract features from the denoised multimodal data to obtain speech features, action features, and physiological features, which are then input into the interaction intent recognition unit. The interaction intent recognition unit adopts a lightweight semantic understanding model based on Transformer to determine the activity level of each modality of data based on the input speech features, action features, and physiological features, and to determine the interaction intent category and intent intensity.
[0024] The feedback signal generation unit uses a pulse width modulation control algorithm to generate multimodal real-time feedback signals, including tactile feedback signals and visual cue signals, based on the type and intensity of the interaction intent.
[0025] Furthermore, the workflow of the multimodal data fusion unit of the AI content generation engine includes the following steps:
[0026] Step 1: Scene type identification:
[0027] Based on the "interaction intent category" and "activity level of each modality" output by the real-time feedback processing module and the "scene parameters" collected by the environment adaptation module, the scene type is determined according to the trained classification and recognition model. The classification and recognition model adopts a lightweight classifier and is obtained through the multimodal interaction-scene mapping sample training set. The samples in the multimodal interaction-scene mapping sample training set are action, voice, physiological and environmental parameter data of different users with actual scene labels.
[0028] Step 2: Basic weight allocation based on scene type:
[0029] Based on the scene type, initial basic weights corresponding to action, voice, physiological and environmental parameters are assigned according to the preset scheme corresponding to the scene type. An attention mechanism fusion algorithm is used to perform feature weighted fusion of action, voice, physiological and environmental data. The weights are dynamically adjusted according to the interaction scene.
[0030] The adaptive content generation unit adopts a model architecture that combines generative adversarial networks and reinforcement learning. It is obtained through a multimodal interaction-content mapping sample training set, in which the samples are action, speech, physiological and environmental parameter data of different users with actual content labels.
[0031] The content optimization unit employs super-resolution reconstruction, frame rate interpolation, and noise suppression algorithms to improve the quality of content output.
[0032] Furthermore, the method for dynamically adjusting weights based on the interaction scenario in step 2 includes the following steps:
[0033] Step 3.1: Calculate the "confidence score" for each modality of data:
[0034] Calculate the confidence score of motion data: Based on the motion data collected by the motion acquisition unit, if the residual of the motion data after Kalman filtering is <0.5mm, the confidence score is in the high range of 0.8-1.0, and the weight is increased by 10%-20% from the base value; if the residual of the motion data after Kalman filtering is >0.3mm, the confidence score is in the low range of 0.3-0.5, and the weight is decreased by 10%-15%; if the residual of the motion data after Kalman filtering is between 0.3-0.5mm, the base weight remains unchanged.
[0035] Calculate the confidence score of the speech data: Based on the signal-to-noise ratio of the speech data collected by the speech acquisition unit, the speech clarity score is calculated after spectral subtraction. If the ambient noise is ≤40dB and the speech command is complete, the confidence score is in the high range of 0.8-1.0, and the weight increases by 10%-20%. If the ambient noise is >60dB, the speech data is considered distorted, and the confidence score is in the low range of 0.2-0.4, and the weight decreases by 20%-30%. If the ambient noise is between 0.3-0.5mm, the base weight remains unchanged.
[0036] Calculating the confidence score of physiological data: Based on the measurement results of the physiological sensing unit, if the heart rate variability is <5% and the skin conductance fluctuation is <0.1μS, the confidence score is in the high range of 0.7-0.9, with the weight increased by 5%-10%; if the heart rate variability is >10% and the skin conductance fluctuation is >0.5μS, the confidence score is in the low range of 0.3-0.5, with the weight decreased by 5%.
[0037] Calculate the confidence score of environmental data: Based on the measurement accuracy of environmental sensors, if the rate of change of environmental parameters is lower than the preset lower threshold, the confidence score is in the high range of 0.7-0.9, and the weight increases by 5%-10%; if the rate of change of environmental parameters is higher than the preset upper threshold, the confidence score is in the low range of 0.2-0.4, and the weight decreases by 5%-10%.
[0038] Attention weight normalization: The "base weight × confidence score" of each modality is used as a temporary weight, and then normalized by the Softmax function to obtain the normalized temporary weights of each modality;
[0039] Step 3.2, Feature Weighted Fusion: Multiply the extracted features of each modality by the corresponding normalized attention weights, and then concatenate the vectors to generate the final 1024-dimensional scene feature vector;
[0040] Step 3.3: Real-time iterative optimization of weights. The current weights are updated after the iteration conditions are met. The iteration triggering conditions are:
[0041] Time-triggered: The weights are iterated once at fixed intervals;
[0042] Event triggering: When the interaction intent changes, environmental parameters exceed the preset environmental change threshold, or physiological state exceeds the preset physiological state change threshold, the weight update is triggered immediately;
[0043] The iterative optimization aims to improve user feedback on content generated by the AI content generation engine, and dynamically adjusts the weighting strategy through a reinforcement learning reward function.
[0044] If the user's subsequent interactions and the generated content have a high degree of matching, the current weight strategy will score high, and the current weight allocation logic will be maintained in subsequent iterations.
[0045] If a user makes a secondary correction interaction, the relevant weights will be adjusted based on the correction behavior.
[0046] Furthermore, the environment adaptation module is used to collect scene parameters using a spatial positioning sensor, an acoustic sensor, and an ambient light sensor;
[0047] The spatial positioning sensor uses a UWB positioning module, combined with an inertial navigation algorithm, to achieve real-time positioning of the system in three-dimensional space; the acoustic sensor uses a noise sensor for environmental noise detection and adaptive gain adjustment of audio content; the ambient light sensor is used to adjust the brightness and contrast of the visual output unit.
[0048] Furthermore, the immersive output module includes a visual output unit, an auditory output unit, and a tactile feedback unit.
[0049] A method for generating digital media content includes the following steps:
[0050] Step S1, System Initialization and Environment Adaptation: After the system starts, the environment adaptation module collects initial ambient light intensity, spatial location, and ambient noise data; the multimodal interaction acquisition module performs self-check and calibration; and the AI content generation engine loads the pre-trained model and configuration parameters to complete system initialization.
[0051] Step S2, Multimodal Interaction Data Acquisition: The multimodal interaction acquisition module synchronously acquires the user's motion trajectory, voice commands, and physiological state data, and transmits them to the real-time feedback processing module through the SPI interface;
[0052] Step S3: Real-time feedback and intent parsing: The real-time feedback processing module performs noise reduction and feature extraction on the collected data, and parses the user's core needs through the interactive intent recognition unit, generates instant feedback signals and transmits them to the immersive output module to realize interactive response;
[0053] Step S4: AI Adaptive Content Generation: The AI content generation engine receives interactive intent data and environmental parameters, performs feature fusion through the multimodal data fusion unit, generates digital media content that matches the scene through the adaptive content generation unit, and improves the quality through the content optimization unit.
[0054] Step S5: Immersive Content Presentation: The visual, auditory, and tactile units of the immersive output module work synchronously to present optimized digital media content through multiple channels, while capturing subsequent user interaction data in real time to form a closed-loop interaction.
[0055] Step S6: Data storage and transmission: The storage and transmission module stores the original interactive data and generated content data in real time, and transmits them to external devices via wired or wireless means according to user settings.
[0056] The present invention, by adopting the above technical solution, has the following beneficial effects:
[0057] 1. Real-time interactive feedback: Employing multimodal data parallel processing and lightweight algorithms, the interactive response latency is ≤150ms and the feedback response time is ≤20ms, solving the problem of lag in traditional system interaction and improving the user's immersive experience;
[0058] 2. Multimodal collaborative fusion: Simultaneously collect action, voice, physiological and environmental data, and dynamically adjust the fusion weight through an attention mechanism. The content generation is over 90% adaptable to user needs and scenarios, expanding the scope of application scenarios. Attached Figure Description
[0059] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0060] The present invention will be further described below. Unless otherwise specified, the materials, reagents, and instruments used in the following embodiments of the present invention are all conventional materials, reagents, and instruments in the art, and are commercially available. Unless otherwise specified, the technical means used in the present invention are methods known to those skilled in the art.
[0061] 1. Technical Solution
[0062] The immersive digital media content generation system based on real-time interactive feedback of the present invention specifically includes a multimodal interactive acquisition module, a real-time feedback processing module, an AI content generation engine, an environment adaptation module, an immersive output module, a storage and transmission module, a power supply module, and a human-computer interaction control module. The connection relationship and function of each module are as follows:
[0063] 1.1 Overall Structure and Connection Relationships
[0064] Mechanical connection: The camera and sensor components of the multimodal interactive acquisition module are fixed to the front of the system via an adjustable bracket, and the display unit and feedback unit of the immersive output module are detachably connected via a snap-fit structure to ensure installation stability and portability;
[0065] Electrical connections: The multimodal interactive acquisition module and the environment adaptation module are connected to the real-time feedback processing module via a PCIe 4.0 interface; the real-time feedback processing module and the environment adaptation module are connected to the AI content generation engine via a high-speed serial bus; the immersive output module, the storage and transmission module, and the human-computer interaction control module are connected to the AI content generation engine via a USB4 interface; the power supply module provides the corresponding voltage to all modules through a power distribution board and supports hot-swapping to switch power supply modes.
[0066] 1.2 Detailed Structure of Each Module
[0067] 1.2.1 Multimodal Interactive Acquisition Module
[0068] The multimodal interaction acquisition module is the system's input unit, used to comprehensively capture user interaction information and status data, including:
[0069] Motion capture unit: The RGB-D depth camera uses a Sony IMX586 sensor with a resolution of 1920×1080 and a frame rate of 60fps. It achieves depth measurement through infrared structured light technology, with a depth range of 0.5-10m, and can accurately extract the coordinates of 25 human joint points (error ≤3mm); The six-axis MEMS motion sensor model MPU9250 integrates a 3-axis gyroscope (measurement range ±2000° / s, zero bias stability ≤2° / h) and a 3-axis accelerometer (measurement range ±16g, nonlinear error ≤0.5%), and simultaneously captures the rapid movement trajectory of the user's hands and limbs;
[0070] Voice acquisition unit: The 4-microphone array uses Knowles SPH0641LM4H microphones with a sampling rate of 48kHz and a bit depth of 24bit. It achieves 360° omnidirectional voice acquisition through beamforming algorithm, effectively suppressing environmental noise (improving the signal-to-noise ratio by 30dB) and supporting long-distance voice command recognition within 5m.
[0071] Physiological sensing unit: The photoelectric heart rate sensor uses the MAX30102 chip, which captures changes in blood flow through a green LED and a photodetector. The measurement range is 50-200 bpm, and the response time is ≤5s. The skin conductance sensor uses the AD8421 instrumentation amplifier, with a measurement range of 0-5μS and a resolution of 0.01μS. It is used to capture changes in skin conductivity caused by the user's emotional fluctuations.
[0072] 1.2.2 Real-time Feedback Processing Module
[0073] The real-time feedback processing module is the core of the interactive response, enabling rapid processing and immediate feedback of multimodal data, including:
[0074] Data preprocessing unit: Utilizes an FPGA chip (Xilinx Artix-7) to perform parallel noise reduction processing on the collected motion, speech, and physiological data (motion data uses Kalman filtering, speech data uses spectral subtraction, and physiological data uses moving average filtering), with a processing latency of ≤10ms;
[0075] Interactive Intent Recognition Unit: Based on a lightweight Transformer model (6 Encoder layers, 8 attention heads), with 512-dimensional input features, trained with 500,000 sets of labeled samples, it supports the recognition of 100+ common interactive intents (such as "switching scenes", "adjusting perspective", "pausing content", etc.), with an intent recognition accuracy of ≥95% and a processing latency of ≤30ms;
[0076] Feedback signal generation unit: Using an STM32H750 microcontroller, it generates tactile vibration signals (frequency and intensity can be dynamically adjusted) and visual cue signals (such as screen highlighting and icon flashing) through a PWM algorithm. The feedback response time is ≤20ms, ensuring that users receive instant interactive feedback.
[0077] 1.2.3 AI Content Generation Engine
[0078] The AI content generation engine is the core computing unit of the system, enabling multimodal data fusion and adaptive content generation, including:
[0079] Multimodal data fusion unit: An attention-based fusion algorithm is used to construct an interactive scene feature vector (1024 dimensions). The weights of each modality are dynamically adjusted according to the scene type: In action-dominant scenes (such as game interaction), the weights of action, speech, physiology, and environment are 0.6, 0.2, 0.1, and 0.1 respectively; in speech-dominant scenes (such as virtual assistant interaction), the weights of speech, action, physiology, and environment are 0.4, 0.3, 0.15, and 0.15 respectively. The signal-to-noise ratio of the fused feature vector is ≥50dB.
[0080] Adaptive Content Generation Unit: Employs a model architecture combining GAN and reinforcement learning. The generator is an improved StyleGAN2 (generating 4K resolution), and the discriminator uses a PatchGAN structure. The reward function of reinforcement learning is dynamically adjusted based on user interaction feedback (action responsiveness, physiological comfort). It supports the generation of multiple content types, including text, images, videos, and 3D models, with a generation latency of ≤100ms and a content similarity matching degree of ≥90% with user intent.
[0081] Content optimization unit: Super-resolution reconstruction uses the ESRGAN model to increase the resolution of generated content from 1080P to 4K, with a detail retention rate of ≥95%; frame rate interpolation uses the RIFE algorithm to increase the content frame rate from 60fps to 120fps, reducing motion blur by 40%; noise suppression uses the BM3D algorithm to effectively remove Gaussian noise and artifacts from the generated content.
[0082] 1.2.4 Environment Adaptation Module
[0083] The environment adaptation module collects environmental parameters of the system to provide a basis for scene adaptation in content generation, including:
[0084] Ambient light sensor: Model BH1750, measurement range 10-10000 lux, accuracy ±5%, response time 120ms, outputs light intensity data via I2C interface, used to adjust the brightness (adjustment range 100-1000cd / m²) and contrast of the visual output unit;
[0085] Spatial positioning sensor: The UWB positioning module uses the Decawave DW1000 chip, with a positioning accuracy of ±1cm, a positioning refresh rate of 100Hz, and a communication distance of ≤50m. Combined with the inertial navigation algorithm (EKF extended Kalman filter), it realizes the real-time positioning of the system in three-dimensional space, with a positioning drift of ≤0.5cm / s.
[0086] Acoustic sensor: Noise sensor model MAX9814, measurement range 0-120dB, accuracy ±0.5dB, frequency response 20Hz-20kHz, used to detect ambient noise intensity and adaptively adjust the gain (adjustment range 0-30dB) and noise reduction level of audio output.
[0087] 1.2.5 Immersive Output Module
[0088] The immersive output module is used to simultaneously present digital media content across multiple channels, creating a fully immersive experience, including:
[0089] Visual output unit: The VR headset display uses a Samsung AMOLED panel with a resolution of 2160×2160 per eye, a refresh rate of 120Hz, a pixel density of 1000PPI, a response time of 1ms, a field of view of 120°, and supports HDR10+ high dynamic range display (dynamic range 1000000:1); the 4K projection module uses DLP technology, with a brightness of 5000 lumens, a contrast ratio of 10000:1, a throw ratio of 1.2:1, and supports keystone correction and autofocus;
[0090] Auditory output unit: The spatial audio processing module uses a Cirrus Logic CS49846 chip, supports 7.1.2 channel surround sound decoding, and achieves three-dimensional spatial sound positioning through head-related transfer function (HRTF) with a positioning accuracy of ±5°, a frequency response of 20Hz-20kHz, and a total harmonic distortion of ≤0.1%; it is equipped with a 3.5mm headphone jack and a Bluetooth 5.2 module, supporting wireless / wired audio output switching;
[0091] Haptic feedback unit: Electromagnetic vibration motor model FFM-1030, vibration frequency 20-500Hz, vibration intensity 0-5N, response time ≤10ms, distributed in the palm and fingertip positions of the interactive handle (4 motors in total); The pneumatic haptic feedback system uses a miniature air pump and airbag, supports 0-20kPa air pressure adjustment, and simulates different tactile sensations (such as pressing, collision).
[0092] 1.2.6 Storage and Transmission Module
[0093] Local storage unit: NVMe M.2 solid-state drive with PCIe 4.0×4 interface, 1TB capacity, read speed ≥3500MB / s, write speed ≥3000MB / s, supports TRIM command and hardware encryption; microSD card slot supports UHS-II interface, supports up to 2TB SDXC card, read and write speed ≥300MB / s;
[0094] Wireless transmission unit: The 5G modem supports SA / NSA dual-mode, is compatible with the Sub-6GHz band, and has a peak download speed of ≥2Gbps and a peak upload speed of ≥500Mbps; the Wi-Fi 6 module supports the 802.11ax protocol, 2.4GHz / 5GHz dual-band, has a peak speed of ≥9.6Gbps, supports MU-MIMO technology, and improves connection stability by 40%.
[0095] Wired transmission unit: USB4 interface with a transmission rate of 40Gbps and supports reverse charging (maximum 100W); HDMI2.1 interface supports 8K@60fps or 4K@120fps video output and supports HDMI eARC audio return function.
[0096] 1.2.7 Power Supply Module
[0097] Designed with dual power supply mode:
[0098] Power supply: It adopts a 21700 lithium battery pack (4 series and 2 parallel, capacity 10000mAh, voltage 14.8V), and outputs 3.3V (for sensors and microcontrollers), 5V (for display and touch module), 12V (for camera and audio module), and 24V (for projection module and air pump) through a DC-DC converter, with a battery life of ≥6 hours (continuous generation of 4K video);
[0099] External power supply: Supports 100-240V wide voltage input, outputs 14.8V / 6A DC power through AC-DC adapter, supports charging while using, charging time ≤2 hours.
[0100] 1.2.8 Human-Computer Interaction Control Module
[0101] Display Unit: The 3.5-inch LCD touchscreen uses an IPS panel with a resolution of 800×480 and a brightness of 300cd / m². It supports multi-touch and is used to display the interaction status (intent recognition result, feedback intensity), environmental parameters (light intensity, noise), content generation progress, and parameter configuration interface.
[0102] Control unit: Includes 6 physical buttons (power button, mode switch button, volume control button, brightness control button, feedback intensity control button, emergency stop button), supporting blind operation; the touch operation area supports gestures such as swiping, clicking, and zooming to adjust parameters such as content generation style and interaction sensitivity;
[0103] Remote control unit: Bluetooth 5.2 module communication distance ≤10m, supports pairing with smartphone APP to realize parameter configuration, content preview, remote control and other functions. The APP supports iOS and Android systems.
[0104] 1.3 Workflow
[0105] The specific steps of the digital media content generation method in this system are as follows:
[0106] S1: System Initialization and Environment Adaptation: After system startup, the environment adaptation module collects initial ambient light intensity, spatial location, and ambient noise data (collection frequencies are 10Hz, 100Hz, and 20Hz, respectively), the multimodal interaction acquisition module performs sensor self-checks and calibrations (calibration time ≤3s); the AI content generation engine loads pre-trained models and user-configured parameters, the real-time feedback processing module initializes the filtering algorithm and intent recognition model, and the system startup initialization is completed;
[0107] S2: Multimodal interaction data acquisition: Users initiate interaction through actions, voice or physiological changes. The multimodal interaction acquisition module simultaneously acquires action trajectory (joint point coordinate sequence), voice commands (audio signals), and physiological state (heart rate, skin conductance response) data. The acquisition delay is ≤50ms. The data is transmitted to the real-time feedback processing module via the SPI interface.
[0108] S3: Real-time Feedback and Intent Resolution: The real-time feedback processing module performs parallel noise reduction on the collected data, extracting motion features (joint movement speed, angle changes), speech features (MFCC coefficients, Mel spectrum), and physiological features (heart rate variability, skin conductance fluctuations); the interactive intent recognition unit resolves the user's core needs based on the feature data, outputting intent category and intensity parameters; the feedback signal generation unit generates corresponding tactile vibration signals and visual cue signals, transmitting them to the immersive output module to achieve instant interactive response (total latency ≤100ms).
[0109] S4: AI Adaptive Content Generation: The AI content generation engine receives interactive intent data and environmental parameters. The multimodal data fusion unit generates scene feature vectors through an attention mechanism. The adaptive content generation unit generates matching digital media content based on the feature vectors (such as adjusting the posture of virtual characters according to user actions and switching video scenes according to voice commands). The content optimization unit performs super-resolution reconstruction, frame rate interpolation, and noise suppression on the generated content to improve content quality.
[0110] S5: Immersive Content Presentation: The visual unit (VR headset or projector) of the immersive output module presents optimized 4K content, the auditory unit outputs spatial sound effects, and the haptic unit provides vibration and air pressure feedback in sync with the content plot. The multi-channel collaborative work builds an immersive experience. At the same time, the multimodal interaction acquisition module continuously captures subsequent user interaction data, and the real-time feedback processing module dynamically adjusts the feedback strategy to form a closed-loop interaction.
[0111] S6: Data Storage and Transmission: The storage and transmission module stores raw interactive data (JSON format), generated content data (JPEG / RAW image format, MP4 / ProRes video format), and configuration parameters in real time with a storage latency of ≤5ms; according to user settings, it transmits data to external devices (such as computers for post-editing, servers for data backup) via 5G / Wi-Fi 6 / USB4 interfaces with a transmission rate of ≥1Gbps.
[0112] This invention provides real-time interactive feedback: employing multimodal data parallel processing and lightweight algorithms, the interactive response latency is ≤150ms, and the feedback response time is ≤20ms, solving the problem of lag in traditional systems and enhancing the user's immersive experience; it simultaneously collects action, voice, physiological, and environmental data, and dynamically adjusts the fusion weights through an attention mechanism, achieving over 90% adaptability of content generation to user needs and scenarios, thus expanding the scope of application scenarios; the specific weight adjustment method is as follows:
[0113] Step 1: Scene type identification (a prerequisite for weight adjustment)
[0114] The system automatically determines the current "interaction scenario type" by analyzing user interaction behavior and environmental parameters in real time, providing a basis for weight allocation.
[0115] Recognition criteria: Based on the "interaction intent category" and "data activity level of each modality" output by the real-time feedback processing module and the "scene parameters" from the environment adaptation module, a scene recognition feature set is constructed.
[0116] Interaction intent categories: such as "switch scene", "adjust view" (action-related intent), "play music", "query information" (voice-related intent);
[0117] Modal data activity: By calculating the rate of change of each modal data (such as the joint movement speed of motion data, the energy value of voice data, and the heart rate variability rate of physiological data), we can determine which type of data is the core driver of the current interaction.
[0118] Environmental parameters: such as ambient noise intensity (when noise > 60dB, the credibility of voice data decreases and the weight is reduced), and spatial motion range (such as when the user moves a large distance in VR games, the weight of motion data is increased).
[0119] Recognition Model: A lightweight classifier (integrated into the multimodal data fusion unit) is used. It is trained with 500,000 sets of multimodal interaction-scene mapping samples, supports real-time scene type determination (determination latency ≤20ms), and outputs scene type labels (such as "action-driven scene", "voice-driven scene", "hybrid interaction scene").
[0120] Step 2: Basic weight allocation based on scenario type
[0121] Based on the scenario type, initial basic weights are first assigned, with the weight range strictly following the limitations specified in the patent document (action: 0.3-0.6, speech: 0.2-0.4, physiological: 0.1-0.2, environment: 0.05-0.15), ensuring that the sum of the weights for each modality is 1. The basic weight allocation for the two core scenarios specified in the patent is shown in Table 1:
[0122] Table 1
[0123] Scene type core mode Basic weight allocation (action / voice / physiology / environment) Application scenario examples Action-driven scene Motion data 0.6 / 0.2 / 0.1 / 0.1 VR games, interactive films (where users advance the plot through body movements), digital twin operations Voice-dominated scenarios Voice data 0.3 / 0.4 / 0.15 / 0.15 Virtual assistant interaction, voice-controlled live streaming, and remote command operation Hybrid Interactive Scenarios Actions + Voice 0.45 / 0.3 / 0.15 / 0.1 Virtual anchor interaction (anchor body movements + audience voice comments), VR social interaction (action communication + voice dialogue) Environmentally sensitive scenarios Environment + Physiology 0.35 / 0.25 / 0.2 / 0.2 Immersive therapeutic content (screen brightness adjusted according to ambient light intensity, music rhythm adjusted according to heart rate).
[0124] Step 3: Fine-tuning of weights based on attention mechanism (core step)
[0125] Based on the basic weight allocation, the weights are fine-tuned a second time through the attention mechanism algorithm to achieve "individual adaptation + real-time optimization" and avoid the problem of a one-size-fits-all approach to basic weights.
[0126] The core logic of the attention mechanism is to simulate the human attention allocation pattern—to assign higher attention weights (i.e., increase the fusion ratio) to modal data that are "more relevant to the current interaction intention and have higher data quality," and vice versa.
[0127] Specific implementation method:
[0128] Calculate the "confidence score" for each modality:
[0129] Motion data confidence: Based on the depth accuracy (±2%) and joint extraction error (≤3mm) of the RGB-D camera, if the motion data noise is low (residual after Kalman filtering <0.5mm), the confidence score is high (0.8-1.0), and the weight is increased by 10%-20% from the base value; if the user's motion amplitude is small and the data change rate is low, the confidence score is low (0.3-0.5), and the weight is decreased by 10%-15%.
[0130] Speech data confidence score: Based on the signal-to-noise ratio of the microphone array (≥60dB), the speech clarity score is obtained after spectral subtraction. If the ambient noise is low (≤40dB) and the speech command is complete, the confidence score is high (0.8-1.0), and the weight is increased by 10%-20%; if the noise is >60dB, the speech data is distorted, the confidence score is low (0.2-0.4), and the weight is decreased by 20%-30%.
[0131] Physiological data confidence: Based on the measurement errors of the heart rate sensor (accuracy ±1 bpm) and the skin conductance sensor (resolution 0.01 μS), if the physiological data is stable (heart rate variability <5%), the confidence score is high (0.7-0.9), and the weight is increased by 5%-10%; if the user's physiological state does not change significantly (skin conductance fluctuation <0.1 μS), the confidence score is low (0.3-0.5), and the weight remains at the base value or decreases by 5%.
[0132] Environmental data confidence level: Based on the measurement accuracy of environmental sensors (light intensity ±5%, noise ±0.5dB, positioning ±1cm), if the environmental parameters are stable (e.g., light intensity change rate <10% / s), the confidence score is high (0.7-0.9), and the weight is increased by 5%-10%; if the environment changes drastically (e.g., sudden strong light, sudden noise change), the confidence score is low (0.2-0.4), and the weight is decreased by 5%-10%.
[0133] Attention weight normalization: The "base weight × confidence score" of each modality is used as a temporary weight, and then normalized by the Softmax function to ensure that the final weight sum is 1, avoiding numerical overflow or weight imbalance.
[0134] Feature weighted fusion: The extracted features of each modality (action features: key point coordinate sequence; speech features: MFCC coefficient; physiological features: heart rate variability; environmental features: light intensity / noise / localization combination vector) are multiplied by the normalized attention weights, and then the vectors are concatenated to generate the final 1024-dimensional scene feature vector (the signal-to-noise ratio of the fused feature vector is ≥50dB).
[0135] Step 4: Real-time iterative optimization of weights (dynamically adjusted closed loop)
[0136] Weight allocation is not completed all at once, but iterates in real time as the interaction progresses, ensuring continuous adaptation to changes in user behavior and scenarios. The iteration logic is as follows:
[0137] Iteration triggering conditions:
[0138] Time-triggered: The weights are iterated every 50ms (synchronized with the multimodal data acquisition cycle);
[0139] Event triggering: When the interaction intent changes (such as switching from "action control" to "voice command"), environmental parameters change abruptly (such as noise increasing from 40dB to 70dB), or physiological state changes drastically (such as heart rate increasing from 70bpm to 120bpm), the weight recalculation is immediately triggered.
[0140] Iterative optimization is based on the following: The optimization objective is the "user feedback effect" of AI-generated content; the weighting strategy is dynamically adjusted using a reinforcement learning reward function.
[0141] If the user's subsequent interactions match the generated content well (e.g., the user intends to "turn around," the generated content switches the perspective synchronously, and the user does not make any secondary adjustments), then the current weight strategy scores highly, and this weight allocation logic is maintained in subsequent iterations.
[0142] If the user makes a secondary correction interaction (such as adjusting the volume by gesture after giving the voice command "play music"), the weight of the "volume control" related features in the voice data will be reduced, while the weight of the action data will be increased.
[0143] 3. Immersive multi-channel presentation: Integrating high refresh rate visuals, spatial audio, and precise haptic feedback to create a comprehensive immersive experience and meet the media industry's demand for high-quality immersive content.
[0144] 4. Adaptive content generation: Based on the model architecture of GAN and reinforcement learning, it supports the dynamic generation of various types of digital media content without human intervention, improving the efficiency and flexibility of content creation;
[0145] 5. High practicality: It integrates local storage, high-speed transmission, dual power supply mode and human-computer interaction control function, supports manual parameter adjustment and remote control, and can be widely used in VR content creation, interactive film and television, virtual live broadcast, digital entertainment and other scenarios, with high commercial value and application prospects.
[0146] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An immersive digital media content generation system based on real-time interactive feedback, characterized in that, It includes a multimodal interaction acquisition module, a real-time feedback processing module, an AI content generation engine, an environment adaptation module, an immersive output module, a storage and transmission module, a power supply module, and a human-computer interaction control module; The multimodal interaction acquisition module is used to simultaneously acquire the user's action data, voice data, and physiological perception multimodal data, and convert the multimodal data into standardized digital signals and send them to the real-time feedback processing module. The real-time feedback processing module is electrically connected to the multimodal interaction acquisition module and includes a data preprocessing unit, an interaction intent recognition unit, and a feedback signal generation unit. The real-time feedback processing module is used to perform real-time noise reduction, feature extraction, and intent parsing on the acquired multimodal data to generate multimodal instantaneous feedback signals. The AI content generation engine is electrically connected to the real-time feedback processing module and the environment adaptation module, and includes a multimodal data fusion unit, an adaptive content generation unit, and a content optimization unit. The AI content generation engine fuses multimodal real-time feedback signals and dynamically generates and optimizes digital media content based on the interaction intent and the environmental parameters collected by the environment adaptation module. The optimized digital media content is then output through the immersive output module. The environment adaptation module includes an ambient light sensor, a spatial positioning sensor, and an acoustic sensor. The environment adaptation module is used to collect ambient light intensity, system spatial position, and environmental noise data and send them to the AI content generation engine to provide an environment adaptation basis for content generation. The immersive output module is electrically connected to the AI content generation engine and includes a visual output unit, an auditory output unit, and a tactile feedback unit, used to present the generated digital media content synchronously through multiple channels. The storage and transmission module is electrically connected to the AI content generation engine and is used to store the original interactive data, generated content data and configuration parameters, and to achieve high-speed data transmission through the 5G / Wi-Fi 6 / USB4 interface; The power supply module is used to provide a stable voltage and supports two modes: continuous power supply and external power supply. The human-computer interaction control module is used to electrically connect with the AI content generation engine and the immersive output module, and to perform parameter configuration, content preview and interaction mode switching.
2. The immersive digital media content generation system based on real-time interactive feedback according to claim 1, characterized in that, The multimodal interaction acquisition module includes an action acquisition unit, an action acquisition unit, and a voice acquisition unit; The motion acquisition unit uses a 1080P RGB-D depth camera and a six-axis MEMS motion sensor, which can simultaneously capture the user's limb movements, posture changes and motion trajectories; The voice acquisition unit uses an array microphone, which supports long-distance voice acquisition and noise suppression; The physiological sensing unit employs a photoelectric heart rate sensor and a skin conductance sensor to capture changes in the user's physiological state.
3. The immersive digital media content generation system based on real-time interactive feedback according to claim 2, characterized in that, The real-time feedback processing module includes a noise reduction unit, a feature extraction unit, an interaction intent recognition unit, and a feedback signal generation unit. The noise reduction unit is used to reduce noise in multimodal data. The feature extraction unit is used to extract features from the denoised multimodal data to obtain speech features, action features, and physiological features, which are then input into the interaction intent recognition unit. The interaction intent recognition unit adopts a lightweight semantic understanding model based on Transformer to determine the activity level of each modality of data based on the input speech features, action features, and physiological features, and to determine the interaction intent category and intent strength. The feedback signal generation unit uses a pulse width modulation control algorithm to generate multimodal real-time feedback signals, including tactile feedback signals and visual cue signals, based on the type and intensity of the interaction intent.
4. The immersive digital media content generation system based on real-time interactive feedback according to claim 3, characterized in that, The workflow of the multimodal data fusion unit of the AI content generation engine includes the following steps: Step 1: Scene type identification: Based on the "interaction intent category" and "activity level of each modality" output by the real-time feedback processing module and the "scene parameters" collected by the environment adaptation module, the scene type is determined according to the trained classification and recognition model. The classification and recognition model adopts a lightweight classifier and is obtained through the multimodal interaction-scene mapping sample training set. The samples in the multimodal interaction-scene mapping sample training set are action, voice, physiological and environmental parameter data of different users with actual scene labels. Step 2: Basic weight allocation based on scene type: Based on the scene type, initial basic weights corresponding to action, voice, physiological and environmental parameters are assigned according to the preset scheme corresponding to the scene type. An attention mechanism fusion algorithm is used to perform feature weighted fusion of action, voice, physiological and environmental data. The weights are dynamically adjusted according to the interaction scene. The adaptive content generation unit adopts a model architecture that combines generative adversarial networks and reinforcement learning. It is obtained through a multimodal interaction-content mapping sample training set, in which the samples are action, speech, physiological and environmental parameter data of different users with actual content labels. The content optimization unit employs super-resolution reconstruction, frame rate interpolation, and noise suppression algorithms to improve the quality of content output.
5. The immersive digital media content generation system based on real-time interactive feedback according to claim 4, characterized in that, The method for dynamically adjusting weights based on the interaction scenario in step 2 includes the following steps: Step 3.1: Calculate the "confidence score" for each modality of data: Calculate the confidence score of motion data: Based on the motion data collected by the motion acquisition unit, if the residual of the motion data after Kalman filtering is <0.5mm, the confidence score is in the high range of 0.8-1.0, and the weight is increased by 10%-20% from the base value; if the residual of the motion data after Kalman filtering is >0.3mm, the confidence score is in the low range of 0.3-0.5, and the weight is decreased by 10%-15%; if the residual of the motion data after Kalman filtering is between 0.3-0.5mm, the base weight remains unchanged. Calculate the confidence score of the speech data: Based on the signal-to-noise ratio of the speech data collected by the speech acquisition unit, the speech clarity score is calculated after spectral subtraction. If the ambient noise is ≤40dB and the speech command is complete, the confidence score is in the high range of 0.8-1.0, and the weight increases by 10%-20%. If the ambient noise is >60dB, the speech data is considered distorted, and the confidence score is in the low range of 0.2-0.4, and the weight decreases by 20%-30%. If the ambient noise is between 0.3-0.5mm, the base weight remains unchanged. Calculating the confidence score of physiological data: Based on the measurement results of the physiological sensing unit, if the heart rate variability is <5% and the skin conductance fluctuation is <0.1μS, the confidence score is in the high range of 0.7-0.9, with the weight increased by 5%-10%; if the heart rate variability is >10% and the skin conductance fluctuation is >0.5μS, the confidence score is in the low range of 0.3-0.5, with the weight decreased by 5%. Calculate the confidence score of environmental data: Based on the measurement accuracy of environmental sensors, if the rate of change of environmental parameters is lower than the preset lower threshold, the confidence score is in the high range of 0.7-0.9, with the weight increased by 5%-10%; if the rate of change of environmental parameters is higher than the preset upper threshold, the confidence score is in the low range of 0.2-0.4, with the weight decreased by 5%-10%. Attention weight normalization: The "base weight × confidence score" of each modality is used as a temporary weight, and then normalized by the Softmax function to obtain the normalized temporary weights of each modality; Step 3.2, Feature Weighted Fusion: Multiply the extracted features of each modality by the corresponding normalized attention weights, and then concatenate the vectors to generate the final 1024-dimensional scene feature vector; Step 3.3: Real-time iterative optimization of weights. The current weights are updated after the iteration conditions are met. The iteration triggering conditions are: Time-triggered: The weights are iterated once at fixed intervals; Event triggering: When the interaction intent changes, environmental parameters exceed the preset environmental change threshold, or physiological state exceeds the preset physiological state change threshold, the weight update is triggered immediately; The iterative optimization aims to improve user feedback on content generated by the AI content generation engine, and dynamically adjusts the weighting strategy through a reinforcement learning reward function. If the user's subsequent interactions and the generated content have a high degree of matching, the current weight strategy will score high, and the current weight allocation logic will be maintained in subsequent iterations. If a user makes a secondary correction interaction, the relevant weights will be adjusted based on the correction behavior.
6. The immersive digital media content generation system based on real-time interactive feedback according to claim 1, characterized in that, The environment adaptation module is used to collect scene parameters using spatial positioning sensors, acoustic sensors, and ambient light sensors. The spatial positioning sensor uses a UWB positioning module, combined with an inertial navigation algorithm, to achieve real-time positioning of the system in three-dimensional space; the acoustic sensor uses a noise sensor for environmental noise detection and adaptive gain adjustment of audio content; the ambient light sensor is used to adjust the brightness and contrast of the visual output unit.
7. The immersive digital media content generation system based on real-time interactive feedback according to claim 1, characterized in that, The immersive output module includes a visual output unit, an auditory output unit, and a tactile feedback unit.
8. The digital media content generation method according to any one of claims 1-7, characterized in that, Includes the following steps: Step S1, System Initialization and Environment Adaptation: After the system starts, the environment adaptation module collects initial ambient light intensity, spatial location, and ambient noise data; the multimodal interaction acquisition module performs self-check and calibration; and the AI content generation engine loads the pre-trained model and configuration parameters to complete system initialization. Step S2, Multimodal Interaction Data Acquisition: The multimodal interaction acquisition module synchronously acquires the user's motion trajectory, voice commands, and physiological state data, and transmits them to the real-time feedback processing module through the SPI interface; Step S3: Real-time feedback and intent parsing: The real-time feedback processing module performs noise reduction and feature extraction on the collected data, and parses the user's core needs through the interactive intent recognition unit, generates instant feedback signals and transmits them to the immersive output module to realize interactive response; Step S4: AI Adaptive Content Generation: The AI content generation engine receives interactive intent data and environmental parameters, performs feature fusion through the multimodal data fusion unit, generates digital media content that matches the scene through the adaptive content generation unit, and improves the quality through the content optimization unit. Step S5: Immersive Content Presentation: The visual, auditory, and tactile units of the immersive output module work synchronously to present optimized digital media content through multiple channels, while capturing subsequent user interaction data in real time to form a closed-loop interaction. Step S6: Data storage and transmission: The storage and transmission module stores the original interactive data and generated content data in real time, and transmits them to external devices via wired or wireless means according to user settings.