Audio-visual linkage system based on AI sound scene regulation and control

By using an AI-based audiovisual linkage system to control soundscapes, environmental acoustics and biological characteristics are collected and analyzed in real time to generate director's intentions, enabling dynamic and personalized control of audiovisual content. This solves the problem of insufficient perception of user emotional state in existing immersive systems and enhances the depth and consistency of the immersive experience.

CN120932687AInactive Publication Date: 2025-11-11SHENZHEN WEIKING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511465769.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-11-11
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120932687A_ABST
    Figure CN120932687A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of man-machine interaction, and discloses an audio-visual linkage system based on AI sound scene regulation and control, and the system comprises a composite sensing assembly which collects environmental acoustics and audience biological characteristics; the emotional state evaluation engine is used for updating the emotional belief state of the audience based on the collected features, and actively calibrating the state by releasing an emotional probe and monitoring biological feature changes when the belief state is high in uncertainty; the hierarchical strategy engine is used for generating a director intention and a collaborative rendering action according to the belief state; and an audiovisual rendering component for executing a rendering action to regulate and control the audiovisual scene in a linkage manner. The method comprises the steps of collecting multi-modal data, evaluating and updating an emotional belief state, judging uncertainty and executing active calibration, performing hierarchical decision to generate a rendering instruction, and finally executing the instruction to realize audio-visual linkage and form closed-loop control. According to the invention, deep, real-time and adaptive coupling of the audio-visual content and the inherent emotion of the audience can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, specifically to an audiovisual linkage system based on AI-based soundscape control. Background Technology

[0002] Immersive technologies, such as virtual reality (VR), augmented reality (AR), and scene creation in theme parks, aim at creating a believable and engaging virtual or augmented environment for users. In these technologies, the synergy between soundscape and visual scene is key to building immersion.

[0003] Current mainstream immersive systems typically use pre-set scripts or simple triggering mechanisms to manage the synchronization of audiovisual content. For example, when a user enters a specific area or performs a specific action, the system plays pre-set audio and video clips. While this model achieves audiovisual linkage to some extent, it is inherently passive and static. The system cannot perceive the user's real-time emotional state and focus of attention, and the audiovisual feedback it provides is often disconnected from the user's intrinsic experience.

[0004] This separation leads to insufficient depth of interaction: regardless of whether the user feels tense, curious, or bored, the system presents content in the same way, failing to create dynamic and personalized emotional resonance. Furthermore, solutions that rely solely on environmental acoustic features for scene matching struggle to distinguish between soundscape changes caused by external environmental variations and acoustic events generated by the user's own behavior, limiting the accuracy and appropriateness of audiovisual mapping. Therefore, how to enable immersive systems to transcend preset logic, accurately perceive the user's deep emotional state, and adaptively and in real-time adjust audiovisual content to form a true interactive loop is a pressing technical problem that needs to be solved in the field.

[0005] Therefore, this invention proposes an audiovisual linkage system based on AI soundscape control to address the shortcomings of existing technologies. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an audiovisual linkage system based on AI soundscape control, which solves the problem of how to enable immersive systems to actively perceive and accurately assess the user's real-time emotional state, and accordingly adaptively and in real-time coordinate and control the visual scene and acoustic environment, thereby solving the problem of separation between audiovisual content and the user's intrinsic experience in existing technologies.

[0007] To achieve the above objectives, the present invention provides the following technical solution: an audiovisual linkage system based on AI soundscape control, the system comprising: Composite sensing components are used to collect environmental acoustic features and audience biometric features in real time; An emotional state assessment engine is used to take the environmental acoustic features and audience biometrics as observations and recursively update a probabilistic belief state that represents the audience's true emotional state through a Bayesian filtering process. A high-order director strategy engine is used to take the temporal sequence of the probabilistic belief states as input and select and generate director intentions in a preset narrative state space through a high-order strategy model trained by reinforcement learning. A low-level tactical execution engine is used to receive the director's intent and generate collaborative rendering actions based on the current probabilistic belief state through a low-level strategy model. An audiovisual rendering component is used to perform the collaborative rendering action. The audiovisual rendering component includes a visual rendering module for generating a visual scene and an active soundscape modulation module for adjusting ambient sound.

[0008] Preferably, the composite sensing component includes an acoustic sensing module and a biosensing module; The acoustic sensing module is used to perform Mel spectrum analysis on the collected ambient sound to obtain an acoustic feature map, and the acoustic sensing module includes a deep residual network, which is used to process the acoustic feature map to output ambient acoustic features. The biosensing module includes at least one non-invasive sensor selected from electroencephalogram (EEG) sensors, electrodermal transfer sensor (EDS) sensors, and photoplethysmography (PPG) sensors. The biosensing module is used to collect the biometric characteristics of the audience. The composite sensing component also includes a data synchronization unit, which is used to attach synchronization timestamps to the environmental acoustic features and the audience's biometric features.

[0009] Preferably, the emotional state assessment engine is used for: The environmental acoustic features and audience biometrics are used as observations, and the probabilistic belief state is recursively updated through a Bayesian filtering process. Furthermore, the emotional state assessment engine is also used to calculate the information entropy of the updated probabilistic belief state, and to control the active emotional calibration module to execute the calibration process when the information entropy is higher than a preset threshold.

[0010] Preferably, the proactive emotion calibration module is used to perform a calibration process, which includes: Control the active soundscape modulation module to release subthreshold acoustic signals as emotion probes; The biosensing module is controlled to monitor changes in the audience's biometrics triggered by the emotion probe; Furthermore, using the changes in the audience's biometrics as observational evidence, a Bayesian update is performed on the probabilistic belief state to generate the updated probabilistic belief state.

[0011] Preferably, the higher-order director strategy engine is used to receive a temporal sequence of probabilistic belief states and generate the director's intention; The advanced director strategy engine is trained through reinforcement learning. The advanced extrinsic rewards used in the reinforcement learning training are calculated based on the peak intensity of audience emotion and duration of attention after a narrative phase ends, as well as the final feedback data after the complete experience ends.

[0012] Preferably, the director's intent is a structured instruction, which includes the target emotional state, the task duration, and constraints on changes in the emotional state.

[0013] Preferably, the low-level tactical execution engine is used to receive the director's intent and probabilistic belief state, and generate collaborative rendering actions; The low-order tactical execution engine is optimized through training, which is based on low-order intrinsic rewards. The calculation of these low-order intrinsic rewards is based on the degree of conformity between the probabilistic belief state and the target probabilistic belief state vector contained in the director's intention, as well as the cost of calling the proactive emotion calibration module.

[0014] Preferably, the active soundscape modulation module is used for: Execute the active soundscape modulation instruction in the collaborative rendering action, wherein the active soundscape modulation instruction includes an ambient soundscape identifier and a volume balance parameter; Furthermore, when the active emotion calibration module performs the calibration process, the active soundscape modulation module is also used to generate and release the emotion probe using digital signal synthesis technology based on a set of specified physical parameters including center frequency, intensity, duration and envelope shape.

[0015] Preferably, the collaborative rendering action is a set of instructions that includes visual rendering instructions, active soundscape modulation instructions, and target execution timestamps; The visual rendering instructions include scene identifiers, lighting parameters, and color configuration schemes, and are used to control the visual rendering module. The active soundscape modulation instruction includes a soundscape identifier, a sound effect trigger event, and a volume parameter. The active soundscape modulation instruction is used to control the active soundscape modulation module. The visual rendering module and the active soundscape modulation module execute the visual rendering instructions and the active soundscape modulation instructions synchronously according to the target execution timestamp.

[0016] This invention also provides an audiovisual linkage method based on AI soundscape control, the method comprising the following steps: S1. Real-time acquisition of environmental acoustic features and audience biometrics; S2. Using the environmental acoustic features and the audience's biometric features as observations, a probabilistic belief state representing the audience's true emotional state is recursively updated through a Bayesian filtering process. S3. Calculate the uncertainty of the probabilistic belief state, and when the uncertainty is higher than a preset threshold, execute an active emotion calibration process. The active emotion calibration process includes releasing a subthreshold acoustic signal as an emotion probe, and using the monitored changes in audience biometrics caused by the emotion probe as new observational evidence to perform a Bayesian update on the probabilistic belief state. S4. Using the time sequence of the probabilistic belief states as input, generate director's intent through a high-order policy model; S5. Receive the director's intent and generate a collaborative rendering action based on the updated probabilistic belief state through a low-order strategy model. S6. Execute the aforementioned collaborative rendering action to coordinate and control the ambient sound and visual scene.

[0017] This invention provides an audiovisual linkage system based on AI-driven soundscape control. It offers the following advantages: 1. This invention establishes an emotional state assessment engine. This engine infers the probabilistic belief state of the viewer based on Bayesian filtering of acoustic and biometric features collected by a composite perception component. Furthermore, when the information entropy of this belief state exceeds a preset threshold, an active emotional calibration module is activated. This module calibrates the belief state by releasing emotional probes and monitoring the resulting changes in biometric features. This approach abandons the traditional deterministic, passive assessment of user states in human-computer interaction. Instead, it acknowledges and quantifies the uncertainty in the assessment process and actively detects and eliminates this uncertainty when necessary, thereby significantly improving the accuracy and robustness of perceiving the viewer's true and deep emotional state.

[0018] 2. This invention employs a hierarchical strategy engine. A high-level director strategy engine generates a macro-level directorial intent based on a long-term narrative goal and a temporal sequence of audience emotional states, while a low-level tactical execution engine is responsible for breaking down this intent into specific collaborative rendering actions based on real-time emotional states. This hierarchical decision-making architecture separates long-term strategic planning from short-term tactical execution, ensuring that each adjustment to the audiovisual content is not an isolated, reactive response, but rather serves a pre-defined, coherent macro-narrative goal. This effectively enhances the narrative depth, artistry, and emotional coherence of the entire immersive experience.

[0019] 3. This invention establishes an audiovisual rendering component that, based on collaborative rendering actions generated by a low-level tactical execution engine, coordinates and controls the active soundscape modulation module and the visual rendering module, and utilizes a unified high-precision clock to ensure spatiotemporal synchronization of audio-visual output. This overcomes the limitation of traditional systems where visual presentation and sound environment are independent, constructing a closed-loop control system where audiovisual content is driven by a unified strategy and strictly synchronized in time and space. This ensures that the visual scene perceived by the audience and the acoustic environment always maintain a high degree of consistency and mutual enhancement, thereby greatly improving the realism of the scene and the user's immersion. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the overall architecture of the audiovisual linkage system based on AI soundscape control according to the present invention; Figure 2 This is a flowchart of the audiovisual linkage method based on AI soundscape control according to the present invention; Figure 3 This is a schematic diagram of the proactive emotion calibration process of the present invention; Figure 4 This is a schematic diagram of the decision-making process of the hierarchical strategy engine of the present invention.

[0021] Among them, 100 is the composite perception component; 110 is the acoustic perception module; 120 is the biosensing module; 200 is the emotional state assessment engine; 210 is the active emotional calibration module; 300 is the high-level director strategy engine; 400 is the low-level tactical execution engine; 500 is the audiovisual rendering component; 510 is the active soundscape modulation module; and 520 is the visual rendering module. Detailed Implementation

[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Reference Figure 1 , Figure 1 This is a schematic diagram of the overall architecture of an audiovisual linkage system based on AI soundscape control according to an embodiment of the present invention. The present invention provides an audiovisual linkage system based on AI soundscape control, which is constructed as a closed-loop control structure, including a composite perception component 100, an emotion state evaluation engine 200, a high-level director strategy engine 300, a low-level tactical execution engine 400, and an audiovisual rendering component 500.

[0024] In one specific embodiment, the composite sensing component 100 is configured to acquire environmental acoustic features and audience biometrics in real time. The composite sensing component 100 includes an acoustic sensing module 110 and a biometric sensing module 120. The acoustic sensing module 110 consists of one or more microphone arrays for acquiring raw ambient sound audio signals. The acoustic sensing module 110 performs preprocessing on the acquired signals and extracts Mel frequency cepstral coefficients (MFCCs) as basic acoustic features through Mel spectral analysis. Subsequently, a pre-trained deep residual network (Res-Net) receives the MFCCs and outputs a high-dimensional acoustic feature vector containing spectral envelope, rhythmic patterns, and environmental sound signature identifiers. The biometric sensing module 120 includes one or more non-invasive sensors, such as electroencephalogram (EEG) sensors, field-skin resonance (GSR) sensors, or photoplethysmography (PPG) sensors, for acquiring audience biometric data. The acoustic sensing module 110 also includes a time synchronization unit to ensure that the acquired biometric data and acoustic feature data are timestamped.

[0025] The emotional state assessment engine 200 is connected to the composite perception component 100, and its input consists of the acoustic feature vector and biometric feature vector output by the composite perception component 100. The emotional state assessment engine 200 uses these feature vectors as observations and recursively updates a probabilistic belief state representing the viewer's true emotional state through a Bayesian filtering process. The probabilistic belief state is a vector whose dimensions correspond to a preset set of emotional states, and each value in the vector represents the probability that the viewer is in the corresponding emotional state.

[0026] The emotional state assessment engine 200 is also configured to quantify the uncertainty of probabilistic belief states and activate an active emotional calibration module 210 when the uncertainty exceeds a preset threshold. The quantification of uncertainty is achieved by calculating information entropy.

[0027] Reference Figure 3 , Figure 3 A calibration procedure executed by the active emotion calibration module 210 is illustrated. This procedure includes: First, the active emotion calibration module 210 instructs the active soundscape modulation module 510 to release a subthreshold acoustic signal serving as an emotion probe. The frequency and intensity of this signal are set to make it difficult for the audience to perceive explicitly. Simultaneously, the active emotion calibration module 210 and the instructing biosensing module 120 monitor changes in the audience's biometrics at high frequency within a very short time window after the release of the emotion probe. Finally, the active emotion calibration module 210 uses the monitored changes in biometrics as new, independent observational evidence to perform a Bayesian update on the current probabilistic belief state, generating an updated belief state with lower uncertainty.

[0028] Reference Figure 4 , Figure 4 The decision-making process of the hierarchical strategy engine in this embodiment is illustrated. The system includes a high-order director strategy engine 300 and a low-order tactical execution engine 400. The high-order director strategy engine 300 is connected to the emotional state evaluation engine 200 and receives the temporal sequence of probabilistic belief states output by the engine. Internally, the high-order director strategy engine 300 is implemented as a reinforcement learning-based agent. Its state space is a set of preset, discrete narrative states, and its action space is the generation of structured director intentions. The director intentions include the target emotional state, the task duration, and constraints on emotional state changes. The engine is trained through reinforcement learning, and the high-order extrinsic rewards used for training are calculated based on the peak intensity of the audience's emotions and the duration of their attention after a narrative phase, as well as the final feedback data after the complete experience.

[0029] The low-level tactical execution engine 400 is connected to the high-level director strategy engine 300 and the emotional state evaluation engine 200. It receives directorial intent from the high-level director strategy engine 300 and probabilistic belief states updated in real-time from the emotional state evaluation engine 200. The low-level tactical execution engine 400 maps the received directorial intent and the current emotional state to a specific collaborative rendering action through a low-level policy model. This collaborative rendering action is a set of instructions, including visual rendering instructions for controlling the visual rendering module 520 and active soundscape modulation instructions for controlling the active soundscape modulation module 510. The policy model of the low-level tactical execution engine 400 is optimized through training, based on a low-level intrinsic reward whose calculation integrates two aspects: the degree of conformity between the audience's emotional state and the target emotional state during the execution of the director's intent; and the cost of invoking the active emotional calibration module 210.

[0030] The audiovisual rendering component 500 is connected to the low-level tactical execution engine 400 to execute the cooperative rendering actions it generates. The audiovisual rendering component 500 includes an active soundscape modulation module 510 and a visual rendering module 520. The active soundscape modulation module 510 has a dual function: First, it executes active soundscape modulation instructions in collaborative rendering to adjust ambient sound; Secondly, upon receiving an instruction from the active emotion calibration module 210, the aforementioned emotion probe is generated and released.

[0031] The visual rendering module 520 is configured as an adaptive scene generator. Based on visual rendering instructions, it calls pre-made resources from a high-definition scene library or performs procedural real-time rendering. This visual rendering module 520 integrates ray tracing algorithms to optimize the material and lighting details of the rendered image. Simultaneously, the visual rendering module 520 also includes a display control module that uses a colorimetric model to dynamically calibrate the output of the display device, ensuring that in one embodiment, the colorimetric calibration accuracy reaches a certain level. And ensure that the switching of visual scenes and the changes in sound scenes are strictly synchronized in time.

[0032] Reference Figure 2 , Figure 2 This is a flowchart of the method of the present invention. In one working cycle, the method first involves the composite sensing component 100 acquiring acoustic and biometric features in real time and sending the feature data to the emotion state assessment engine 200. The emotion state assessment engine 200 recursively updates the probabilistic belief state and continuously calculates its information entropy. If the information entropy is higher than a threshold, an active emotion calibration process is executed to reduce uncertainty. The high-order director strategy engine 300 generates a macroscopic directorial intent based on the temporal sequence of belief states and sends it to the low-order tactical execution engine 400. The low-order tactical execution engine 400 combines this intent with the latest belief state to generate specific collaborative rendering actions and sends them to the audiovisual rendering component 500. The audiovisual rendering component 500 executes rendering instructions, coordinating and controlling the ambient sound and visual scene. The environmental changes caused by the rendering results and their impact on the audience's state are collected by the composite sensing component 100 as new input in the next time step, thus forming a continuously running closed loop. In one specific embodiment, the end-to-end latency from sound scene recognition to rendering response is less than 50 milliseconds.

[0033] Reference Figure 1 , Figure 1 This is a schematic diagram of the overall architecture of an audiovisual linkage system based on AI soundscape control according to an embodiment of the present invention. The specific implementation methods of each core component in the system will be described in detail below.

[0034] The composite sensing component 100 in this embodiment functions to collect environmental acoustic features and audience biometric features in parallel and in real time, and combine these feature data into a unified, time-synchronized multimodal data stream as input to the subsequent emotional state assessment engine 200. Physically and functionally, the composite sensing component 100 consists of an acoustic sensing module 110 and a biometric sensing module 120.

[0035] The acoustic sensing module 110 is configured to acquire ambient sound and extract ambient acoustic features. In one specific implementation, the hardware of the acoustic sensing module 110 consists of a multi-channel microphone array deployed within the interaction space to capture sound signals from different directions, facilitating subsequent sound source separation or noise reduction processing and improving the signal-to-noise ratio of the input signal. The acoustic sensing module 110 processes the acquired raw digital audio signals to extract feature vectors that characterize the sound content.

[0036] The feature extraction process begins with signal preprocessing. The raw audio stream is segmented into consecutive, partially overlapping short frames, for example, each frame is 25 milliseconds long with a frame shift of 10 milliseconds. A window function (e.g., a Hamming window) is applied to each frame to reduce spectral leakage. Subsequently, a Fast Fourier Transform (FFT) is performed on the windowed signal frames to obtain the power spectrum of that frame.

[0037] Next, Mel spectrum analysis is performed. The power spectrum obtained in the previous steps is passed through a set of Mel filter banks. This filter bank is characterized by a dense distribution of filters in the low-frequency region and a sparse distribution in the high-frequency region, to simulate the auditory perception characteristics of the human ear. After passing the power spectrum of each frame through this filter bank, the logarithm of the output energy of each filter is taken to obtain the logarithmic filter bank energy.

[0038] Finally, a discrete cosine transform (DCT) is performed on the logarithmic energy spectrum, and its lower-order coefficients are taken as the Mel frequency cepstral coefficients (MFCCs) of that frame. The MFCC vectors of multiple consecutive frames are combined to form a two-dimensional acoustic feature map, which is expanded in the time and coefficient dimensions.

[0039] To further analyze the acoustic features, the acoustic perception module 110 employs a pre-trained deep residual network (Res-Net) model. The two-dimensional MFCC feature map generated in the preceding steps is fed into the Res-Net model as input data. This model, through its deep convolutional layer structure, automatically learns and extracts higher-order, more abstract patterns from the MFCC map. The output of the Res-Net model is a high-dimensional acoustic feature vector. In this embodiment, the deep residual network (Res-Net) model is chosen because its unique residual-learning structure effectively solves the problems of vanishing or exploding gradients in deep neural networks by introducing shortcut connections. This allows the network to be built very deep, enabling it to learn and capture complex dependencies and abstract patterns across long time dimensions in the MFCC feature map, which is crucial for accurately identifying persistent environmental sound signatures or complex rhythmic patterns.

[0040] Each dimension of this vector corresponds to a different acoustic attribute, such as the shape of the spectral envelope, the rhythmic pattern of the sound signal, and environmental soundprint identifiers that can identify specific environments (such as streets or forests). This acoustic feature vector is the final output of the acoustic perception module 110.

[0041] The biosensing module 120 is configured to collect the biometrics of the audience. In one specific implementation, the biosensing module 120 integrates one or more non-invasive physiological sensors. These sensors may include an electroencephalogram (EEG) sensor for measuring electrical activity in the cerebral cortex; a skin conductance (GSR) sensor, also known as an electrical activity of the skin (EDA) sensor, for measuring changes in skin conductivity to reflect autonomic nervous system activity; and a photoplethysmography (PPG) sensor for measuring blood volume pulsation to extract heart rate and heart rate variability (HRV).

[0042] The biosensing module 120 filters and processes the raw biosignals collected from various sensors to extract quantified biofeature parameters, such as brainwave power in specific frequency bands (Alpha, Beta), amplitude and frequency of skin conductance response, and time-domain or frequency-domain indicators of heart rate variability.

[0043] To ensure the consistency of multimodal data, the biosensing module 120 also includes a data acquisition and synchronization unit. This unit attaches a synchronization timestamp from a high-precision system clock to the acoustic feature vector received from the acoustic sensing module 110 and the biometric parameters extracted from its various biosensors. This timestamp allows the system to precisely align data from different sources and with different sampling rates onto the same timeline. In each processing cycle, the system aggregates the latest acoustic feature vector and a set of biometric parameters into a single, multi-dimensional composite feature vector based on the timestamp. This composite feature vector is then passed to the emotion state assessment engine 200 for further processing.

[0044] Reference Figure 1 and Figure 3 The emotional state assessment engine 200 receives a time-synchronized multidimensional composite feature vector from the composite sensing component 100 within the system. The audience's emotional state is an internal latent variable that cannot be directly observed, while the acoustic and biometric features provided by the composite sensing component 100 are indirect external representations of this internal state, carrying uncertainty and noise. Therefore, the emotional state assessment engine 200 in this embodiment employs a Bayesian filtering framework to mathematically rigorously and recursively infer the probability distribution of the audience's emotional state from these incomplete and uncertain observational data.

[0045] The engine first defines the audience's emotions as a discrete set of states. Each element It represents a specific emotional category, such as focus, relaxation, tension, or joy. At any given moment... The engine maintains a probabilistic belief state vector. This vector represents the time at time... The audience is in every emotional state The probability. Updating this belief state is a recursive process, involving two steps: prediction and correction. Based on time... state of belief The engine first predicts the time using a state transition model. The prior beliefs. Then, when the moment is received... The composite feature vector (i.e., the observed value) At that time, the engine corrects prior beliefs using an observation model to obtain the time. Posterior belief state This recursive update process is described by the following formula: ; In the formula, and They represent the times at time 1 and 2 respectively. and Emotional state; It is after the update at the moment In an emotional state The probability of; At any moment In an emotional state The probability of; At any moment Observations received from the composite sensing component 100; It is a state transition model, representing the change in emotional state from... Transferred to The probability of this model is based on prior psychological knowledge or statistical data. It is an observational model, representing the emotional state as When the eigenvector is observed The probability is obtained by training the model on labeled data; It is a normalization factor whose value ensures that the sum of the probabilities of all emotional states is 1.

[0046] During system operation, belief state vector The probability distribution within the system can change. In some cases, the probability values ​​concentrate on a few states, indicating a high degree of certainty about the current emotion. In other cases, the probability values ​​are scattered across multiple states, making it impossible for the system to definitively determine the viewer's emotion. To quantify this uncertainty, the Emotional State Assessment Engine 200 calculates the current belief state. Information entropy The value of information entropy directly reflects the randomness or ambiguity of the probability distribution. Given a belief state... The formula for calculating its information entropy is: ; In the formula, It is a belief state vector The middle corresponds to the emotional state. The probability value; It represents the total number of emotional states.

[0047] The emotional state assessment engine has a preset information entropy threshold. Information entropy is calculated at each time step. Then, it will be compared with the threshold. If the calculated information entropy is higher than the threshold, it indicates that the system's confidence in assessing the audience's current emotional state is too low. At this time, the system activates its internal proactive emotional calibration module 210 to perform a proactive emotional calibration process. The goal of this process is to proactively acquire additional information to quickly reduce the uncertainty of the assessment.

[0048] Reference Figure 3 The specific steps of the active emotion calibration process are as follows. First, the active emotion calibration module 210 sends a command to the active soundscape modulation module 510 in the audiovisual rendering component 500, controlling it to generate and release a preset subthreshold acoustic signal as an emotion probe, denoted as... The frequency, loudness, and duration of the signal were set below the conscious perception threshold of the audience, but were sufficient to trigger potential physiological and psychological reactions.

[0049] In releasing emotional probes Simultaneously, the active emotion calibration module 210 instructs the biosensing module 120 to monitor changes in specific biometrics at a high sampling rate within a short time window immediately following probe release, denoted as... The amount of change It is a vector whose components can be parameters such as the amplitude of skin conductance response, energy changes in specific frequency bands of heart rate variability, etc.

[0050] Finally, the proactive emotion calibration module 210 will analyze the monitored changes in biometrics. As an independent and highly informative observational piece of evidence regarding the current state of high entropy belief... Perform a special Bayesian update to obtain an updated belief state with significantly reduced uncertainty. The calibration update process is described by the following formula: ; In the formula, The audience was in an emotional state before calibration. The probability of; It is the emotional state of the audience after calibration. The probability of; It is an emotion probe The amount of changes in biological characteristics that are triggered and actually monitored; It is a pre-trained probe response model that provides responses based on the known emotional state of the audience. Under these conditions, apply emotional probes It can trigger changes in specific biological characteristics. The conditional probability; It is the normalization factor, and its calculated value is: This is used to ensure the updated belief state vector The sum of all probabilities is 1. Updated belief state. It will replace the original This serves as the basis for subsequent decision-making by the strategy engine.

[0051] To establish the above observation model With probe response model Training data collection is required beforehand. In a specific data collection process, a group of participants are invited to experience a series of calibrated audiovisual stimuli that evoke specific emotions (such as joy, tension, sadness, etc.). During this process, the system simultaneously records the multimodal feature data (constituting the observations) collected by the composite perception component 100. (This is combined with the subjects' immediate emotional state self-reports after the stimulus ended to form an observation model for training.) The labeled dataset. For probe response models. This involves applying specific emotional probes to subjects when they are in different emotional baseline states. And record the amount of unconscious changes in their biological characteristics. Thus establishing different emotional states The probe response probability relationship.

[0052] Reference Figure 1 and Figure 4The decision-making system in this embodiment consists of a hierarchical strategy engine, which includes a high-level director strategy engine 300 and a low-level tactical execution engine 400. This hierarchical structure decouples long-term macro-narrative planning from short-term real-time tactical execution.

[0053] The advanced director strategy engine 300 is designed for long-term, strategic decision-making. In a specific implementation, this engine is implemented as a reinforcement learning-based agent. The agent's input is a time series of probabilistic belief states output by the emotion state evaluation engine 200. Its internal state space is defined as a set of discrete, pre-defined macro-narrative states, such as calm exploration, tension build-up, emotional climax, and relaxing reflection. The system maps the input belief state sequence to one of these narrative states.

[0054] The action of this higher-order agent is to generate a structured directorial intent, denoted as... The director's intention It is a data structure, specifically in the form of: ; In the formula, It is a target probabilistic belief state vector that defines the distribution of the audience's emotional state expected to be achieved at this stage of the narrative; It is the effective duration of the director's intention; It is a set of constraints in the process of emotional state change, such as limiting the maximum rate of change of the emotional state vector over time to ensure the smoothness of emotional transition.

[0055] This high-order agent is trained offline and optimized online using reinforcement learning. Its training is based on a reward function. It is a high-level extrinsic reward, which is within a director's intention. Duration Calculations are performed after completion. The calculation method is as follows: ; In the formula, It is the peak intensity of the target emotional component in the belief state during this stage, i.e. ,in, It is the target emotional state; It is the cumulative duration during which the audience is in a focused or related emotional state within this phase, and this value is calculated based on the belief state vector; and These are weighting coefficients used to balance different reward components. Furthermore, the final feedback data collected after the entire interactive experience concludes, such as user subjective rating questionnaires, is also used to further optimize the agent's strategy offline.

[0056] The function of the low-level tactical execution engine 400 is to receive the director's intentions issued by the high-level director strategy engine 300. And based on the real-time updated belief state This is broken down into a series of specific, real-time rendering instructions. This low-level tactical execution engine 400 can also be implemented as a reinforcement learning agent. At each time step... The input to the agent is a statement containing the current director's intent. and current belief state tuples.

[0057] The low-level agent's action is a collaborative rendering action. This action is one that includes visual rendering instructions. and active soundscape modulation commands The instruction set. Visual rendering instructions. This includes the scene identifier to be rendered, lighting parameters, color scheme, etc. Active sound and background modulation commands. This includes the background soundscape identifier to be played, the instructions to trigger specific sound effect events, and the volume parameters of each sound source.

[0058] The training and optimization of this low-order agent are based on a low-order intrinsic reward function. The reward is given at each time step. The calculation is performed to guide the agent's immediate behavior. This reward function combines the approach to the goal with the cost of resource usage. Its calculation method is as follows: ; In the formula, At any moment The current belief state vector; Is this the director's current intention? The target belief state vector defined in [the document / reference]; It is the square of the Euclidean distance between two vectors, and a negative value of this term rewards the degree of closeness between the current state and the target state. It is an indicator variable, if at time... If the proactive emotional calibration process is triggered, then ,otherwise . This is a positive constant representing the cost of invoking a single proactive emotion calibration process. This term exists to prevent the system from using emotion probes too frequently, thus striking a balance between achieving accurate assessments and minimizing potential audience distraction.

[0059] In another alternative embodiment, the low-order tactical execution engine 400 can also be implemented through supervised learning. In this approach, a training dataset needs to be constructed, containing (director's intent, current belief state) as input features and (cooperative rendering actions) designed by human experts as output labels. Then, a model such as a multi-output regressive neural network is used for training. After training, the model can directly map high-order intents and current states to specific rendering instructions. Compared to reinforcement learning, this training process is more direct, but it relies on high-quality expert-annotated data.

[0060] Reference Figure 1 The audiovisual rendering component 500 is the final execution layer of the closed-loop control system in this embodiment. Its function is to receive and execute collaborative rendering action instructions from the low-level tactical execution engine 400. The audiovisual rendering component 500 consists of an active soundscape modulation module 510 and a visual rendering module 520, which work together to generate spatiotemporally synchronized audiovisual output that conforms to the system's decision-making intent.

[0061] The active soundscape modulation module 510 performs two independent functions triggered by different upstream modules within the system. First, it performs routine ambient soundscape rendering. It receives the active soundscape modulation instruction portion from the collaborative rendering action instructions of the low-level tactical execution engine 400. This instruction includes ambient soundscape identifiers, trigger events for specific sound effects, and volume balance parameters for multi-channel audio. Based on this, the active soundscape modulation module 510 loads corresponding audio samples from a soundscape resource library and mixes and processes them through a spatial audio renderer to generate an immersive, multi-channel auditory environment.

[0062] Secondly, the active soundscape modulation module 510 precisely generates and releases the emotion probe signal. It receives a specific request from the active emotion calibration module 210 within the emotion state assessment engine 200. This request specifies the precise physical parameters of the emotion probe signal, including center frequency, intensity (decibels), duration, and envelope shape. Based on these parameters, the active soundscape modulation module 510 generates the subthreshold acoustic signal using digital signal synthesis technology and controls the output speaker to release it into the environment.

[0063] The function of the visual rendering module 520 is to generate the final visual image based on the visual rendering instructions issued by the low-level tactical execution engine 400. Internally, the visual rendering module 520 is implemented as an adaptive scene generator. When the received visual rendering instruction contains a scene identifier, the generator calls the corresponding pre-made 3D model, texture, material, and lighting data from a local high-definition scene library. When the instruction contains a set of procedural parameters, the generator uses techniques such as fractal algorithms or L-systems to perform procedural real-time rendering to generate non-repetitive, dynamic visual scenes.

[0064] To enhance the realism of the rendered output, the visual rendering module 520 employs a ray tracing algorithm in the final image compositing stage. For each pixel in the scene, the algorithm traces the path of light rays backward from the virtual camera, simulating the reflection, refraction, and shadow effects that occur when light interacts with the surfaces of objects in the scene. This process significantly improves the texture of the rendered image, the smoothness of the light and shadow transitions, and the overall physical realism.

[0065] The visual rendering module 520 also includes a display control unit to ensure the accuracy of output colors and the spatiotemporal synchronization of audiovisual content. This display control unit employs a standard colorimetric model, such as the CIE-1976 (Lab*) color space, to dynamically calibrate the output of the display device. Before sending the rendered frame to the display device, the display control unit compares the pixel colors in the frame buffer with the target color profile defined in the instructions and calculates the color difference between the two. Color difference The calculation formula is: ; In the formula, and These are the coordinates of the target color and the actual output color in the Lab* color space, respectively. The display control unit adjusts the output by generating and applying a color correction lookup table (LUT) to ensure the average color difference. The value is no greater than 1.5. To achieve spatiotemporal synchronization, both the visual rendering module 520 and the active soundscape modulation module 510 are synchronized with a unified high-precision system clock. Each collaborative rendering action command issued by the low-level tactical execution engine 400 carries a target execution timestamp. The two modules precisely start their respective rendering processes according to this timestamp, ensuring that the switching of visual scenes and changes in soundscape occur at the same time.

[0066] Reference Figure 2 , Figure 2 This is a flowchart of an audiovisual linkage method based on AI soundscape control according to an embodiment of the present invention. This section will be combined with... Figures 1 to 4The system components shown in the diagram illustrate the specific implementation process of this method. This method embodies a continuously running, closed-loop control process encompassing perception, evaluation, decision-making, and execution.

[0067] At the start of a work cycle, the first step in this process is perception. The composite perception component 100, through its acoustic perception module 110 and biosensing module 120, continuously and in parallel acquires sound signals from the environment and physiological signals from the audience. These raw signals are processed within the modules into a time-synchronized, multi-dimensional composite feature vector and sent to the emotional state assessment engine 200.

[0068] The second step is evaluation. After receiving the composite feature vector, the emotional state evaluation engine 200 uses it as new observational evidence to perform a Bayesian update on the probabilistic belief state vector representing the audience's emotional state, thereby obtaining the latest posterior probability distribution reflecting the emotional state at the current moment.

[0069] The third step is uncertainty assessment and calibration. After updating the belief state, the emotion state assessment engine 200 immediately calculates the information entropy of the belief state vector. This information entropy value is compared with a preset threshold. If the value is lower than or equal to the threshold, it indicates that the system has sufficient confidence in the assessment of the current emotion state, and the process proceeds directly to the next step. If the value is higher than the threshold, it indicates that the uncertainty of the assessment result is too high, and the system will trigger an active emotion calibration process. In this calibration process, the emotion state assessment engine 200 instructs the active soundscape modulation module 510 to release an emotion probe, and based on the amount of biometric changes triggered by the probe monitored by the biosensing module 120, performs an additional Bayesian update on the belief state to obtain a calibrated belief state with lower information entropy.

[0070] The fourth step is high-level decision-making. The high-level director strategy engine 300 continuously receives and caches the temporal sequence of belief states (processed or uncalibrated) output by the emotion state assessment engine 200. This engine matches this sequence with its internal narrative state model and, based on its reinforcement learning strategy, selects and generates a macro-level, long-term director intention. This director intention defines the narrative goal for the next stage and is then distributed to the low-level tactical execution engine 400.

[0071] The fifth step is low-level decision-making. The low-level tactical execution engine 400 receives the director's intent from the high-level engine and combines it with the latest real-time belief state from the emotional state assessment engine 200. Through its low-level strategy model, it instantly generates a specific and executable collaborative rendering action instruction.

[0072] The sixth step is execution and rendering. The audiovisual rendering component 500 receives and parses the collaborative rendering action instruction. Its internal active soundscape modulation module 510 and visual rendering module 520 execute their respective rendering tasks in parallel according to the synchronization timestamp carried in the instruction, and output ambient sound and visual scenes that match the content of the instruction.

[0073] The seventh step is the closed loop. The audio-visual content output by the audiovisual rendering component 500 constitutes a new environment and acts on the audience, potentially altering their emotional state. In the next time step, the composite perception component 100 will collect the acoustic features of this new environment and the audience's latest biometrics. This new data will serve as input to initiate the next round of the perception-evaluation-decision-execution loop, thus forming a continuously adaptive closed-loop control system.

[0074] In a specific embodiment, each stage of this technical solution has defined performance indicators. From the moment the acoustic signal enters the acoustic sensing module 110 to the point where the deep residual network within the acoustic sensing module 110 outputs a usable acoustic feature vector, the latency is no more than 25 milliseconds. From the moment the low-order tactical execution engine 400 issues a collaborative rendering action command to the audiovisual rendering component 500 completing the switching and presentation of the corresponding audio-visual scene, the response time is less than 50 milliseconds. These indicators ensure that the system can respond promptly to changes in the environment and the audience's state.

[0075] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An audiovisual linkage system based on AI-driven soundscape control, characterized in that, The system includes: Composite sensing components are used to collect environmental acoustic features and audience biometric features in real time; An emotional state assessment engine is used to take the environmental acoustic features and audience biometrics as observations and recursively update a probabilistic belief state that represents the audience's true emotional state through a Bayesian filtering process. A high-order director strategy engine is used to take the temporal sequence of the probabilistic belief states as input and select and generate director intentions in a preset narrative state space through a high-order strategy model trained by reinforcement learning. A low-level tactical execution engine is used to receive the director's intent and generate collaborative rendering actions based on the current probabilistic belief state through a low-level strategy model. An audiovisual rendering component is used to perform the collaborative rendering action. The audiovisual rendering component includes a visual rendering module for generating a visual scene and an active soundscape modulation module for adjusting ambient sound.

2. The audiovisual linkage system based on AI soundscape control according to claim 1, characterized in that, The composite sensing component includes an acoustic sensing module and a biosensing module; The acoustic sensing module is used to perform Mel spectrum analysis on the collected ambient sound to obtain an acoustic feature map, and the acoustic sensing module includes a deep residual network, which is used to process the acoustic feature map to output ambient acoustic features. The biosensing module includes at least one non-invasive sensor selected from electroencephalogram (EEG) sensors, electrodermal transfer sensor (EDS) sensors, and photoplethysmography (PPG) sensors. The biosensing module is used to collect the biometric characteristics of the audience. The composite sensing component also includes a data synchronization unit, which is used to attach synchronization timestamps to the environmental acoustic features and the audience's biometric features.

3. The audiovisual linkage system based on AI soundscape control according to claim 1, characterized in that, The emotional state assessment engine is used for: The environmental acoustic features and audience biometrics are used as observations, and the probabilistic belief state is recursively updated through a Bayesian filtering process. Furthermore, the emotional state assessment engine is also used to calculate the information entropy of the updated probabilistic belief state, and to control the active emotional calibration module to execute the calibration process when the information entropy is higher than a preset threshold.

4. The audiovisual linkage system based on AI soundscape control according to claim 3, characterized in that, The proactive emotion calibration module is used to execute a calibration process, which includes: Control the active soundscape modulation module to release subthreshold acoustic signals as emotion probes; The control biosensing module monitors the amount of change in the audience's biometrics triggered by the emotion probe; Furthermore, using the changes in the audience's biometrics as observational evidence, a Bayesian update is performed on the probabilistic belief state to generate the updated probabilistic belief state.

5. The audiovisual linkage system based on AI soundscape control according to claim 1, characterized in that, The higher-order director strategy engine is used to receive a temporal sequence of probabilistic belief states and generate the director's intention. The advanced director strategy engine is trained through reinforcement learning. The advanced extrinsic rewards used in the reinforcement learning training are calculated based on the peak intensity of audience emotion and the duration of attention after a narrative phase ends, as well as the final feedback data after the complete experience ends.

6. The audiovisual linkage system based on AI soundscape control according to claim 5, characterized in that, The director's intent is a structured instruction that includes the target emotional state, the duration of the task, and constraints on changes in the emotional state.

7. The audiovisual linkage system based on AI soundscape control according to claim 1, characterized in that, The low-level tactical execution engine is used to receive the director's intent and probabilistic belief states, and generate collaborative rendering actions; The low-order tactical execution engine is optimized through training, which is based on low-order intrinsic rewards. The calculation of these low-order intrinsic rewards is based on the degree of conformity between the probabilistic belief state and the target probabilistic belief state vector contained in the director's intention, as well as the cost of calling the proactive emotion calibration module.

8. The audiovisual linkage system based on AI soundscape control according to claim 4, characterized in that, The active soundscape modulation module is used for: Execute the active soundscape modulation instruction in the collaborative rendering action, wherein the active soundscape modulation instruction includes an ambient soundscape identifier and a volume balance parameter; Furthermore, when the active emotion calibration module performs the calibration process, the active soundscape modulation module is also used to generate and release the emotion probe using digital signal synthesis technology based on a set of specified physical parameters including center frequency, intensity, duration and envelope shape.

9. The audiovisual linkage system based on AI soundscape control according to claim 1, characterized in that, The collaborative rendering action is a set of instructions that includes visual rendering instructions, active soundscape modulation instructions, and target execution timestamps; The visual rendering instructions include scene identifiers, lighting parameters, and color configuration schemes, and are used to control the visual rendering module. The active soundscape modulation instruction includes a soundscape identifier, a sound effect trigger event, and a volume parameter. The active soundscape modulation instruction is used to control the active soundscape modulation module. The visual rendering module and the active soundscape modulation module execute the visual rendering instructions and the active soundscape modulation instructions synchronously according to the target execution timestamp.

10. An audiovisual linkage method based on AI-driven soundscape control, applied to the system described in any one of claims 1-9, characterized in that, The method includes the following steps: S1. Real-time acquisition of environmental acoustic features and audience biometrics; S2. Using the environmental acoustic features and the audience's biometric features as observations, a probabilistic belief state representing the audience's true emotional state is recursively updated through a Bayesian filtering process. S3. Calculate the uncertainty of the probabilistic belief state, and when the uncertainty is higher than a preset threshold, execute an active emotion calibration process. The active emotion calibration process includes releasing a subthreshold acoustic signal as an emotion probe, and using the monitored changes in audience biometrics caused by the emotion probe as new observational evidence to perform a Bayesian update on the probabilistic belief state. S4. Using the time sequence of the probabilistic belief states as input, generate director's intent through a high-order policy model; S5. Receive the director's intent and generate a collaborative rendering action based on the updated probabilistic belief state through a low-order strategy model. S6. Execute the aforementioned collaborative rendering action to coordinate and control the ambient sound and visual scene.