Immersive audio-video follow-up adjustment method and system

By building a multi-dimensional perception system and deep neural network model, collecting and analyzing user and audio and video content data in real time, and generating a dynamic audio parameter adjustment solution, the problem of audio effect mismatch in existing technologies is solved, and precise adaptation and personalized adjustment of the immersive audio experience are achieved.

CN120832119AInactive Publication Date: 2025-10-24SHENZHEN ZIDOO TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511326369.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-10-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies find it difficult to accurately adapt audio effects when users change dynamically, resulting in the destruction of the immersive sound field and poor audio experience, and are unable to meet diverse immersive audio needs.

Method used

By building a multi-dimensional perception system to collect user head movement, body posture, physiological state and audio and video content feature data in real time, using deep neural network models to perform information fusion and spatiotemporal correlation analysis, generating dynamic audio parameter adjustment solutions, and combining user feedback to build a closed-loop adjustment system, we can achieve deep collaborative adaptation of audio with user status and content.

Benefits of technology

It significantly improves the naturalness and adaptability of the immersive audio experience, takes into account both universality and individual differences, provides personalized audio adjustment, and solves the problem of scene mismatch in traditional audio adjustment methods under dynamic changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832119A_ABST
    Figure CN120832119A_ABST
Patent Text Reader

Abstract

The invention is applicable to the field of intelligent audio adjustment, and provides an immersive audio-video follow-up adjustment method and system, and the method comprises the steps: constructing a multi-dimensional perception system, and collecting multi-source information in real time; carrying out fusion processing on the collected multi-source information based on a deep neural network model, mining a dynamic mapping relation between a user state and the video content through space-time correlation analysis, and identifying a user interaction intention and an emotional tone and a space scene attribute of the video content; according to a fusion processing result, calling a dynamic parameter adjustment engine, and generating an audio parameter adjustment scheme in real time; a user experience feedback closed loop is constructed, a visual attention area of a user is collected through eye movement tracking equipment, and personalized adjustment preference parameters are generated; according to the method and the device, the audio effect is accurately matched with the user state and the audio and video content, the naturalness and the adaptability of immersive experience are remarkably improved, universality and individual differences are considered, and the audio experience which is more suitable for scenes and needs of the user is brought to the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of intelligent audio adjustment, and particularly relates to an immersive audiovisual audio follow-up adjustment method and system. BACKGROUND

[0002] In the field of immersive audiovisual experience, with the continuous improvement of users' pursuit of immersive experience, the requirements for audio effects are increasingly demanding. On the one hand, in simulating real audio environment, the prior art is difficult to accurately adapt to the dynamic changes of users. When the user moves the body or changes the head orientation during the process of watching a movie or playing a game, the spatial positioning and effect of the audio cannot be timely and accurately adjusted, resulting in the destruction of the originally created immersive sound field, the deviation of the sound orientation and the stereo sense, and the serious influence on the user experience.

[0003] On the other hand, different audiovisual contents have unique audio characteristics. From the exciting action movie music to the delicate dialogue scene, from the rhythm of the game sound effect to the melodious music program, the traditional audio adjustment method cannot dynamically optimize the audio parameters according to the real-time changes of the content, and it is difficult to highlight the audio details and emotional expression of various contents, and it is difficult to fully meet the diversified needs of users for immersive audio experience.

[0004] In view of the above technical problems, the present application aims to solve the above-mentioned problems in the prior art, and provides a method for realizing immersive audiovisual audio follow-up adjustment, so as to significantly improve the immersion and adaptability of audio experience. SUMMARY

[0005] The present application aims to provide an immersive audiovisual audio follow-up adjustment method and system, which aims to solve the problems proposed in the background art.

[0006] The present application is realized in this way. On the one hand, an immersive audiovisual audio follow-up adjustment method, the method comprising: constructing a multi-dimensional perception system to collect multi-source information in real time; The multi-source information includes the head six-degree-of-freedom motion data, the body posture dynamic change data, the physiological state parameters, and the audio feature data and video frame feature data of the currently played audiovisual content of the user; Based on the deep neural network model, the collected multi-source information is fused and processed, the dynamic mapping relationship between the user state and the audiovisual content is mined through the spatio-temporal correlation analysis, and the user interaction intention and the emotional tone and spatial scene attribute of the audiovisual content are identified; According to the fusion processing result, a dynamic parameter adjustment engine is called, and an immersive experience evaluation model is combined to generate an audio parameter adjustment scheme in real time; The audio parameters include three-dimensional space sound image positioning parameters, dynamic range compression ratios, multi-channel gain distribution coefficients and psychoacoustic optimization parameters, which are used to realize deep collaborative adaptation of audio effects, user states and video and audio content, and construct an immersive audio sound field with scene adaptive characteristics. A user experience feedback closed loop is constructed, visual attention areas of the user are collected through an eye tracking device, and personalized adjustment preference parameters are generated in combination with adjustment instructions triggered by the user.

[0007] As a further scheme of the application, the multi-dimensional perception system is constructed, and the real-time collection of multi-source information specifically includes: Six-degree-of-freedom motion data of the head of the user are collected through a millimeter wave radar and an inertial measurement unit; Dynamic change data of the body posture of the user and micro-motion characteristics of the body surface are collected through a distributed pressure sensor array and infrared thermal imaging technology; Physiological state parameters of the heart rate variability and the respiration rate of the user are collected through wearable biosensors; Real-time spectrum waterfall diagrams, sound pressure level dynamic curves of an audio stream, and depth map sequences and dynamic object motion vector fields of video frames are synchronously collected through a video and audio content analysis interface.

[0008] As a further scheme of the application, the multi-source information collected is fused and processed based on a deep neural network model, a dynamic mapping relationship between the user state and the video and audio content is mined through spatio-temporal correlation analysis, and the user interaction intention and the emotional tone and spatial scene attribute of the video and audio content are identified, and the specific steps include: A transformer model with an attention mechanism is used to perform spatio-temporal alignment on the multi-source information; Physiological state parameters and audio feature data are correlated and analyzed through a pre-trained emotion recognition sub-model, and real-time emotional state labels of the user are output; A scene classifier is constructed, and the spatial scene type and acoustic characteristics of the video and audio content are identified based on video frame feature data and audio spatial features.

[0009] As a further scheme of the application, a dynamic parameter adjustment engine is called according to the fusion processing result, a preset immersive experience evaluation model is combined, and an audio parameter adjustment scheme is generated in real time, and the specific steps include: When it is detected that the head of the user performs a large turning action, a sound image positioning compensation value is calculated in real time by the dynamic parameter adjustment engine, multi-channel delay difference and gain difference are adjusted, and the azimuth and elevation of the spatial sound image are synchronously deflected with the user's visual angle; If the video and audio content is identified as a thrilling scene and the physiological parameters of the user show a nervous state, the high-frequency gain is reduced, the low-frequency dynamic range is increased, and pre-echo suppression processing is synchronously realized; For the open field scene, a virtual sound amplification algorithm is called to improve the diffusion coefficient of the multi-channel signal, and the coverage range of the sound field is dynamically adjusted according to the distance change of the user's body posture.

[0010] As a further scheme of the present application, when the video and audio content is identified as a thriller scene and the user's physiological parameters show a tense state, the high-frequency gain is reduced, the low-frequency dynamic range is increased, and the pre-echo suppression processing is simultaneously implemented, specifically comprising: The adjustment engine receives the scene and user state label, analyzes the related parameters of the high-frequency gain, low-frequency dynamic range and pre-echo suppression; Separate the high-frequency signal, linearly attenuate it with a digital attenuator according to the threshold, and control the signal peak value; Separate the low-frequency signal and calculate the dynamic range; Expand the dynamic range to the preset interval, increase the gain for weak signals, and retain the amplitude for strong signals; Based on the preset leading audio frame, identify potential pre-echoes, and use an adaptive filter for selective attenuation; Compare the processed audio parameters with the standard audio parameters, and correct the excessive attenuation or distortion; If the user's tense state is relieved, gradually return to the regular adjustment parameters corresponding to the thriller scene.

[0011] As a further scheme of the present application, for the open field scene, a virtual sound amplification algorithm is called to improve the diffusion coefficient of the multi-channel signal, and the coverage range of the sound field is dynamically adjusted according to the distance change of the user's body posture, specifically comprising: After recognizing the open field scene, activate the virtual sound amplification algorithm, and simultaneously obtain the original diffusion coefficient of the multi-channel signal and the distance parameter of the user's body posture; Analyze the spatial characteristics of the multi-channel signal, determine the diffusion coefficient adjustment interval, and increase the diffusion coefficient by adding differential reverberation to monitor and ensure that the diffusion effect meets the acoustic characteristics of the open field; Quantify the distance parameter of the user's body posture to generate a comprehensive distance parameter, and establish a mapping relationship model between the comprehensive distance parameter and the sound field coverage range; According to the mapping relationship model, calculate the target sound field coverage radius, adjust the gain and delay parameters of each channel to match the target radius, and update the parameters according to the user's real-time movement state to realize continuous adjustment of the sound field range; Compare the actual sound field data with the expected effect, and if the deviation is out of limit, correct the parameters until the scene and user state are adapted.

[0012] As a further scheme of the present application, the user experience feedback closed loop is constructed, the user's visual attention area is collected through eye tracking equipment, and personalized adjustment preference parameters are generated by combining the user's active triggered adjustment instructions, specifically comprising: An eye movement tracking device is started to capture the eye movement trajectory of a user in real time, and a visual attention area is analyzed, while an audio adjustment instruction actively triggered by the user through a gesture, voice or physical key is collected; The visual attention area coordinates and the residence time length are extracted, the audio adjustment instruction actively adjusted by the user is subjected to semantic analysis or action recognition, and is converted into a standardized adjustment parameter requirement; An association model of feedback data and audio parameters is established, and the matching relationship between the visual attention area and the audio features is analyzed; The frequency and amplitude of the active adjustment instruction of the user are counted, the adjustment tendency of the user to different scenes is extracted, and an initial preference feature set is formed; Based on the initial preference feature set, combined with the basic adjustment parameters of the current audio-visual scene, personalized adjustment preference parameters are generated through a weighting algorithm; The personalized adjustment preference parameters are input into a deep neural network model for incremental training, and the understanding of the subjective demand of the user by the model is updated.

[0013] As a further scheme of the application, in a further aspect, an immersive audio-visual audio follow-up adjustment system, a multi-dimensional perception system module is used to build a multi-dimensional perception system and collect multi-source information in real time; An identification module is used to fuse the collected multi-source information based on a deep neural network model, mine the dynamic mapping relationship between the user state and the audio-visual content through space-time correlation analysis, and identify the user interaction intention and the emotional tone and spatial scene attribute of the audio-visual content; A first generation module is used to generate an audio parameter adjustment scheme in real time according to the fusion processing result, call a dynamic parameter adjustment engine, and combine a preset immersive experience evaluation model; A second generation module is used to build a user experience feedback closed loop, collect the visual attention area of the user through an eye movement tracking device, combine the adjustment instruction actively triggered by the user, and generate personalized adjustment preference parameters.

[0014] As a further scheme of the application, the multi-dimensional perception system module specifically includes: A first acquisition unit is used to collect six-degree-of-freedom motion data of the head of the user through a millimeter wave radar and an inertial measurement unit; A second acquisition unit is used to collect dynamic change data of the body posture of the user and body surface micro-motion features by using a distributed pressure sensor array and an infrared thermal imaging technology; A third acquisition unit is used to collect physiological state parameters of the heart rate variability and the respiration rate of the user by using a wearable biological sensor; A fourth acquisition unit is used to synchronously collect a real-time frequency spectrum waterfall diagram, a sound pressure level dynamic curve of an audio stream, and a depth map sequence and a dynamic object motion vector field of a video frame through an audio-visual content analysis interface.

[0015] The application provides an immersive audiovisual audio follow-up adjustment method and system, which realizes accurate adaptation of audio effects, user state and audiovisual content, significantly improves the naturalness and adaptability of immersive experience, considers universality and individual differences, and brings users more scene-adapted and self-demand-adapted audio experience. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 It is a main flowchart of an immersive audiovisual audio follow-up adjustment method.

[0017] Figure 2 It is a flowchart of constructing a multi-dimensional perception system to collect multi-source information in real time in an immersive audiovisual audio follow-up adjustment method.

[0018] Figure 3 It is a flowchart of fusing and processing the collected multi-source information based on a deep neural network model, mining the dynamic mapping relationship between the user state and the audiovisual content through space-time correlation analysis, identifying the user interaction intention, and the emotional keynote and spatial scene attribute of the audiovisual content in an immersive audiovisual audio follow-up adjustment method.

[0019] Figure 4 It is a flowchart of calling a dynamic parameter adjustment engine according to the fusion processing result, combining a preset immersive experience evaluation model, and generating an audio parameter adjustment scheme in real time in an immersive audiovisual audio follow-up adjustment method.

[0020] Figure 5 It is a flowchart of reducing the high-frequency gain and increasing the low-frequency dynamic range at the same time, and simultaneously realizing pre-echo suppression processing when the audiovisual content is identified as a thrilling scene and the user physiological parameters show a nervous state in an immersive audiovisual audio follow-up adjustment method.

[0021] Figure 6 It is a flowchart of calling a virtual sound amplification algorithm to improve the diffusion coefficient of multi-channel signals and dynamically adjusting the coverage range of the sound field according to the distance change of the user's body posture in an immersive audiovisual audio follow-up adjustment method.

[0022] Figure 7 It is a flowchart of constructing a user experience feedback closed loop and generating personalized adjustment preference parameters by combining the user's visual attention area collected by an eye tracking device with the user's active triggered adjustment instructions in an immersive audiovisual audio follow-up adjustment method.

[0023] Figure 8 It is a main structure diagram of an immersive audiovisual audio follow-up adjustment system.

[0024] Figure 9 It is a structure block diagram of a multi-dimensional perception system module in an immersive audiovisual audio follow-up adjustment system. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0026] The specific implementation of the present invention is described in detail below with reference to specific embodiments.

[0027] The present invention provides an immersive video and audio follow-up adjustment method and system, which solves the technical problems in the background technology.

[0028] like Figure 1 FIG. 1 is a main flow chart of an immersive audio and video follow-up adjustment method provided by an embodiment of the present invention. The immersive audio and video follow-up adjustment method includes: Step S100: Build a multi-dimensional perception system to collect multi-source information in real time; The multi-source information includes the user's head six-degree-of-freedom motion data, body posture dynamic change data, physiological state parameters, and audio feature data and video frame feature data of the currently playing audio and video content; Step S200: fusion processing of the collected multi-source information based on a deep neural network model, mining the dynamic mapping relationship between user status and audio and video content through spatiotemporal correlation analysis, and identifying the user's interaction intention as well as the emotional tone and spatial scene attributes of the audio and video content; Step S300: Based on the fusion processing results, the dynamic parameter adjustment engine is called, and combined with the preset immersive experience evaluation model, an audio parameter adjustment plan is generated in real time; The audio parameters include three-dimensional spatial sound image positioning parameters, dynamic range compression ratio, multi-channel gain distribution coefficient and psychoacoustic optimization parameters, which are used to achieve deep collaborative adaptation of audio effects with user status and audio and video content, and build an immersive audio sound field with scene-adaptive characteristics; Step S400: Building a user experience feedback closed loop, using an eye tracking device to collect the user's visual focus area, and combining it with the adjustment instructions actively triggered by the user to generate personalized adjustment preference parameters; The application is applied, through the construction of multi-dimensional perception system, aims to capture the user state and audio-visual content characteristics, through the collection of head movement, body posture, physiological parameters and audio-visual features, to provide multi-dimensional data basis for subsequent adjustment, avoid the adjustment deviation caused by single information. Based on deep neural network fusion processing multi-source information, through the time and space correlation analysis to mine the dynamic relationship between user and content, break through the limitation of traditional independent analysis, accurately identify the user interactive intention and audio-visual attribute, provide decision basis for adjustment. Calling the dynamic parameter adjustment engine to generate scheme, the analysis result is converted into specific audio parameter adjustment, realizes the deep adaptation of audio and user state, content, solves the scene mismatch problem of traditional static adjustment. At the same time, the construction of feedback loop generates personalized parameters through capturing user visual attention and active instructions, makes up for the individual differences of general adjustment, so as to realize long-term optimization.

[0029] As shown in Figure 2 As a preferred embodiment of the application, the construction of multi-dimensional perception system, real-time collection of multi-source information specifically includes: Step S101: collecting six-degree-of-freedom motion data of the user's head through millimeter wave radar and inertial measurement unit; Step S102: collecting dynamic change data of user body posture and surface micro-motion characteristics by using distributed pressure sensor array and infrared thermal imaging technology; Step S103: collecting physiological state parameters of heart rate variability and respiration rate of the user by using wearable biosensor; Step S104: synchronously collecting real-time frequency spectrum waterfall diagram, sound pressure level dynamic curve of audio stream and depth map sequence and dynamic object motion vector field of video frame through audio-visual content analysis interface.

[0030] In the application of the embodiment, the six-degree-of-freedom motion data of the head is collected by using millimeter wave radar and inertial measurement unit (IMU) fusion, the purpose is to improve the accuracy and stability of head motion tracking through high-frequency sampling (such as more than 100Hz) and multi-sensor complementation, to provide accurate input for sound image positioning adjustment. The body posture and surface micro-motion are captured by using distributed pressure sensor array and infrared thermal imaging technology, which can make up for the limitation of pure visual recognition, accurately perceive the details such as user sitting posture change, and provide basis for sound field range adjustment. The heart rate variability and respiration rate are collected by using wearable biosensor, aiming to objectively reflect the user's emotional state (such as tension, relaxation) through physiological signals, and avoid the one-sidedness of relying only on behavior judgment. Through the audio-visual content analysis interface, the audio frequency spectrum, sound pressure level and video depth map are synchronously collected, the features are extracted from the content source, and the adjustment is matched with the style and scene of the audio-visual itself.

[0031] As shown in Figure 3As shown, as a preferred embodiment of the present invention, the method of fusing the collected multi-source information based on the deep neural network model, mining the dynamic mapping relationship between user status and audio-visual content through spatiotemporal correlation analysis, and identifying the user interaction intention and the emotional tone and spatial scene attributes of the audio-visual content specifically includes: Step S201: Using an attention mechanism-based converter model to perform spatiotemporal alignment on multi-source information; Step S202: Using the pre-trained emotion recognition sub-model, the physiological state parameters are correlated with the audio feature data and a real-time emotional state label of the user is output; Step S203: constructing a scene classifier to identify the spatial scene type and acoustic characteristics of the audio and video content based on the video frame feature data and the audio spatial features; In this embodiment, a transformer model employing an attention mechanism aligns multimodal data in space and time. By strengthening key relationships (such as head movement and video depth information), this approach addresses temporal and spatial asynchrony issues with multi-source data and improves fusion accuracy. A pre-trained emotion recognition sub-model correlates physiological parameters with audio features, converting abstract physiological signals into directly adjustable emotion labels, providing a basis for scenario-based adjustments (e.g., thriller scenes). A scene classifier is constructed to identify spatial scene types and acoustic characteristics, accurately distinguishing between indoor, outdoor, open, and enclosed scenes, providing scene attribute support for targeted adjustments in open areas and other scenarios.

[0032] like Figure 4 As shown, as a preferred embodiment of the present invention, the method of calling the dynamic parameter adjustment engine based on the fusion processing result and combining the preset immersive experience evaluation model to generate the audio parameter adjustment scheme in real time specifically includes: Step S301: When a significant head turn is detected, the dynamic parameter adjustment engine calculates the sound image localization compensation value in real time. By adjusting the multi-channel delay and gain differences, the azimuth and elevation angles of the spatial sound image are deflected synchronously with the user's viewing angle. Step S302: If the video content is identified as a thrilling scene and the user's physiological parameters indicate a state of tension, the high-frequency gain is reduced while the low-frequency dynamic range is increased, thereby simultaneously implementing pre-echo suppression processing; Step S303: For open-space scenarios, a virtual sound reinforcement algorithm is invoked to increase the diffusion coefficient of the multi-channel signal and dynamically adjust the coverage of the sound field according to the distance change of the user's body posture; It should be understood that the immersive experience evaluation model covers four aspects: perceived immersion (visual clarity, audio positioning, etc.), emotional response (emotions identified through physiological data and facial expressions), cognitive load (eye tracking, task completion reflecting cognitive consumption), and behavioral performance (interaction frequency, success rate, etc. reflecting user engagement); when a large head turn of the user is detected, sound image synchronization deflection is achieved by adjusting the multi-channel delay difference and gain difference, the purpose is to simulate the characteristics of the sound field in the real world that the sound image position changes when the viewing angle changes, to solve the problem of immersion break caused by traditional fixed sound image. When a horror scene is identified and the user is nervous, high and low frequency adjustment and pre-echo suppression are used to adapt to the user's physiological state and avoid high-frequency harshness or low-frequency deficiency that aggravates discomfort. For open sites, a virtual sound amplification algorithm is called to improve the diffusion coefficient and dynamically adjust the sound field range, simulate the acoustic characteristics of real open spaces, and solve the distortion problem of closed space adjustment scheme in open scenes.

[0033] As shown in Figure 5 As a preferred embodiment of the present application, when the video and audio content is identified as a horror scene and the user's physiological parameters show a nervous state, the high-frequency gain is reduced and the low-frequency dynamic range is increased, and the pre-echo suppression process is simultaneously implemented, which specifically includes: Step S3021: The adjustment engine receives the scene and user state label, and analyzes the related parameters of high-frequency gain, low-frequency dynamic range, and pre-echo suppression; Step S3022: Separate the high-frequency signal, linearly attenuate it with a digital attenuator according to the threshold, and control the signal peak value; Step S3023: Separate the low-frequency signal and calculate the dynamic range; Step S3024: Expand the dynamic range to the preset interval, increase the gain for weak signals, and retain the amplitude for strong signals; Step S3025: Based on the preset leading audio frame, identify potential pre-echo, and use an adaptive filter for selective attenuation; Step S3026: Collect the processed audio parameters and compare them with the standard audio parameters to correct excessive attenuation or distortion; Step S3027: If the user's nervous state is relieved, gradually return to the regular adjustment parameters corresponding to the horror scene; The engine analysis parameter is adjusted in the application of the embodiment, aiming to determine the target of high and low frequency and pre-echo suppression, and avoid blind adjustment; the high frequency signal is separated and linearly attenuated, which can reduce the stimulation of sharp sound effect on nervous users; the low frequency dynamic range is expanded, which can not only retain the impact of thrilling scenes, but also enhance details through weak signal gain; the pre-echo suppression can eliminate sound distortion and improve clarity; the feedback calibration is used to avoid the imbalance of hearing caused by excessive adjustment; the parameters are backed up after the user's nervous state is relieved, which can realize smooth transition when the state changes and avoid sudden discomfort. The specific standard for relieving the nervous state is: physiological signal recovery: the indexes such as heart rate variability and breathing rate collected by wearable biological sensors return to the baseline level of the user in the thrilling scene (such as the heart rate fluctuation amplitude is reduced to within ± 10% of the pre-scene range, and the breathing rate returns to a stable state); behavior characteristics are stable: the distributed pressure sensor and infrared thermal imaging technology detect that the user's body posture tends to be stable (such as no sudden limb tension, curling and other micro-motions), the head movement amplitude is reduced and the rhythm is smooth; emotional label conversion: the pre-trained emotion recognition sub-model combines physiological and behavioral data to output the nervous label to neutral emotional labels such as calm or focused, and the state lasts for a certain period of time (such as 3-5 seconds).

[0034] As shown in Figure 6 As a preferred embodiment of the present application, for open site scenes, a virtual sound amplification algorithm is called to improve the diffusion coefficient of multi-channel signals, and the coverage range of the sound field is dynamically adjusted according to the distance change of the user's body posture, which specifically includes: Step S3031: After recognizing the open site scene, the virtual sound amplification algorithm is activated, and the original diffusion coefficient of the multi-channel signal and the distance parameter of the user's body posture are obtained; Step S3032: Analyze the spatial characteristics of the multi-channel signal, determine the diffusion coefficient adjustment interval, and improve the diffusion coefficient by adding differential reverberation to monitor and ensure that the diffusion effect meets the acoustic characteristics of the open site; Step S3033: Quantize the distance parameter of the user's body posture to generate a comprehensive distance parameter, and establish a mapping relationship model between the comprehensive distance parameter and the coverage range of the sound field; Step S3034: Calculate the target sound field coverage radius according to the mapping relationship model, adjust the gain and delay parameters of each sound channel to match the target radius, and update the parameters according to the real-time movement state of the user to realize continuous adjustment of the sound field range; Step S3035: Compare the actual sound field data with the expected effect, and if the deviation is out of limit, correct the parameters until the scene and user state are adapted; In the application of the embodiment, first, the virtual sound amplification algorithm is activated after the open space scene is identified, and the original diffusion coefficient of the multi-channel signal and the distance parameter of the user's body posture are synchronously obtained, with the purpose of determining the starting point and core basis of adjustment. The original diffusion coefficient determines the basic sound diffusion capability, and the distance parameter of the user (such as the equivalent distance corresponding to the straight-line distance from the device and the body pitch angle) provides a key input for subsequent personalized adaptation, avoiding blind adjustment that deviates from the actual position of the user. Second, the spatial characteristics (such as the intensity and phase relationship of each channel) of the multi-channel signal are analyzed, and the diffusion coefficient adjustment interval is determined. The purpose of adding differential reverberation to improve the diffusion coefficient is to simulate the natural sound diffusion and gentle attenuation of the acoustic characteristics in the open space, solving the problem of lack of spaciousness caused by excessive sound focusing in the traditional closed space adjustment scheme. At the same time, the diffusion effect is monitored to ensure that it meets the scene characteristics, laying the foundation for subsequent sound field range adjustment. Third, the distance parameter of the user's body posture is quantified to generate a comprehensive distance parameter, and a mapping relationship model between the comprehensive distance parameter and the sound field coverage range is established. This step aims to convert the dynamic posture of the user (such as leaning forward to approach the device or leaning backward to move away) into a calculable sound field adjustment index, breaking through the limitations of the traditional fixed sound field range, and allowing the sound field to actively adapt to the user's position changes. Subsequently, the target sound field coverage radius is calculated according to the mapping model, the gain and delay parameters of each channel are adjusted to match the radius, and the parameters are updated in real time according to the user's movement, realizing continuous adjustment of the sound field range. The significance of this is to ensure that the user is always in the optimal coverage area of the sound field during movement, avoiding the problem of strong and weak sound or positional distortion caused by position changes. Finally, the actual sound field data is collected and compared with the expected effect. If the deviation exceeds the limit, the parameters are corrected until the scene and user state are adapted. This feedback mechanism is to compensate for the difference between theoretical calculation and actual listening, and to improve the adjustment accuracy through dynamic calibration.

[0035] As shown in Figure 7 As a preferred embodiment of the present application, the user experience feedback closed loop is constructed, the visual attention area of the user is collected by the eye tracking device, and the personalized adjustment preference parameter is generated by combining the user's active triggering adjustment instruction, which specifically includes: Step S401: Start the eye tracking device to capture the user's eye movement trajectory in real time, analyze the visual attention area, and collect the audio adjustment instruction actively triggered by the user through gestures, voice or physical buttons; Step S402: Extract the visual attention area coordinates and dwell time, perform semantic analysis or action recognition on the audio adjustment instruction actively adjusted by the user, and convert it into a standardized adjustment parameter requirement; Step S403: Establish an association model between feedback data and audio parameters, and analyze the matching relationship between the visual attention area and the audio characteristics; Step S404: count the frequency and amplitude of the user's active adjustment instructions, extract the user's adjustment tendency for different scenes, and form an initial preference feature set; Step S405: based on the initial preference feature set, combine the basic adjustment parameters of the current audio-visual scene, and generate personalized adjustment preference parameters through a weighting algorithm; Step S406: input the personalized adjustment preference parameters into a deep neural network model for incremental training, and update the model's understanding of the user's subjective needs; In application, the dynamic parameter adjustment engine preferentially calls the personalized preference parameters in the subsequent audio adjustment process, so that the output result adapts to the user's hearing preference, while continuously receiving new feedback data, realizing closed-loop continuous iteration, collecting eye movement trajectories and active instructions, aiming to directly obtain the user's focus of attention (such as the characters in the picture) and adjustment preference (such as adjusting the volume), making up for the shortcomings of passive perception; preprocessing data and converting into standardized requirements can eliminate noise and ambiguity of original data, facilitating subsequent analysis; establishing a feedback and audio parameter association model can mine visual attention (such as preferring clear voice when focusing on dialogue) and adjustment tendency (such as preferring strong low frequency in action scenes), forming initial preference features; a weighting algorithm generates personalized parameters to balance general adjustment and individual differences; inputting the personalized adjustment preference parameters into a deep neural network model for incremental training can make the adjustment adapt to the user's habits in the long term.

[0036] As shown in Figure 8 As another preferred embodiment of the present application, on the other hand, an immersive audio-visual audio follow-up adjustment system, the system comprises: A multi-dimensional perception system module 100 is used to build a multi-dimensional perception system and collect multi-source information in real time; An identification module 200 is used to fuse the collected multi-source information based on a deep neural network model, mine the dynamic mapping relationship between the user state and the audio-visual content through spatio-temporal correlation analysis, and identify the user's interactive intention and the emotional tone and spatial scene attribute of the audio-visual content; A first generation module 300 is used to call a dynamic parameter adjustment engine according to the fusion processing result, combine a preset immersive experience evaluation model, and generate an audio parameter adjustment scheme in real time; A second generation module 400 is used to build a user experience feedback closed loop, collect the user's visual attention area through an eye tracking device, combine the user's active triggered adjustment instruction, and generate personalized adjustment preference parameters.

[0037] In the application, the multi-dimensional perception system module 100 is used to construct a multi-dimensional perception system, real-time collect multi-source information, fuse the collected multi-source information based on a deep neural network model, mine the dynamic mapping relationship between the user state and the audio-visual content through spatio-temporal correlation analysis, and the recognition module 200 identifies the user interaction intention and the emotional tone and spatial scene attribute of the audio-visual content. According to the fusion processing result, the dynamic parameter adjustment engine is called, the first generation module 300 combines the preset immersive experience evaluation model to generate an audio parameter adjustment scheme in real time, and the second generation module 400 constructs a user experience feedback closed loop. The visual attention area of the user is collected through an eye tracking device, and the personalized adjustment preference parameter is generated in combination with the adjustment instruction triggered by the user.

[0038] As shown in Figure 9 , as another preferred embodiment of the application, the multi-dimensional perception system module 100 specifically includes: The first acquisition unit 101 is configured to collect six-degree-of-freedom motion data of the user's head through a millimeter wave radar and an inertial measurement unit. The second acquisition unit 102 is configured to collect dynamic change data of the user's body posture and body surface micro-motion features by using a distributed pressure sensor array and infrared thermal imaging technology. The third acquisition unit 103 is configured to collect physiological state parameters of heart rate variability and breathing frequency of the user by using a wearable biological sensor. The fourth acquisition unit 104 is configured to synchronously collect real-time spectrum waterfall diagram, sound pressure level dynamic curve of an audio stream, and depth map sequence and dynamic object motion vector field of a video frame through an audio-visual content analysis interface.

[0039] In the application, the first acquisition unit 101 collects six-degree-of-freedom motion data of the user's head through a millimeter wave radar and an inertial measurement unit, the second acquisition unit 102 collects dynamic change data of the user's body posture and body surface micro-motion features by using a distributed pressure sensor array and infrared thermal imaging technology, the third acquisition unit 103 collects physiological state parameters of heart rate variability and breathing frequency of the user by using a wearable biological sensor, and the fourth acquisition unit 104 synchronously collects real-time spectrum waterfall diagram, sound pressure level dynamic curve of an audio stream, and depth map sequence and dynamic object motion vector field of a video frame through an audio-visual content analysis interface.

[0040] The above-mentioned embodiments of the present application provide an immersive audiovisual audio follow-up adjustment method and an immersive audiovisual audio follow-up adjustment system, and a multi-dimensional perception system is constructed to comprehensively capture user states and audiovisual content characteristics, to provide a multi-dimensional data basis for subsequent adjustment by collecting head movement, body posture, physiological parameters and audiovisual characteristics, and to avoid adjustment deviation caused by single information. Based on deep neural network fusion processing of multi-source information, the dynamic relationship between the user and the content is mined through spatio-temporal correlation analysis, the limitations of traditional independent analysis are broken through, the user interaction intention and the audiovisual attribute are accurately identified, and a decision basis is provided for adjustment. The dynamic parameter adjustment engine is called to generate a scheme, the analysis result is converted into specific audio parameter adjustment, the depth adaptation of the audio and the user state and the content is realized, and the scene mismatch problem of the traditional static adjustment is solved. Meanwhile, the feedback closed loop is constructed to generate personalized parameters by capturing the user's visual attention and active instructions, to make up for the individual differences of the general adjustment, to realize long-term optimization, to accurately adapt the audio effect to the user state and the audiovisual content, and to significantly improve the naturalness and adaptability of the immersive experience, to balance the generality and individual differences, and to bring the user a more scene and self-demand matching audio experience.

[0041] In order to enable the above-mentioned method and system to run smoothly, the system can include more or less components than those described above, or combine certain components, or different components, such as input and output devices, network access devices, buses, processors and memories, etc.

[0042] The processor can be a central processing unit, and can also be other general-purpose processors, digital signal processors, application-specific integrated circuits, ready-to-program gate arrays or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The above-mentioned processor is the control center of the system, which is connected with various parts by various interfaces and lines.

[0043] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description simple, all possible combinations of the technical features in the above-mentioned embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0044] The above-mentioned embodiments only express several embodiments of the present application, and the description is more specific and detailed, but it cannot be understood as a limitation on the scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which belong to the protection scope of the present application. Therefore, the protection scope of the present application patent should be subject to the appended claims.

[0045] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An immersive audio-video-audio follow adjustment method, characterized in that, The method comprises: constructing a multi-dimensional perception system to collect multi-source information in real time; the multi-source information includes the user's head six-degree-of-freedom motion data, body posture dynamic change data, physiological state parameters, and current audio feature data and video frame feature data of the playing video content; based on a deep neural network model, the collected multi-source information is fused and processed, the dynamic mapping relationship between the user state and the video content is mined through space-time correlation analysis, the user interaction intention and the emotional tone and spatial scene attribute of the video content are identified; according to the fusion processing result, a dynamic parameter adjustment engine is called, and an immersive experience evaluation model is combined to generate an audio parameter adjustment scheme in real time; the audio parameters include three-dimensional space sound image positioning parameters, dynamic range compression ratio, multi-channel gain distribution coefficient and psychoacoustic optimization parameters, which are used to realize the deep collaborative adaptation of audio effect, user state and video content, and construct an immersive audio sound field with scene adaptive characteristics; a user experience feedback closed loop is constructed, the user's visual attention area is collected through an eye tracking device, and personalized adjustment preference parameters are generated combined with the user's active triggered adjustment instructions.

2. The immersive audio-video frequency following adjustment method of claim 1, wherein, The method comprises: collecting the user's head six-degree-of-freedom motion data through a millimeter wave radar and an inertial measurement unit; adopting a distributed pressure sensor array and infrared thermal imaging technology to collect the dynamic change data of the user's body posture and the micro-motion characteristics of the body surface; using wearable biosensors to collect the user's heart rate variability and respiratory rate physiological state parameters; through a video content analysis interface, the real-time frequency spectrum waterfall diagram, the sound pressure level dynamic curve of the audio stream, and the depth map sequence and dynamic object motion vector field of the video frame are synchronously collected.

3. The immersive audio-video frequency following adjustment method of claim 1, wherein, The method comprises: adopting a transformer model with an attention mechanism to perform space-time alignment on the multi-source information; through a pre-trained emotion recognition sub-model, the physiological state parameters and the audio feature data are analyzed to output the user's real-time emotional state label; constructing a scene classifier to identify the spatial scene type and acoustic characteristics of the video content based on the video frame feature data and the audio spatial feature.

4. The immersive audio-video frequency following adjustment method of claim 1, wherein, The method comprises: when detecting that the user's head performs a large turning action, the dynamic parameter adjustment engine calculates a sound image positioning compensation value in real time, adjusts the multi-channel delay difference and gain difference, and makes the azimuth and elevation of the spatial sound image deflect with the user's visual angle; if the video content is identified as a thrilling scene and the user's physiological parameters show a nervous state, the high-frequency gain is reduced, the low-frequency dynamic range is increased, and pre-echo suppression processing is simultaneously implemented. For open field scenes, a virtual sound amplification algorithm is called to improve the diffusion coefficient of multi-channel signals, and the coverage range of the sound field is dynamically adjusted according to the distance change of the user's body posture.

5. The immersive audio-video frequency following adjustment method of claim 4, wherein, If the video and audio content is identified as a thriller scene and the user's physiological parameters show a tense state, the high-frequency gain is reduced, the low-frequency dynamic range is increased, and the pre-echo suppression processing is simultaneously implemented, specifically including: The adjustment engine receives scene and user state labels, analyzes the related parameters of high-frequency gain, low-frequency dynamic range, and pre-echo suppression; Separate the high-frequency signal, linearly attenuate it with a digital attenuator according to the threshold, and control the signal peak value; Separate the low-frequency signal and calculate the dynamic range; Expand the dynamic range to the preset interval, increase the gain for weak signals, and preserve the amplitude for strong signals; Based on the preset leading audio frame, identify potential pre-echoes, and use an adaptive filter for selective attenuation; Collect the processed audio parameters and compare them with the standard audio parameters to correct excessive attenuation or distortion; If the user's tense state is relieved, gradually return to the regular adjustment parameters corresponding to the thriller scene.

6. The immersive audio-video frequency following adjustment method of claim 4, wherein, For open field scenes, a virtual sound amplification algorithm is called to improve the diffusion coefficient of multi-channel signals, and the coverage range of the sound field is dynamically adjusted according to the distance change of the user's body posture, specifically including: After identifying the open field scene, activate the virtual sound amplification algorithm, and simultaneously obtain the original diffusion coefficient of multi-channel signals and the distance parameter of the user's body posture; Analyze the spatial characteristics of multi-channel signals to determine the diffusion coefficient adjustment interval, and increase the diffusion coefficient by adding differential reverberation to monitor and ensure that the diffusion effect meets the acoustic characteristics of open field scenes; Quantify the distance parameter of the user's body posture to generate a comprehensive distance parameter, and establish a mapping relationship model between the comprehensive distance parameter and the sound field coverage range; Calculate the target sound field coverage radius based on the mapping relationship model, adjust the gain and delay parameters of each channel to match the target radius, and update the parameters based on the user's real-time movement state to achieve continuous adjustment of the sound field range; Compare the actual sound field data with the expected effect, and if the deviation exceeds the limit, correct the parameters until they adapt to the scene and user state.

7. The immersive audio-video frequency following adjustment method of claim 1, wherein, The user experience feedback closed loop is constructed by collecting the user's visual attention area through eye tracking equipment, and combining the user's active triggering adjustment instructions to generate personalized adjustment preference parameters, specifically including: Start the eye tracking equipment to capture the user's eye movement trajectory in real time, analyze the visual attention area, and simultaneously collect the user's active triggering audio adjustment instructions through gestures, voice, or physical buttons; Extract the coordinates and duration of the visual attention area, and perform semantic analysis or action recognition on the user's active audio adjustment instructions to convert them into standardized adjustment parameter requirements; Establish an association model between feedback data and audio parameters to analyze the matching relationship between the visual attention area and audio features; Statistically analyze the frequency and amplitude of the user's active adjustment instructions to extract the user's adjustment inclination for different scenes, forming an initial preference feature set; Based on the initial preference feature set, combine the current video and audio scene's basic adjustment parameters to generate personalized adjustment preference parameters through a weighted algorithm; Input the personalized adjustment preference parameters into the deep neural network model for incremental training to update the model's understanding of the user's subjective needs.

8. An immersive audiovisual audio follow-up adjustment system, characterized in that, The system comprises: A multi-dimensional perception system module for constructing a multi-dimensional perception system and collecting multi-source information in real time; An identification module for fusing the collected multi-source information based on a deep neural network model, mining the dynamic mapping relationship between user states and audio-visual content through spatio-temporal correlation analysis, and identifying the user interaction intention, the emotional tone of the audio-visual content, and the spatial scene attribute; A first generation module for generating an audio parameter adjustment scheme in real time according to the fusion processing result, calling a dynamic parameter adjustment engine, and combining a preset immersive experience evaluation model; A second generation module for constructing a user experience feedback closed loop, collecting the user's visual focus area through an eye tracking device, combining the user's active triggered adjustment instruction, and generating personalized adjustment preference parameters.

9. The immersive audio-visual audio follow adjustment system of claim 8, wherein, The multi-dimensional perception system module specifically comprises: A first collection unit for collecting six-degree-of-freedom motion data of the user's head through a millimeter wave radar and an inertial measurement unit; A second collection unit for collecting dynamic change data of the user's body posture and body surface micro-motion characteristics using a distributed pressure sensor array and infrared thermal imaging technology; A third collection unit for collecting physiological state parameters of the user's heart rate variability and breathing frequency using wearable biological sensors; A fourth collection unit for synchronously collecting real-time frequency spectrum waterfall diagrams, sound pressure level dynamic curves of audio streams, and depth map sequences and dynamic object motion vector fields of video frames through an audio-visual content analysis interface.

Citation Information

Cited By

  • Ear clip type earphone adaptive control method for real-time music style identification

    CN121442237A

  • Method and system for efficiently carrying out audio and video collaborative management

    CN121792774A

  • Immersive cinema dynamic content adaptation control method based on multi-modal fusion

    CN121956638A

  • Play control method, device and equipment of 5D dynamic cinema and storage medium

    CN122172531A