Large model driven cockpit multi-occupant individual sound field management method and device and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
AI Technical Summary
In multi-occupant cabin environments, the continuous changes in occupant posture lead to spatial mismatch in independent sound fields and enhanced crosstalk. Existing sound field management methods are unable to respond in a timely manner, affecting privacy and auditory consistency.
By adopting a large model-driven cockpit multi-occupant independent sound field management method, occupant status data is collected and preprocessed in real time, a continuous and reliable attitude representation is constructed, a multi-occupant attitude evolution prediction model is established, a spatial description of the listening area and acoustic prior constraints are generated, an independent sound field allocation strategy is generated in combination with audio object type, and adaptive monitoring and adjustment are performed during the sound field execution process.
It achieves feedforward adjustment to changes in occupant posture, reduces spatial mismatch in the listening area, improves the stability and consistency of the single sound field, and enhances voice privacy protection and listening quality.
Smart Images

Figure CN122435946A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound field management technology, specifically to a method, device, and storage medium for managing the sound field of a large-scale model-driven cockpit with multiple occupants. Background Technology
[0002] With the continuous development of intelligent cockpits and in-vehicle infotainment systems, in-vehicle audio application scenarios are becoming increasingly diverse. Audio content such as navigation prompts, voice interaction, communication calls, and multimedia entertainment in the cockpit exhibits characteristics of concurrency, multi-source, and personalization. In multi-occupant environments, different occupants have significant differences in their attention areas, listening positions, and interaction needs regarding audio content. The cockpit sound field is gradually evolving from traditional whole-cabin playback towards zoned, directional, and personalized sound.
[0003] For example, the invention patent with publication number CN118042398B discloses a method for optimizing the sound field in a vehicle, an audio system, electronic devices, and a storage medium, relating to the field of sound reproduction technology. This application first calculates the actual acoustic response of each seat in the car's cabin to preset audio signals played by each speaker, and then calculates the ideal acoustic response of the optimal listening position in the listening room to preset audio signals played by each speaker. Based on the match between the actual and ideal acoustic responses, the sound field at each seat in the cabin is reconstructed, resulting in a reconstructed sound field for each seat. This allows the subsequent audio system to control each speaker to reproduce the audio signal based on the reconstructed sound field results at each seat. At this point, the actual acoustic response at each seat will match the ideal acoustic response in the listening room, meaning that each seat can balance the sound field size with the acoustic energy / sound field balance between the passenger's ears, thus ensuring the acoustic experience for passengers at each seat in the cabin.
[0004] For example, invention patent CN117041857A discloses a method, system, vehicle, and storage medium for adjusting the position and volume of the cabin sound field. The method includes the following steps: collecting noise from different directions outside the vehicle at preset intervals; obtaining the equalization parameter range and volume range of the cabin sound field; calculating the sound field equalization parameter value based on the noise from different directions outside the vehicle and the equalization parameter range of the cabin sound field, and adjusting the audio signal output by the cabin speaker based on the sound field equalization parameter value, thus achieving sound field position adjustment; calculating the volume value of the cabin speaker based on the noise from different directions outside the vehicle and the volume range of the cabin sound field, and adjusting the volume output by the cabin speaker based on the volume value, thus achieving volume adjustment. This invention can automatically adjust the position and volume of the cabin sound field according to the magnitude and location of external noise, allowing the driver and passengers to enjoy a better audio experience.
[0005] However, during actual driving, occupants' head and body postures are constantly changing, including turning their heads, leaning forward, leaning to the side, and turning their heads to talk. This causes spatial mismatch in the independent sound fields constructed based on the initial calibration position or static acoustic model. Existing sound field management methods mostly use passive compensation, correcting only after sound leakage is detected. This makes it difficult to respond promptly to continuous posture changes, leading to drift of independent sound field boundaries, enhanced crosstalk, and affecting privacy and auditory consistency. The problem is particularly prominent when multiple occupants are active simultaneously.
[0006] Therefore, in order to address the above issues, there is an urgent need for large-scale model-driven cockpit multi-occupant single-sound field management methods, equipment, and storage media. Summary of the Invention
[0007] Technical problems to be solved
[0008] To address the shortcomings of existing technologies, this invention provides a method, device, and storage medium for managing the independent sound field of a multi-occupant cockpit driven by a large model, which solves the problems of spatial mismatch of independent sound fields, enhanced crosstalk, and delayed passive compensation caused by continuous changes in occupant posture in multi-occupant cockpits.
[0009] Technical solution
[0010] To achieve the above objectives, the present invention provides the following technical solution: a large-model driven cockpit multi-occupant single-sound field management method, comprising the following steps: S1, real-time acquisition of occupant status data in the cockpit, preprocessing of the occupant status data, evaluation of occupant head posture changes based on the preprocessed occupant status data, construction of a continuous and reliable posture representation, and judgment and smoothing processing to form a reliable state feature set; S2, based on the reliable state feature set, construction of a multi-occupant posture evolution prediction model, prediction of the evolution trend of occupant posture over time, and generation of audio matching the occupant posture changes based on the prediction results. S3, Based on the spatial description of the occupant's listening area and acoustic prior constraints, and combined with the currently activated audio object type, establish the association between the audio object and the occupant, generate a multi-occupant independent sound field allocation strategy, and perform adjustment and smoothing processing according to priority when sound field conflicts are detected; S4, Based on the independent sound field allocation strategy and acoustic prior constraints, perform 3D audio rendering with zoning constraints on each audio object, and monitor and evaluate the sound energy leakage status of the target sound area and non-target sound area during the sound field execution process, and trigger adaptive suppression and parameter update when sound field interference trends are detected.
[0011] Furthermore, the real-time acquisition of occupant status data within the cabin and the specific process of preprocessing this data are as follows: Real-time acquisition of occupant status data includes: acquiring the three-dimensional spatial position and head posture angle of each occupant's head through cabin vision, and performing head key point detection by the cabin vision perception unit; generating a corresponding key point response heatmap for each head key point, and using the maximum response value in the key point response heatmap as the detection confidence level of the head key point; acquiring pressure distribution data for each seat through seat occupancy sensors, and performing a pressure-weighted average calculation on the pressure distribution data for each seat to obtain the seat occupancy data. The system identifies the center of pressure and calculates the probability of occupancy based on the deviation between the current pressure distribution data and the empty seat pressure distribution. It also adds timestamps to the occupant status data using a unified time reference and performs resampling on data from different sampling frequencies according to a unified sampling period. Median filtering is used to perform preliminary smoothing and anomaly suppression on the head posture angle sequence. Numerical range constraints and anomaly suppression are applied to the detection confidence and occupancy probability, and minimum-maximum normalization is performed on the occupant status data. Finally, a single-sound field management database is established, and the original and pre-processed occupant status data are written into this database.
[0012] Furthermore, based on the preprocessed occupant state data, the changes in occupant head posture are evaluated, a continuous and reliable posture representation is constructed, and judgment and smoothing processes are performed to form a reliable state feature set. The specific process is as follows: Based on the preprocessed head posture angle, the second-order time difference of the head posture angle is calculated to obtain the posture angular acceleration; based on the sliding time window, the median of the posture angular acceleration is taken, and the median absolute deviation of the posture angular acceleration is calculated; the median of the posture angular acceleration is subtracted from the current posture angular acceleration, and divided by the sum of the median absolute deviation of the posture angular acceleration and the smallest positive number to obtain the normalized value of the posture angular acceleration, and the normalized value of the posture angular acceleration is compared with zero to perform non-negativity processing; the non-negativity posture angular acceleration normalization is then performed. The negative of the value is used as the exponent for natural exponentiation to obtain the attitude change suppression value. The detection confidence, seat occupancy probability, and attitude change suppression value are multiplied to obtain the attitude continuity confidence value. The attitude continuity confidence value of each crew member is calculated, and validity is determined: when the attitude continuity confidence value is greater than or equal to the confidence threshold, the current head 3D spatial position and head attitude angle are determined to be valid attitude data; when the attitude continuity confidence value is less than the confidence threshold, the current head 3D spatial position and head attitude angle are determined to be insufficiently confident, and the head 3D spatial position and head attitude angle are smoothed based on median filtering and moving average filtering using a sliding time window. The attitude continuity confidence value and the smoothed data are written into the single-sound field management database.
[0013] Furthermore, based on the reliable state feature set, a multi-occupant attitude evolution prediction model is constructed. The specific process for predicting the evolution trend of occupant attitude over time is as follows: For each occupant, a multi-occupant attitude temporal feature set is constructed based on the preprocessed head 3D spatial position, head attitude angle, and attitude continuity reliability value recorded in the monophonic field management database; the historical multi-occupant attitude temporal feature set is input into the multimodal temporal modeling algorithm according to the occupant index for temporal modeling, learning the dynamic evolution law of head 3D spatial position and head attitude angle changing over time, and constructing a multi-occupant attitude evolution prediction model; the attitude evolution prediction model infers the attitude changes of each occupant within the time window and outputs the predicted head 3D spatial position and predicted head attitude angle of each occupant in the future.
[0014] Furthermore, the specific process of generating a spatial description of the hearing area and acoustic prior constraints that match the occupant's posture changes based on the prediction results is as follows: For each occupant, acoustic model prior parameters are generated based on the predicted three-dimensional spatial position of the head and the predicted head posture angle; a head coordinate system is constructed based on the head posture angle, and the structural vectors of both ears relative to the head center are rotated and superimposed with the head center position to obtain the predicted positions of both ears, with the midpoint of the predicted positions of both ears as the center position of the target hearing area. Simultaneously, based on the spatial dispersion of the center position of the target hearing area within the prediction time window, the spatial coverage of the target hearing area is determined; within the spatial coverage of the target hearing area, acoustic prior parameters are generated around the target hearing area. Candidate sampling points are generated at the center of the listening area. For each candidate sampling point, the corresponding incident direction parameters are calculated based on the spatial orientation relationship between the candidate sampling point and the speaker position. The corresponding HRTF is obtained from the HRTF database based on the incident direction. The HRTF corresponding to the candidate sampling point is fused based on the spatial distance from the candidate sampling point to the center of the target listening area to obtain the target listening area HRTF prior. The equivalent propagation path parameters from the speaker to the target listening area are calculated based on the speaker position and the center of the target listening area. The predicted three-dimensional spatial position of each occupant's head, the predicted head attitude angle, and the corresponding acoustic model prior parameters are written into the independent sound field management database.
[0015] Furthermore, based on the spatial description of the occupants' listening areas and acoustic prior constraints, and combined with the currently activated audio object type, an association relationship between audio objects and occupants is established, generating a multi-occupant independent sound field allocation strategy. The specific process of performing adjustment and smoothing processing according to priority when sound field conflicts are detected is as follows: Obtain the target listening area center position, target listening area spatial coverage, and corresponding acoustic model prior parameters for each occupant at the current moment; simultaneously obtain the currently activated audio object information and identify the type of each audio object; based on the audio object type and the occupant's sitting posture pressure center position, target listening area center position, and spatial coverage, a multi-occupant independent sound field allocation strategy is established. The system establishes spatial and functional relationships between audio objects and occupants. Based on these relationships, it determines the priority order of each audio object within the occupant's individual sound field. Combining the target listening zone center position, spatial coverage, and prior acoustic model parameters for each occupant, it generates an independent sound field allocation strategy for each audio object. After generating the independent sound field allocation strategy, it performs a consistency check to identify sound field overlap. When sound field overlap is detected, it adjusts the corresponding independent sound field allocation strategy according to the priority order of the audio objects. Simultaneously, it performs time smoothing processing on the changes in the independent sound field allocation strategy over time and outputs the results.
[0016] Furthermore, based on the independent sound field allocation strategy and acoustic prior constraints, the specific process of performing 3D audio rendering with zoning constraints on each audio object is as follows: The center position, spatial coverage, and acoustic model prior parameters of the target listening area corresponding to each occupant are read, and the independent sound field allocation strategy and its spatial and functional relationship with the occupant are read for each audio object; the independent sound field allocation strategy and acoustic model prior parameters are aligned according to the occupant index and the audio object index to construct a 3D audio rendering input set, limiting the rendering constraints of each audio object within the corresponding occupant's target listening area; based on the 3D audio rendering input set, 3D audio rendering processing is performed on each audio object in the vehicle-mounted DSP, including: performing binaural rendering on the audio object based on the target listening area HRTF prior, and performing propagation path compensation on the rendering signal based on the equivalent propagation path parameters; simultaneously, zoning rendering is performed on the target listening areas corresponding to different occupants according to the independent sound field allocation strategy, so that each audio object is mapped to the independent sound field output of the corresponding occupant.
[0017] Furthermore, during the sound field execution process, the acoustic energy leakage status of the target sound area and non-target sound area is monitored and evaluated. The specific process of triggering adaptive suppression and parameter updates when a sound field interference trend is detected is as follows: After completing the 3D audio rendering output, for each occupant, the corresponding target sound area audio signal and non-target sound area audio signal are collected through the in-vehicle microphone array. Power spectral density estimation is performed on the collected audio signals, and noise suppression processing is applied to the power spectral density to obtain the denoised target sound area spectrum and the denoised non-target sound area spectrum. Energy integration is performed on the denoised target sound area spectrum and the denoised non-target sound area spectrum within the frequency band to obtain the target sound area energy and the non-target sound area energy. The non-target sound area energy is divided by the target sound area energy, and the logarithm of the ratio is taken to obtain the logarithmic energy leakage ratio. Energy normalization processing is performed on the denoised target sound area spectrum and the denoised non-target sound area spectrum within the frequency band to obtain the corresponding normalized energy leakage ratio. The spectral distribution is calculated based on the Jensen-Shannon divergence between the target sound zone spectrum and the non-target sound zone spectrum using a normalized spectral distribution, yielding the spectral distribution difference value. This difference value is then added to a minimum positive number, and the square root is taken to obtain the normalized denominator. The logarithmic energy leakage ratio is divided by the normalized denominator, and the opposite is used as the exponent for natural exponentiation. The result of natural exponentiation is added to a constant and the reciprocal is taken to obtain the zone crosstalk leakage drive value. The zone crosstalk leakage drive value for each occupant is calculated and compared with an allowable threshold. When the zone crosstalk leakage drive value exceeds the allowable threshold, a closed-loop suppression control for the individual sound field is triggered. This includes dynamically updating the individual sound field allocation strategy and adaptively adjusting the corresponding acoustic model prior parameters. Simultaneously, time smoothing is performed on the changes in the individual sound field allocation strategy and acoustic model prior parameters over time. The updated independent sound field allocation strategy and acoustic model prior parameters are then written back to the individual sound field management database.
[0018] The second aspect of this invention provides a large-model driven cockpit multi-occupant single-sound-field management device, comprising: an occupant state acquisition and preprocessing module, used to acquire occupant state data in the cockpit in real time, preprocess the occupant state data, evaluate changes in occupant head posture based on the preprocessed occupant state data, construct a continuous and reliable posture representation, and perform judgment and smoothing processing to form a reliable state feature set; and an attitude prediction and sound field adaptation module, used to construct a multi-occupant posture evolution prediction model based on the reliable state feature set, predict the evolution trend of occupant posture over time, and combine the prediction results to generate a listening area spatial description and acoustic prior that match the changes in occupant posture. Constraints; Multi-occupant independent sound field strategy and arbitration module, used to establish the association between audio objects and occupants based on the spatial description of the occupant's listening area and acoustic prior constraints, combined with the currently active audio object type, generate multi-occupant independent sound field allocation strategy, and perform adjustment and smoothing processing according to priority when sound field conflicts are detected; 3D audio rendering and closed-loop control module, used to perform 3D audio rendering with zonal constraints on each audio object based on independent sound field allocation strategy and acoustic prior constraints, and monitor and evaluate the acoustic energy leakage status of target sound area and non-target sound area during sound field execution, and trigger adaptive suppression and parameter update when sound field interference trend is detected.
[0019] A third aspect of the present invention provides a storage medium for managing the multi-occupant single sound field in a large model-driven cockpit, comprising: the storage medium having one or more programs, the one or more programs being executed by one or more processors to implement a method for managing the multi-occupant single sound field in a large model-driven cockpit.
[0020] Beneficial effects
[0021] The present invention has the following beneficial effects:
[0022] (1) This invention, by modeling and predicting the continuous changes in the occupant's posture, introduces posture evolution information before the generation and execution of the sound field, realizing the transformation from "post-compensation" to "pre-matching", effectively reducing the spatial mismatch of the listening area caused by the dynamic behavior of the occupant turning his head and leaning forward, and improving the stability and consistency of the sound field.
[0023] (2) The present invention comprehensively evaluates the continuous reliability of occupant posture data by integrating posture change features, visual detection confidence and seat occupancy probability, and performs smoothing processing on low confidence states to avoid abnormal posture or transient noise directly affecting sound field decision-making, thereby improving the reliable operation capability in complex cabin environments.
[0024] (3) This invention establishes spatial and functional associations based on the spatial description of the occupant's listening area and the type of audio object, and introduces a conflict arbitration and priority adjustment mechanism, which can achieve reasonable independent sound field allocation in scenarios with multiple occupants and multiple sound sources, and enhance the adaptability to complex application scenarios.
[0025] (4) This invention continuously monitors and evaluates the sound energy leakage status of the target sound zone and non-target sound zone, and automatically triggers parameter adaptive update when a sound field interference trend is detected, thereby realizing closed-loop control in the sound field execution process, effectively suppressing crosstalk enhancement problems, and improving the voice privacy protection capability and listening quality in the cockpit.
[0026] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0027] Figure 1 Flowchart of a large-model-driven multi-occupant single-sound-field management method for cockpits;
[0028] Figure 2 Structural diagram of a multi-occupant independent sound field management system for a large-scale model-driven cockpit;
[0029] Figure 3 A comparison chart of the distribution of crosstalk leakage driver values in different zones;
[0030] Figure 4 A schematic diagram of crosstalk leakage monitoring and closed-loop suppression in a single-sound field;
[0031] Figure 5 This is a schematic diagram of a multi-occupant attitude evolution prediction model and its feedforward acoustic field drive. Detailed Implementation
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. As those skilled in the art will understand, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Please see Figures 1-5 This invention provides a technical solution: a method for managing the independent sound field of a multi-occupant cockpit driven by a large model, such as... Figure 1As shown, the process includes the following steps: S1, real-time acquisition of occupant status data in the cabin, preprocessing of the occupant status data, evaluation of occupant head posture changes based on the preprocessed occupant status data, construction of a continuous and reliable posture representation, and judgment and smoothing processing to form a reliable state feature set; S2, based on the reliable state feature set, construction of a multi-occupant posture evolution prediction model, prediction of the evolution trend of occupant posture over time, and generation of a listening zone spatial description and acoustic prior constraints that match the occupant posture changes based on the prediction results; S3, based on the occupant's listening zone spatial description and acoustic prior constraints, combined with the currently activated audio object type, establishing the association between audio objects and occupants, generating a multi-occupant independent sound field allocation strategy, and performing adjustment and smoothing processing according to priority when sound field conflicts are detected; S4, based on the independent sound field allocation strategy and acoustic prior constraints, performing 3D audio rendering with partition constraints on each audio object, and monitoring and evaluating the sound energy leakage status of the target sound zone and non-target sound zone during the sound field execution process, and triggering adaptive suppression and parameter updates when a sound field interference trend is detected.
[0034] Specifically, the real-time acquisition of occupant status data within the cabin and the preprocessing of this data are as follows: Real-time acquisition of occupant status data includes: obtaining the three-dimensional spatial position and head attitude angles of each occupant's head through cabin vision. The cabin vision system consists of camera modules positioned on the dashboard, roof, or A-pillar area. Based on a calibrated cabin coordinate system, it outputs the three-dimensional spatial position parameters of the head in the cabin coordinate system and the corresponding attitude angle parameters, including at least pitch, yaw, and roll angles. The cabin vision perception unit then performs head key point detection, generating a corresponding key point response heatmap for each key point. Key points include, but are not limited to, the positions of the eyes, nose tip, mouth, and ears. The key point response heatmap is composed of pixel-level response distributions output by a neural network model, and the maximum response value in the key point response heatmap is used as the detection confidence level of the head key point to characterize the reliability of the current key point detection results. The data is then processed by the seat occupancy... Seat sensors collect pressure distribution data for each seat. Seat occupancy sensors are pressure array sensors installed in the seat cushion area, outputting pressure values corresponding to multiple sampling points. Pressure weighted average calculation is performed on the pressure distribution data of each seat to obtain the center position of the sitting posture pressure. The center position of the sitting posture pressure is obtained by multiplying the spatial coordinates of each pressure sampling unit with the corresponding pressure value and summing them after normalization. It is used to characterize the actual center position of the occupant's sitting posture on the seat and serves as auxiliary constraint information for determining the occupant's spatial position. The occupancy probability is calculated based on the deviation between the current pressure distribution data and the empty seat pressure distribution. The occupancy probability is calculated by the ratio of the total pressure value of the current pressure distribution to the calibrated empty seat pressure distribution. It is used to determine whether the current seat is occupied. The empty seat pressure distribution is a reference pressure distribution pre-collected and stored in the state of no occupants, used to distinguish between the actual seat occupancy state and the instantaneous disturbance or noise state. Based on a unified time reference, timestamps are added to the occupant status data. The unified time reference is provided by the vehicle's main control unit, and resampling processing is performed on data with different sampling frequencies according to a unified sampling period. The resampling processing is implemented through interpolation to ensure the alignment of occupant status data from different sources on the time axis. Median filtering is used to perform preliminary smoothing and anomaly suppression processing on the head posture angle sequence. The median filtering is based on a sliding time window to process the head posture angle sequence and is used to eliminate abnormal jump values caused by instantaneous occlusion, detection jitter, or changes in illumination. The length of the sliding time window is preferably 3 to 9 frames, corresponding to a time range of approximately 0.05 seconds to 0.30 seconds. The specific window length is configured according to the cockpit visual sampling frequency. For example, when the sampling frequency is 30Hz, a median filtering window length of 5 to 7 frames is preferred to ensure the anomaly suppression effect while avoiding excessive smoothing of normal head posture changes.Numerical range constraints and anomaly suppression are applied to the detection confidence and seat occupancy probability. The numerical range constraints limit the detection confidence and seat occupancy probability to a valid range, and truncation is performed when the values exceed the valid range to avoid outliers interfering with subsequent attitude assessment. Min-maximum normalization is applied to the occupant status data. The normalization process is based on the minimum and maximum values of the corresponding features obtained from historical operation phases, mapping data of different dimensions to a unified numerical range, which facilitates subsequent feature fusion and model processing. A monophonic field management database is established, in which raw and preprocessed occupant status data are written. This database centrally stores the raw data, preprocessed data, and associated index information of each occupant at different times, providing a unified data support foundation for subsequent occupant attitude assessment, attitude evolution prediction, and monophonic field control.
[0035] In this implementation scheme, through multi-source sensor fusion and data preprocessing under a unified time reference, stable, continuous, and quantifiable acquisition of occupant head spatial position, posture changes, and seat occupancy status is achieved, effectively suppressing data unreliability issues caused by detection jitter, momentary occlusion, and noise disturbance. At the same time, through standardized data alignment, smoothing, and normalization, a clear and time-consistent occupant status data foundation is constructed, providing highly reliable and feasible data support for subsequent attitude reliability assessment, attitude evolution prediction, and precise control of the independent sound field, significantly improving the stability and adaptability of cockpit sound field management in dynamic driving scenarios.
[0036] Specifically, based on the preprocessed occupant state data, the occupant's head posture changes are evaluated, a continuous and reliable posture representation is constructed, and judgment and smoothing processes are performed to form a reliable state feature set. The specific process is as follows: Based on the preprocessed head posture angle, the second-order time difference of the head posture angle is calculated to obtain the posture angular acceleration. The second-order time difference is calculated based on the head posture angle sequence of adjacent sampling times and is used to characterize the acceleration degree of head posture changes, thereby reflecting whether the occupant has behavioral characteristics such as rapid head turning or sudden posture changes. Based on a sliding time window, the median of the posture angular acceleration is taken, and the median absolute deviation of the posture angular acceleration is calculated. The sliding time window is a time interval covering multiple consecutive sampling times, preferably 5 to 40 frames, corresponding to 0.3 seconds to 2 seconds of sampling data, used to characterize the typical change level of the posture angular acceleration within a local time range. The median is used to characterize the steady-state change level within the time window, and the median absolute deviation is used to characterize the dispersion and fluctuation amplitude of the posture angular acceleration.The normalized value of the attitude angular acceleration is obtained by subtracting the median of the attitude angular acceleration from the current attitude angular acceleration and dividing by the sum of the median absolute deviation and the minimum positive value. The minimum positive value is introduced to avoid numerical instability when the median absolute deviation is close to zero. The normalized value is then compared to zero to perform non-negativity processing. If the normalized value is less than zero, it is set to zero to retain only attitude abrupt changes significantly above the steady-state level, suppressing the impact of low-amplitude normal attitude adjustments on subsequent judgments. The opposite of the non-negative attitude angular acceleration normalized value is then taken. The number is used as an exponent for natural exponential operation to obtain the attitude change suppression value. The natural exponential operation is used to map the amplitude of attitude angle change to a continuous and smooth suppression factor, so that the more drastic the attitude change, the stronger the corresponding suppression effect, thereby reducing the weight of attitude data in the reliability assessment. The result of the exponential operation is limited to the interval [0,1]. The detection confidence, the occupancy probability, and the attitude change suppression value are multiplied to obtain the attitude continuous reliability value. Among them, the detection confidence is used to reflect the visual reliability of the current head attitude detection result, the occupancy probability is used to reflect the physical reliability of the current occupant's presence, and the attitude change suppression value is used to reflect the visual reliability of the current head attitude detection result, the occupancy probability is used to reflect the physical reliability of the current occupant's presence, and the attitude change suppression value is used to reflect the physical reliability of the current occupant's presence. To reflect the stability of attitude changes in temporal continuity, a reliable representation of attitude continuity is constructed by multiplying these three factors (space, physical, and temporal) to comprehensively measure the reliability of the current attitude data in the spatial, physical, and temporal dimensions. The reliability value of each crew member's attitude continuity is calculated, and a validity determination is performed: when the reliability value is greater than or equal to a reliability threshold, the current 3D spatial position of the head and the head posture angle are considered valid attitude data, and this data is directly used for subsequent attitude evolution modeling and auditory region generation; when the reliability value is less than the reliability threshold, the current 3D spatial position of the head and the head posture angle are deemed insufficiently reliable, and further determination is based on... Median filtering and moving average filtering within a sliding time window smooth the head's three-dimensional spatial position and head posture angle. Median filtering suppresses aberrations caused by transient occlusion or detection failure, while moving average filtering maintains the overall trend continuity of posture changes, generating smooth posture estimation results for subsequent processing. Continuous and reliable posture values and smoothed data are written into the monophonic field management database and stored in association with occupant and time indices, providing a continuous and stable data input foundation for subsequent occupant posture evolution prediction, dynamic adjustment of the listening area, and monophonic field closed-loop control.
[0037] The specific formula for the attitude continuity confidence value is as follows:
[0038] ;
[0039] In the formula, The confidence value represents the continuous attitude, which is used to comprehensively evaluate the stability and reliability of occupant head posture data over time. By fusing visual detection reliability, the validity of actual occupant seating, and the continuity constraint of attitude changes, a higher confidence value is output when the attitude change is stable, the detection is reliable, and the occupant is real. When there are sudden attitude changes, abnormal shaking, or non-real seating, the confidence value is significantly reduced through an exponential suppression mechanism, thereby providing robust and filterable state inputs for subsequent attitude prediction, auditory zone modeling, and sound field control. This indicates the detection confidence level, which reflects the reliability of the cockpit vision system's identification of key points on the occupant's head, and avoids visual false detections or occlusions interfering with attitude judgment. It represents the probability of seat occupancy, used to characterize whether the corresponding seat is actually occupied by a passenger, and to suppress invalid posture data generated under conditions of empty seats, accidental triggering, or non-human interference. It represents the angular acceleration of the head, which is used to describe the instantaneous drasticness of changes in head posture and is the core dynamic feature for judging whether abrupt changes or abnormal shaking of the posture have occurred. This represents the median of the attitude angular acceleration, which is used as a steady-state reference level for attitude changes within the current time window to reduce the impact of extreme values on the evaluation results. It represents the median absolute deviation of attitude angular acceleration, used to measure the normal fluctuation range of attitude change amplitude, and to provide an adaptive scaling benchmark for abnormal changes. This represents a very small positive number, used to prevent the denominator from being zero and to enhance the stability of numerical calculations. It is not involved in attitude semantics judgment, and its preferred value is [value missing]. .
[0040] In this implementation scheme, by combining temporal analysis of occupant head posture changes with multi-source credible constraint fusion, continuous and reliable evaluation and adaptive smoothing of dynamic posture data are achieved. This effectively suppresses the impact of instantaneous jitter, occlusion, and abnormal mutations on the stability of posture recognition. While ensuring timely response to real posture changes, it improves the continuity and reliability of occupant state representation, providing a stable and feasible data foundation for subsequent posture evolution prediction and precise control of the single sound field. This significantly enhances the adaptability and robustness of cockpit sound field management in dynamic driving scenarios.
[0041] Specifically, based on the reliable state feature set, a multi-occupant posture evolution prediction model is constructed. The specific process for predicting the evolution trend of occupant posture over time is as follows: For each occupant, based on the preprocessed head three-dimensional spatial position, head posture angle, and posture continuous reliability value recorded in the monophonic field management database, a multi-occupant posture temporal feature set is constructed. The posture temporal feature set is organized into a multi-dimensional input tensor arranged in time order during the model input stage. The feature dimension includes at least: head three-dimensional spatial position vector, head posture angle vector, and corresponding posture continuous reliability value. Each feature is spliced in the time dimension to form a unified temporal input structure, which is used to simultaneously characterize the spatial state, change trend, and data reliability information of the occupant posture. The historical multi-occupant posture time-series feature set is input into a multimodal temporal modeling algorithm according to the occupant index for temporal modeling. A multimodal temporal modeling algorithm based on long short-term memory network is adopted. The occupant's head three-dimensional spatial position, head posture angle, and posture continuous reliability value at multiple consecutive time points are used as input feature sequences to model and learn the evolution relationship of occupant posture in the time dimension. The algorithm learns the dynamic evolution law of head three-dimensional spatial position and head posture angle changing with time and constructs a multi-occupant posture evolution prediction model. The multi-occupant posture evolution prediction model is used to make forward inferences on the trend of occupant posture changes before the sound field control is executed. Its output results serve as feedforward inputs for subsequent listening zone space generation, acoustic model prior parameter construction, and multi-occupant single sound field allocation strategy arbitration. This is used to adjust the sound field spatial constraints in advance, which is different from the passive compensation method that only corrects after sound energy leakage occurs. This reduces the risk of sound field mismatch and crosstalk, and improves the stability and response timeliness of single sound field control. Furthermore, the multi-occupant posture evolution prediction model is a lightweight temporal model trained offline and deployed on the vehicle. During the online phase, it only performs forward inference and does not involve online training updates, thus meeting the vehicle's computing power and real-time requirements. The model training process optimizes parameters based on historical posture temporal data, while the model inference process performs forward inference based on the posture temporal features within the current time window. The time window preferably covers continuous posture data within 0.5 to 3 seconds to ensure prediction stability while also considering the response speed to posture changes. The posture evolution prediction model infers the posture changes of each occupant within the time window, outputting the predicted three-dimensional spatial position and predicted head posture angle for each occupant in the future. The future prediction time is a short-term prediction time after the current time, preferably within the range of 0.1 to 1 second, reflecting the continuous evolution trend of the occupant's head posture in a short period. This provides forward-looking posture reference information for subsequent auditory zone spatial generation and adaptive adjustment of the monophonic field, realizing a shift in monophonic field control from a "post-correction" to a "prediction-driven" feedforward adjustment mode.
[0042] like Figure 5The diagram illustrates the multi-occupant posture evolution prediction model and its feedforward sound field driving mechanism. It shows the feedforward process of multi-occupant posture evolution prediction in single-sound field control: First, the historical posture temporal feature set of the occupants is used as the model input. This feature set consists of multi-dimensional information such as the three-dimensional spatial position of the head, head posture angle, and posture continuity reliability values, arranged chronologically to comprehensively characterize the spatial state, changing trend, and data reliability of the occupant posture. Then, the LSTM-based multimodal posture evolution prediction model models the temporal features, learning the dynamic evolution of occupant posture over time through a state memory and update mechanism, and outputting short-term predicted posture results, including the predicted three-dimensional head position and posture angle within a selectable range of 0.1 to 1 second. This prediction result serves as the feedforward input, directly used for the generation of the listening zone space, the construction of acoustic model priors, and the adjustment and arbitration of the independent sound field allocation strategy. This allows for pre-adaptation to occupant posture changes before sound field execution, realizing a shift from passive compensation to a prediction-driven single-sound field adjustment mechanism.
[0043] In this implementation scheme, by introducing multi-occupant attitude temporal modeling based on long short-term memory networks, joint temporal learning of occupant head spatial position, attitude changes, and their reliability is achieved. This can accurately characterize the attitude evolution trend within a short timescale while ensuring prediction stability. By making forward-looking predictions of future attitudes, reliable prior references are provided for the dynamic generation of the listening area space and the adaptive control of the individual sound field. This effectively reduces the impact of rapid attitude changes on the sound field matching accuracy, thereby improving the adaptability and responsiveness of multi-occupant cockpit sound field management in dynamic scenarios.
[0044] Specifically, the process of generating a spatial description of the hearing area and acoustic prior constraints that match the changes in occupant posture based on the prediction results is as follows: For each occupant, acoustic model prior parameters are generated based on the predicted three-dimensional spatial position of the head and the predicted head posture angle. Both the predicted three-dimensional spatial position and the predicted head posture angle are located in a unified cockpit coordinate system to ensure the consistency and feasibility of subsequent spatial calculations. A head coordinate system is constructed based on the head posture angle, with the head center position as the origin and the coordinate axis direction determined by the head posture angle. The structural vectors of both ears relative to the head center are rotated and superimposed with the head center position to obtain the predicted ear positions. The structural vectors of both ears relative to the head center are calibrated fixed vectors used to represent the relative geometric position of the ears in the head coordinate system. The midpoint of the predicted ear positions is taken as the target hearing area center position to approximately represent the geometric center of the hearing area. Based on the spatial dispersion of the target listening area center position within the prediction time window, the spatial coverage of the target listening area is determined. The spatial dispersion of the target listening area center position is obtained by statistically analyzing the three-dimensional spatial distribution of the target listening area center position within the prediction time window. Preferably, the spatial fluctuation of the listening area center over time is quantified by calculating the variance of the target listening area center position in each spatial dimension. In specific implementation, the spatial coverage can be abstractly represented as the region defined by the spatial statistical characteristics of the target listening area center position within the prediction time window. Preferably, it is described in the following way: the spatial coverage is defined as a spherical or approximately spherical region centered on the time mean of the target listening area center position and bounded by the spatial dispersion radius. The spatial dispersion radius is calculated based on the three-dimensional position variance of the target listening area center position within the prediction time window and is used to characterize the amplitude range of the listening area center fluctuation over time.Within the spatial coverage area of the target listening zone, candidate sampling points are generated around the center of the target listening zone. These candidate sampling points are uniformly distributed within the target listening zone according to spatial resolution, used for discretizing the listening zone space. For each candidate sampling point, based on its spatial orientation relative to the speaker position, the corresponding incident direction parameters are calculated. These parameters include azimuth and elevation angles, used to describe the incident direction characteristics of sound waves propagating from the speaker to the candidate sampling point. The corresponding HRTF is obtained from the HRTF database based on the incident direction. The HRTF database is an initialized general standard header related transfer function library used to reflect the binaural acoustic response characteristics under different incident directions. The HRTF corresponding to the candidate sampling points is fused based on the spatial distance from the candidate sampling point to the center of the target listening zone. The fusion process involves analyzing different... The HRTF of candidate sampling points is weighted and summed according to the corresponding distance weights, so that sampling points closer to the center of the listening area have a higher weight in the fusion result, and the HRTF prior of the target listening area is obtained, which serves as the spatial acoustic constraint condition for subsequent 3D audio rendering. Based on the speaker position and the center position of the target listening area, the equivalent propagation path parameters from the speaker to the target listening area are calculated. The equivalent propagation path parameters are used to characterize the distance, delay and attenuation characteristics of the sound signal from the speaker to the center of the target listening area to support subsequent propagation path compensation processing. The predicted three-dimensional spatial position of each occupant's head, the predicted head posture angle, and the corresponding acoustic model prior parameters are written into the single sound field management database and stored in association with the occupant index and time index, which provides unified and traceable acoustic prior data support for multi-occupant single sound field strategy generation, 3D audio rendering and closed-loop adjustment.
[0045] In this implementation scheme, by mapping the occupant attitude prediction results to a quantifiable spatial description of the listening area and acoustic prior constraints, the dynamic adaptive generation of the listening area center position, spatial coverage, and acoustic response parameters is achieved. By modeling the spatial discrete characteristics of the listening area center changing over time, and combining HRTF fusion and propagation path constraints, the acoustic priors can be updated in real time with changes in occupant attitude. This provides a stable, continuous, and forward-looking spatial acoustic foundation for subsequent 3D audio rendering and single-field control, significantly improving the accuracy and robustness of multi-occupant cockpit sound field management in dynamic scenarios.
[0046] Specifically, based on the spatial description of the occupant's listening area and acoustic prior constraints, combined with the currently active audio object type, the association between audio objects and occupants is established, generating a multi-occupant independent sound field allocation strategy. When sound field conflicts are detected, adjustment and smoothing are performed according to priority. The specific process is as follows: The vehicle's main controller retrieves the target listening area center position, target listening area spatial coverage, and corresponding acoustic model prior parameters for each occupant from the database according to the occupant index and timestamp. Simultaneously, it acquires the currently active audio object information and identifies the type of each audio object. The audio object information is provided by the vehicle's audio management component, which interfaces with the audio focus of the vehicle's operating system and can output audio object identifiers, input stream types, sampling rates, channel configurations, volume envelopes, and audio object information. The identification of audio object types preferably employs a function label mapping table to map navigation, calls, media, and prompts to preset types, thereby avoiding semantic understanding uncertainty. Based on the type of the audio object and the seat pressure center position, target listening area center position, and spatial coverage of each occupant, a spatial and functional association relationship is established between the audio object and the occupant. Specifically, the spatial association relationship is implemented as follows: each audio object is bound to a default interactive seat, and the seat pressure center position corresponding to the seat is read as the seat space representative point; the Euclidean distance from the seat space representative point to the target listening area center position of each occupant is calculated, and the target listening area spatial coverage is used to determine whether the occupant is within the serviceable space of the audio object; if so, a spatial association is established. The functional association relationship is implemented as follows: function indicators such as "call answerer seat," "voice wake-up triggered seat," and "navigation driver and passenger seat display bound seats" are obtained from the vehicle system interaction status and used as the target occupant candidate set for the audio object; when a function indicator exists, a strong functional association is established within the candidate set to ensure the certainty of the attribution of functions such as calls and navigation. When spatial association and functional association conflict, it is preferable to prioritize functional association and use spatial association as a secondary factor in the merging and determination process. This ensures that objects with clear objectives, such as calls and voice recordings, are not misassigned due to spatial proximity.The priority order of each audio object in the occupant's independent sound field is determined based on the spatial and functional relationship. The priority order is implemented in real time using a regularized queue at the vehicle end: the audio objects are preferably sorted in the basic order of "safety-related > call > navigation > system prompts > media playback", and then further sorted within the same type of object according to "current interaction seat matching degree" and "spatial association strength" to ensure that the priority determination is executable and reproducible at the millisecond level. Combined with the target listening area center position, spatial coverage range and acoustic model prior parameters corresponding to each occupant, an independent sound field allocation strategy is generated for each audio object. The specific implementation of the independent sound field allocation strategy is to output a set of parameter packages that can be directly sent to the DSP rendering link. The parameter package includes at least: audio object identifier, target occupant index, target listening area center position, target listening area spatial coverage range parameters, target listening area HRTF prior index, equivalent propagation path parameters and energy allocation ratio of the audio object. The parameter package is stored in a structured data format and sent to the DSP through inter-process communication. After generating independent sound field allocation strategies, a consistency check is performed on these strategies to identify sound field overlap. The specific implementation plan for the consistency check is as follows: For any two independent sound field allocation strategies, it is determined whether their target listening area spatial coverage areas geometrically overlap. If the overlap ratio exceeds an overlap ratio threshold, sound field overlap is considered to exist. The geometric overlap can be calculated by comparing the center distance between listening area envelopes (e.g., spheres or ellipsoids) with a scale threshold to ensure that the computational complexity meets the real-time requirements of the vehicle end. When sound field overlap is detected, the corresponding independent sound field allocation strategies are adjusted according to the priority order of the audio objects. Specifically, the adjustment scheme is as follows: the target listening area center position of high-priority audio objects remains unchanged, and for low-priority audio objects, the "shrink spatial coverage area" operation is performed sequentially. One or more of the following methods are used: "reduce the energy allocation ratio" and "change the rendering direction or virtual location of the sound source" to make different independent sound fields spatially separable; when the overlap of low-priority objects cannot be eliminated after adjustment, the low-priority objects are switched to shared sound field output to ensure the privacy and auditory consistency of high-priority objects; at the same time, the change of independent sound field allocation strategy over time is smoothed, preferably using exponential smoothing update, so that the parameter change rate between adjacent time moments is controlled, avoiding jitter at the listening zone boundary and sudden volume changes, and outputting the final independent sound field allocation strategy into the independent sound field management database and synchronously into the DSP rendering input queue so that the next rendering frame can read the latest strategy; at the same time, the conflict detection results, adjustment action type and adjusted parameters are recorded for subsequent model iteration and fault tracking.
[0047] In this implementation scheme, by integrating occupant posture perception, posture evolution prediction, and independent sound field allocation control, it is possible to pre-construct a listening zone description and acoustic constraints that match the occupant spatial state when multiple occupants are present and their postures are constantly changing. Based on the spatial and functional relationship of audio objects, a stable and distinguishable independent sound field allocation strategy is generated. When sound field overlap occurs, orderly arbitration and smooth adjustment are achieved, thereby effectively suppressing crosstalk and sound field drift, improving the privacy, stability, and overall listening consistency of multiple occupants listening independently in the cabin, and possessing good vehicle-side feasibility and real-time operational reliability.
[0048] Specifically, the 3D audio rendering process based on independent sound field allocation strategies and acoustic prior constraints for each audio object is as follows: The target listening area center position, spatial coverage, and acoustic model prior parameters for each occupant are read, along with the independent sound field allocation strategy for each audio object and its spatial and functional relationship with the occupant. The target listening area center position, spatial coverage, and acoustic model prior parameters are all derived from the independent sound field management database and are jointly retrieved using occupant and time indexes to ensure consistency of the read data under the same time reference, avoiding sound field shifts or rendering errors caused by data asynchrony. The independent sound field allocation strategy clarifies the mapping relationship between each audio object and different occupants and their corresponding spatial constraint range, while the spatial and functional relationship limits the audio object to participate in rendering only within the independent sound field of the corresponding occupant. The independent sound field allocation strategy and acoustic model prior parameters are aligned according to the occupant index and audio object index to construct a 3D audio rendering input set, which limits the rendering constraints of each audio object within the target listening area of the corresponding occupant. The alignment process is achieved by establishing a multi-dimensional mapping table of occupant index, audio object index, and acoustic parameter index, so that each audio object only calls the listening area parameters and acoustic prior constraints corresponding to its associated occupant during the rendering stage, thereby avoiding crosstalk or spatial aliasing problems caused by rendering audio objects across listening areas. The rendering constraints include at least the center position of the target listening area, spatial coverage, HRTF prior parameters, and equivalent propagation path parameters, which are used to limit the spatial positioning, directionality, and energy distribution of audio objects. Based on the 3D audio rendering input set, 3D audio rendering processing is performed on each audio object in the vehicle-mounted DSP. This includes: performing binaural rendering on the audio object based on the target listening zone HRTF prior, and performing propagation path compensation on the rendered signal based on the equivalent propagation path parameters. The vehicle-mounted DSP is the vehicle audio processing unit, used to perform real-time processing on the audio signal with a fixed frame length and a fixed sampling rate. The binaural rendering process is achieved by convolving the audio object signal with the corresponding left and right ear HRTFs respectively, so as to reconstruct the sound field effect that conforms to the spatial characteristics of the target listening zone at the occupant's two ears. Propagation path compensation is used to compensate for the difference in propagation distance between different speakers and the center of the target listening zone. The compensation methods include delay compensation and amplitude attenuation compensation to improve the accuracy of sound image localization and auditory consistency.Simultaneously, based on the independent sound field allocation strategy, partition rendering is performed on the target listening areas corresponding to different occupants, so that each audio object is mapped to the independent sound field output of the corresponding occupant. The partition rendering is achieved by configuring independent sound field rendering channels for different occupants in the DSP. The rendering channels are isolated from each other in terms of parameter calls and output mapping, so that the same audio object can be selectively rendered or suppressed in the independent sound fields corresponding to different occupants. After the rendering is completed, the audio signal is sent to the corresponding speaker or speaker combination through the vehicle audio output link, thereby forming an independent sound field output effect for different occupants in physical space.
[0049] In this implementation scheme, by aligning the independent sound field allocation strategy with the acoustic prior constraints corresponding to the occupants, 3D audio rendering with zoning constraints is implemented on the audio objects in the vehicle-side DSP. This ensures that each audio object is effectively presented only within the target listening area of the corresponding occupant, thereby reducing cross-zone crosstalk interference. At the same time, by combining propagation path compensation and zoning rendering mechanisms, the accuracy of sound image positioning and the consistency of listening experience are improved while ensuring real-time performance, thereby achieving stable output and fine control of independent sound fields in multi-occupant scenarios.
[0050] Specifically, during the sound field execution process, the acoustic energy leakage status of the target sound zone and non-target sound zone is monitored and evaluated. When a sound field interference trend is detected, adaptive suppression and parameter updates are triggered. The specific process is as follows: After the 3D audio rendering output is completed, for each occupant, the corresponding target sound zone audio signal and non-target sound zone audio signal are collected through the in-vehicle microphone array. The in-vehicle microphone array consists of multiple microphone units fixedly deployed in the cabin ceiling, pillars, or around the seats. Each microphone unit completes spatial position calibration during the initialization phase to distinguish the target sound zone sampling channel corresponding to the target listening area space and the non-target sound zone sampling channel located in other areas, so as to ensure the spatial orientation of the sound zone sampling. Consistency and uniformity can be achieved by mapping the spatial relationship between the geometric position of the microphone array and the spatial envelope of the listening area. Power spectral density estimation is performed on the acquired audio signal, preferably using a short-time Fourier transform combined with a window function to stably characterize the sound energy distribution in the time-frequency domain. Noise suppression processing is then applied to the power spectral density to reduce the impact of ambient background noise, inherent in-vehicle noise, and sensor background noise on subsequent evaluations. Frequency domain threshold suppression is preferred. This yields the denoised target sound region spectrum and the denoised non-target sound region spectrum. Energy integration is then performed on the denoised target sound region spectrum and the denoised non-target sound region spectrum within their respective frequency bands to obtain the energy of the target sound region. Non-target sound zone energy; where the frequency band is the set frequency band of interest, used to cover the effective sound energy distribution range of the main audio objects in the cabin; the setting of the frequency band of interest is based on the main spectral distribution characteristics of voice broadcasts, call voices and entertainment audio in the vehicle scene, and is comprehensively determined in combination with the spectral characteristics of in-vehicle background noise, so as to suppress low-frequency structural noise and high-frequency invalid noise, while highlighting the frequency components that have a significant impact on the hearing and privacy of the occupants; preferably, the lower limit frequency of the frequency band of interest is set to 200Hz to 400Hz, and the upper limit frequency is set to 6kHz to 8kHz, wherein: the frequency band below 200Hz mainly includes vehicle body structure vibration, road noise and low-frequency engine noise, which contribute less to speech and directional hearing; Frequency bands above 8kHz mainly include high-frequency environmental noise and sensor noise, which contribute little to the stability of sound leakage determination. In practice, the focus frequency band can be adaptively adjusted according to the type of audio object. When the audio object is speech, a frequency band range of 300Hz to 6kHz is preferred; when the audio object is music or multimedia audio, a frequency band range of 200Hz to 8kHz is preferred, so as to ensure the accuracy of the assessment while taking into account the spectral characteristics of different audio content. The frequency band parameters are configured during the initialization phase and stored in the independent sound field management database as a unified reference for sound energy leakage assessment and spectral distribution analysis, so as to ensure the comparability and consistency of assessment results between different times and different occupants.The energy of the non-target sound region is divided by the energy of the target sound region, and the logarithm of the ratio is taken to obtain the logarithmic energy leakage ratio. The logarithmic energy leakage ratio characterizes the overall energy leakage intensity of the non-target sound region relative to the target sound region, and can suppress scale instability caused by excessive energy amplitude differences. The denoised target sound region spectrum and the denoised non-target sound region spectrum are respectively subjected to energy normalization within their respective frequency bands. Normalization is achieved by dividing the energy at each frequency point by the total energy within the corresponding frequency band, eliminating absolute energy differences and retaining only the spectral morphology distribution characteristics, resulting in the corresponding normalized spectral distribution. Based on the normalized spectral distribution, the spectrum of the target sound region and the spectrum of the non-target sound region are calculated. The Jensen-Shannon divergence between the spectra yields the spectral distribution difference value, which is used to quantify the similarity between the target and non-target sound regions at the spectral structure level, reflecting whether crosstalk exhibits significant interference characteristics in spectral morphology. The spectral distribution difference value is added to a minimum positive number, and the square root is taken to obtain the normalized denominator. By taking the square root, the growth trend of the spectral difference is transformed from a linear relationship to a sub-linear relationship, thereby suppressing the excessive amplification effect of the spectral difference on the normalized denominator when the spectral difference is large. This improves the numerical stability, adjustment sensitivity, and controllability of the crosstalk leakage drive value under different spectral difference conditions. The logarithmic energy leakage ratio is divided by the normalized denominator, and the opposite of the denominator is used as the exponent for natural exponentiation. This exponential mapping method nonlinearly fuses energy leakage and spectral similarity, significantly amplifying cases with high leakage and similar spectra, while suppressing cases with low leakage or large spectral differences. The result of the natural exponentiation is added to a constant and the reciprocal is taken to obtain the partitioned crosstalk leakage drive value. The partitioned crosstalk leakage drive value is a dimensionless index used to uniformly characterize the crosstalk interference intensity level in the current monophonic field state. The crosstalk leakage drive value for each occupant's zone is calculated and compared with the allowable threshold. When the crosstalk leakage drive value exceeds the allowable threshold, which is the calibrated safety upper limit, it is used to distinguish between acceptable minor leakage and interference states requiring intervention, triggering the single-field closed-loop suppression control. This includes: dynamically updating the independent sound field allocation strategy and adaptively adjusting the corresponding acoustic model prior parameters. The adaptive adjustment includes fine-tuning the target listening area spatial coverage, HRTF weight, and propagation path compensation parameters to suppress crosstalk diffusion trends. Simultaneously, time smoothing is performed on the changes in the independent sound field allocation strategy and acoustic model prior parameters over time. Time smoothing is used to avoid the impact of frequent and drastic parameter changes on auditory stability, preferably achieved through exponential smoothing. The closed-loop updated independent sound field allocation strategy and acoustic model prior parameters are then written back to the single-field management database to form a traceable historical adjustment record and provide data support for subsequent sound field rendering iterations and system state analysis.
[0051] The specific formula for the partition crosstalk leakage driver value is as follows:
[0052] ;
[0053] In the formula, The driving value represents the crosstalk leakage of the zone, which is used to quantify the degree of sound energy leakage and potential interference risk between the target sound zone and the non-target sound zone. By comprehensively comparing the relative energy levels of the non-target sound zone and the target sound zone and the structural differences in their spectral distribution, the crosstalk trend is continuously evaluated during the sound field execution. When the energy of the non-target sound zone increases significantly and its spectral distribution gradually approaches that of the target sound zone, the driving value increases rapidly, thereby triggering the adaptive suppression and update of the independent sound field parameters, realizing closed-loop control of sound field interference. This represents the difference in spectral distribution, used to measure the degree of difference in spectral structure between the target sound area and the non-target sound area. When the spectra of the two are similar, it indicates that the crosstalk has a stronger perceptual interference. This represents the spectrum of the non-target sound region after denoising, reflecting the actual distribution of sound energy received within the non-target region. It is a direct observation for assessing the intensity of crosstalk leakage. This represents the target sound region spectrum after denoising, used to describe the desired sound energy distribution within the target sound region, and serves as a reference benchmark for crosstalk evaluation. The frequency band is used to limit the effective acoustic frequency band of interest in crosstalk assessment, so as to avoid irrelevant frequency bands from interfering with the judgment results. Represents extremely small positive numbers, used to prevent the denominator from being zero and to enhance the stability of numerical calculations. It does not participate in physical or acoustic semantic judgments, and its preferred value is [value missing]. .
[0054] In this embodiment, Table 1 is a data table of crosstalk leakage drive values for each zone. The table details the target sound zone energy, non-target sound zone energy, spectral distribution difference value, and zone crosstalk leakage drive value for each of the five samples. Specifically, sample 1 has a target sound zone energy of 1.20, a non-target sound zone energy of 0.06, a spectral distribution difference value of 0.42, and a zone crosstalk leakage drive value of 0.00973; sample 2 has a target sound zone energy of 1.10, a non-target sound zone energy of 0.15, a spectral distribution difference value of 0.36, and a zone crosstalk leakage drive value of 0.03487; sample 3... The target sound zone energy is 0.95, the non-target sound zone energy is 0.28, the spectral distribution difference value is 0.25, and the zone crosstalk leakage drive value is 0.07993; for sample 4, the target sound zone energy is 0.90, the non-target sound zone energy is 0.45, the spectral distribution difference value is 0.18, and the zone crosstalk leakage drive value is 0.16332; for sample 5, the target sound zone energy is 0.85, the non-target sound zone energy is 0.72, the spectral distribution difference value is 0.10, and the zone crosstalk leakage drive value is 0.37171.
[0055] surface Partition Crosstalk Leakage Driver Value Data Table
[0056] like Figure 3 The figure shows a comparison of the distribution of crosstalk leakage drive values in different zones. The horizontal axis represents the sample number, the vertical axis represents the crosstalk leakage drive value in different zones, and the dashed line represents the preset allowable threshold. (This is in conjunction with Table 1 and...) Figure 3 It can be seen that the crosstalk leakage drive values of samples 1 to 3 are all significantly lower than the allowable threshold, indicating that the sound field leakage is within an acceptable range and the isolation effect of the single sound field is good. The crosstalk leakage drive value of sample 4 is slightly higher than the allowable threshold, reflecting that the sound field has begun to show a perceptible risk of crosstalk. The crosstalk leakage drive value of sample 5 significantly exceeds the allowable threshold, indicating that the crosstalk leakage in the non-target sound zone is serious, and has reached the point where intervention is needed to trigger the single sound field closed-loop suppression and parameter adjustment. Overall, this reflects the effective ability of this indicator to distinguish the degree of sound field interference and the significance of threshold judgment.
[0057] like Figure 4 The diagram illustrates the monitoring and closed-loop suppression of crosstalk leakage in a single sound field. Using the two-dimensional space of the cabin as a reference coordinate, it shows the formation, leakage, and closed-loop control relationship of independent sound fields for multiple occupants. Speakers are located on one side of the cabin, radiating audio signals towards the target sound zone, forming the main acoustic energy coverage within that zone. These signals are sampled by microphones deployed within the target sound zone to characterize the actual acoustic energy state of the target listening area. Simultaneously, some acoustic energy propagates along the spatial path into non-target sound zones, causing crosstalk leakage. This leakage is collected by microphones in the non-target sound zones to assess the degree of interference between sound fields. Crosstalk assessment is performed based on the acoustic energy differences collected between the target and non-target sound zones. When an increasing trend in crosstalk leakage is detected, a closed-loop suppression and parameter update process is triggered, adaptively adjusting the independent sound field allocation strategy and acoustic parameters to suppress the diffusion of acoustic energy into non-target sound zones and maintain the spatial stability and auditory consistency of each occupant's independent sound field.
[0058] In this implementation scheme, a dual-dimensional joint evaluation of energy and spectrum for both target and non-target sound areas is conducted to achieve refined monitoring of crosstalk leakage during sound field execution. Based on a nonlinear fusion index of logarithmic energy leakage ratio and spectral distribution difference, a criterion that is highly sensitive to sound field interference trends and numerically stable is constructed. When an anomaly is detected, the independent sound field allocation strategy and acoustic prior parameters are adaptively updated in conjunction with a time smoothing mechanism to achieve timely suppression of crosstalk and stable auditory control. This effectively improves the privacy, robustness, and long-term operational stability of the independent sound field in dynamic multi-occupant scenarios.
[0059] Reference Figure 2As shown, the second aspect of the present invention provides a large model-driven cockpit multi-occupant single-sound field management device, applied to the aforementioned large model-driven cockpit multi-occupant single-sound field management method, comprising: an occupant state acquisition and preprocessing module, used to acquire occupant state data in the cockpit in real time, preprocess the occupant state data, evaluate the changes in occupant head posture based on the preprocessed occupant state data, construct a continuous and reliable posture representation, and perform judgment and smoothing processing to form a reliable state feature set; and an attitude prediction and sound field adaptation module, used to construct a multi-occupant attitude evolution prediction model based on the reliable state feature set, predict the evolution trend of occupant posture over time, and generate a model matching the occupant posture changes based on the prediction results. The system includes: a listening zone spatial description and acoustic prior constraints; a multi-occupant independent sound field strategy and arbitration module, which establishes the association between audio objects and occupants based on the occupant's listening zone spatial description and acoustic prior constraints, combined with the currently active audio object type, generates a multi-occupant independent sound field allocation strategy, and performs adjustment and smoothing processing according to priority when sound field conflicts are detected; and a 3D audio rendering and closed-loop control module, which performs 3D audio rendering with zoning constraints on each audio object based on the independent sound field allocation strategy and acoustic prior constraints, monitors and evaluates the acoustic energy leakage status of the target sound zone and non-target sound zone during sound field execution, and triggers adaptive suppression and parameter updates when sound field interference trends are detected.
[0060] In this implementation scheme, through the collaborative design of multi-source occupant state perception, attitude evolution prediction, independent sound field strategy arbitration, and closed-loop sound field control, the system achieves forward-looking management and adaptive adjustment of the sound field in the multi-occupant cockpit. It can stably construct and dynamically maintain an independent sound field in complex scenarios with continuous changes in occupant attitude and concurrent multiple audio objects, effectively suppressing sound field conflicts and crosstalk leakage, and improving the privacy, listening consistency, robustness, and feasibility of the cockpit audio system.
[0061] A third aspect of the present invention provides a storage medium for managing the multi-occupant single sound field in a large model-driven cockpit, comprising: the storage medium having one or more programs, the one or more programs being executed by one or more processors to implement a method for managing the multi-occupant single sound field in a large model-driven cockpit.
[0062] In this implementation scheme, the aforementioned storage medium solidifies the multi-occupant single-sound field management process driven by a large model in a programmatic manner, enabling key functions such as occupant state perception, attitude prediction, sound field allocation, and closed-loop suppression to be uniformly deployed and run stably on a general-purpose processor platform. It has good portability, scalability, and engineering feasibility, making it easy to reuse and upgrade in different vehicle models and hardware architectures, thereby improving the implementation efficiency and application value of the cockpit multi-occupant single-sound field management system.
[0063] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0064] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. As those skilled in the art will understand, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for managing the independent sound field of a multi-occupant cockpit driven by a large model, characterized in that: Includes the following steps: S1 collects occupant status data in the cabin in real time, preprocesses the occupant status data, evaluates changes in occupant head posture based on the preprocessed occupant status data, constructs a continuous and reliable posture representation, and performs judgment and smoothing processing to form a reliable state feature set. S2, based on the credible state feature set, constructs a multi-occupant posture evolution prediction model to predict the evolution trend of occupant posture over time, and combines the prediction results to generate a spatial description of the listening area and acoustic prior constraints that match the changes in occupant posture. S3, based on the spatial description of the occupant's listening area and acoustic prior constraints, combined with the currently activated audio object type, establishes the association between the audio object and the occupant, generates a multi-occupant independent sound field allocation strategy, and performs adjustment and smoothing processing according to priority when sound field conflicts are detected. S4 performs 3D audio rendering with partition constraints on each audio object based on an independent sound field allocation strategy and acoustic prior constraints. During the sound field execution, it monitors and evaluates the sound energy leakage status of the target sound zone and non-target sound zone. When a sound field interference trend is detected, it triggers adaptive suppression and parameter updates.
2. The method for managing the sound field of a large-scale model-driven cockpit with multiple occupants according to claim 1, characterized in that, The specific process for real-time acquisition of occupant status data in the cabin and preprocessing of the occupant status data is as follows: Real-time occupant status data is collected, including: acquiring the three-dimensional spatial position and head posture angle of each occupant's head through cockpit vision, and performing head key point detection by the cockpit vision perception unit. A corresponding key point response heatmap is generated for each head key point, and the maximum response value in the key point response heatmap is used as the detection confidence of the head key point; pressure distribution data of each seat is collected through seat occupancy sensors, and pressure weighted average calculation is performed on the pressure distribution data of each seat to obtain the center position of the sitting pressure, and the probability of occupying a seat is calculated based on the deviation between the current pressure distribution data and the empty seat pressure distribution. Based on a unified time reference, timestamps are added to the occupant status data, and resampling is performed on data with different sampling frequencies according to a unified sampling period; median filtering is used to perform preliminary smoothing and anomaly suppression on the head posture angle sequence; numerical range constraints and anomaly suppression are performed on the detection confidence and seat occupancy probability, and minimum-maximum normalization is performed on the occupant status data; a single-sound field management database is established, and the original and preprocessed occupant status data are written into the single-sound field management database.
3. The method for managing the sound field of a large-scale model-driven cockpit with multiple occupants according to claim 1, characterized in that, The specific process of evaluating changes in occupant head posture based on preprocessed occupant state data, constructing a continuous and reliable posture representation, and performing judgment and smoothing processing to form a reliable state feature set is as follows: Based on the preprocessed head attitude angle, the second-order time difference of the head attitude angle is calculated to obtain the attitude angular acceleration; based on the sliding time window, the median of the attitude angular acceleration is taken, and the median absolute deviation of the attitude angular acceleration is calculated. The normalized value of the attitude angular acceleration is obtained by subtracting the median of the attitude angular acceleration from the current attitude angular acceleration and dividing by the sum of the median absolute deviation of the attitude angular acceleration and the smallest positive number. The normalized value of the attitude angular acceleration is then compared with zero to perform non-negativity processing. The negative of the non-negative normalized value of the attitude angular acceleration is taken as the exponent and natural exponential operation is performed to obtain the attitude mutation suppression value. The attitude continuity confidence value is obtained by multiplying the detection confidence, the occupancy probability and the attitude mutation suppression value. Calculate the attitude continuity confidence value for each occupant and perform a validity determination: When the attitude continuity confidence value is greater than or equal to the confidence threshold, the current head three-dimensional spatial position and head attitude angle are determined to be valid attitude data; When the attitude continuity confidence value is less than the confidence threshold, it is determined that the confidence of the current head three-dimensional spatial position and head attitude angle is insufficient, and the head three-dimensional spatial position and head attitude angle are smoothed based on the median filtering and moving average filtering of the sliding time window. The attitude continuity confidence value and the smoothed data are written into the single-field management database.
4. The method for managing the sound field of a large-scale model-driven cockpit with multiple occupants according to claim 1, characterized in that, The specific process of constructing a multi-occupant attitude evolution prediction model based on a reliable state feature set to predict the evolution trend of occupant attitude over time is as follows: For each occupant, a multi-occupant attitude temporal feature set is constructed based on the preprocessed head 3D spatial position, head attitude angle, and attitude continuity reliability value recorded in the monophonic field management database. The historical multi-occupant attitude temporal feature set is then input into a multimodal temporal modeling algorithm according to the occupant index for temporal modeling. The algorithm learns the dynamic evolution law of the head 3D spatial position and head attitude angle over time and constructs a multi-occupant attitude evolution prediction model. The attitude evolution prediction model infers the attitude changes of each occupant within the time window and outputs the predicted head 3D spatial position and predicted head attitude angle of each occupant in the future.
5. The method for managing the sound field of a large-scale model-driven cockpit with multiple occupants according to claim 1, characterized in that, The specific process of generating a spatial description of the listening area and acoustic prior constraints that match the changes in occupant posture by combining the prediction results is as follows: For each occupant, acoustic model prior parameters are generated based on the predicted three-dimensional spatial position of the head and the predicted head posture angle: a head coordinate system is constructed based on the head posture angle, and the structural vectors of the two ears relative to the head center are rotated and superimposed with the head center position to obtain the predicted position of the two ears. The midpoint of the predicted position of the two ears is taken as the center position of the target hearing area. At the same time, the spatial coverage of the target hearing area is determined based on the spatial dispersion of the center position of the target hearing area within the prediction time window. Within the spatial coverage area of the target listening area, candidate sampling points are generated around the center of the target listening area; For each candidate sampling point, the corresponding incident direction parameters are calculated based on the spatial orientation relationship between the candidate sampling point and the speaker position. The corresponding HRTF is obtained from the HRTF database based on the incident direction. The HRTF corresponding to the candidate sampling point is fused based on the spatial distance from the candidate sampling point to the center of the target listening area to obtain the target listening area HRTF prior. The equivalent propagation path parameters from the speaker to the target listening area are calculated based on the speaker position and the center of the target listening area. The predicted three-dimensional spatial position of each occupant's head, the predicted head attitude angle, and the corresponding acoustic model prior parameters are written into the independent sound field management database.
6. The method for managing the sound field of a large-scale model-driven cockpit with multiple occupants according to claim 1, characterized in that, The specific process of establishing the association between audio objects and occupants based on the spatial description of the occupant's listening area and acoustic prior constraints, combined with the currently activated audio object type, generating a multi-occupant independent sound field allocation strategy, and performing adjustment and smoothing processing according to priority when sound field conflicts are detected is as follows: The system obtains the target hearing zone center position, target hearing zone spatial coverage, and corresponding acoustic model prior parameters for each occupant at the current moment. Simultaneously, it obtains the information of currently activated audio objects and identifies the type of each audio object. Based on the type of audio object and the seated pressure center position, target hearing zone center position, and spatial coverage of each occupant, it establishes the spatial and functional relationship between audio objects and occupants. Based on the spatial and functional relationship, the priority order of each audio object in the occupant's independent sound field is determined, and combined with the target listening zone center position, spatial coverage and acoustic model prior parameters of each occupant, an independent sound field allocation strategy corresponding to each audio object is generated. After generating the independent sound field allocation strategy, a consistency check is performed on the independent sound field allocation strategy to identify sound field overlap. When sound field overlap is detected, the corresponding independent sound field allocation strategy is adjusted according to the priority order of the audio objects. At the same time, time smoothing processing is performed on the changes of the independent sound field allocation strategy over time, and then the result is output.
7. The method for managing the sound field of a large model-driven cockpit with multiple occupants according to claim 1, characterized in that, The specific process of 3D audio rendering based on independent sound field allocation strategy and acoustic prior constraints, which applies partition constraints to each audio object, is as follows: Read the target listening area center position, spatial coverage and acoustic model prior parameters for each occupant, and read the independent sound field allocation strategy and spatial and functional relationship with the occupant for each audio object; align the independent sound field allocation strategy and acoustic model prior parameters according to the occupant index and audio object index to construct a 3D audio rendering input set, and limit the rendering constraints of each audio object in the corresponding occupant target listening area. Based on the 3D audio rendering input set, 3D audio rendering processing is performed on each audio object in the vehicle-mounted DSP, including: performing binaural rendering on the audio object based on the target listening area HRTF prior, and performing propagation path compensation on the rendering signal based on the equivalent propagation path parameters; at the same time, performing partition rendering on the target listening area corresponding to different occupants according to the independent sound field allocation strategy, so that each audio object is mapped to the independent sound field output of the corresponding occupant.
8. The method for managing the sound field of a large model-driven cockpit with multiple occupants according to claim 1, characterized in that, The specific process of monitoring and evaluating the acoustic energy leakage status of the target and non-target acoustic regions during sound field execution, and triggering adaptive suppression and parameter updates when a sound field interference trend is detected, is as follows: After completing the 3D audio rendering output, for each occupant, the corresponding target sound zone audio signal and non-target sound zone audio signal are collected through the in-vehicle microphone array. The power spectral density of the collected audio signals is estimated, and noise suppression processing is performed on the power spectral density to obtain the denoised target sound zone spectrum and the denoised non-target sound zone spectrum. Energy integration is performed on the denoised target sound region spectrum and the denoised non-target sound region spectrum within the frequency band to obtain the target sound region energy and the non-target sound region energy. The non-target sound region energy is divided by the target sound region energy, and the logarithm of the obtained ratio is taken to obtain the logarithmic energy leakage ratio. Energy normalization is performed on the denoised target sound region spectrum and the denoised non-target sound region spectrum within the frequency band to obtain the corresponding normalized spectral distribution. The Jensen-Shannon divergence between the target sound region spectrum and the non-target sound region spectrum is calculated based on the normalized spectral distribution to obtain the spectral distribution difference value. The spectral distribution difference value is added to the smallest positive number and the square root is taken to obtain the normalized denominator term. Divide the logarithmic energy leakage ratio by the normalized denominator, take the opposite of the number as the exponent for natural exponentiation, add the result of natural exponentiation to the constant 1, and take the reciprocal to obtain the partition crosstalk leakage drive value. Calculate the crosstalk leakage drive value for each occupant and compare it with the allowable threshold. When the crosstalk leakage drive value exceeds the allowable threshold, trigger the independent sound field closed-loop suppression control, including: dynamically updating the independent sound field allocation strategy and adaptively adjusting the corresponding acoustic model prior parameters. Simultaneously, time smoothing processing is performed on the changes of independent sound field allocation strategy and acoustic model prior parameters over time; and the closed-loop updated independent sound field allocation strategy and acoustic model prior parameters are written back to the independent sound field management database.
9. A large-scale model-driven cockpit multi-occupant single-sound field management device, characterized in that: include: The occupant status acquisition and preprocessing module is used to acquire occupant status data in the cabin in real time, preprocess the occupant status data, evaluate the changes in occupant head posture based on the preprocessed occupant status data, construct a continuous and reliable posture representation, and perform judgment and smoothing processing to form a reliable status feature set. The attitude prediction and sound field adaptation module is used to build a multi-occupant attitude evolution prediction model based on a reliable state feature set, predict the evolution trend of occupant attitude over time, and generate a listening area spatial description and acoustic prior constraints that match the changes in occupant attitude by combining the prediction results. The multi-occupant independent sound field strategy and arbitration module is used to establish the association between audio objects and occupants based on the spatial description of the occupant's listening area and acoustic prior constraints, combined with the currently active audio object type, to generate a multi-occupant independent sound field allocation strategy, and to perform adjustment and smoothing processing according to priority when sound field conflicts are detected. The 3D audio rendering and closed-loop control module is used to perform 3D audio rendering with partition constraints on each audio object based on independent sound field allocation strategy and acoustic prior constraints. During the sound field execution, it monitors and evaluates the sound energy leakage status of the target sound zone and non-target sound zone, and triggers adaptive suppression and parameter update when sound field interference trend is detected.
10. A large-scale model-driven cockpit multi-occupant independent sound field management storage medium, characterized in that: include: The storage medium has one or more programs, which are executed by one or more processors to implement the large model-driven cockpit multi-occupant single-sound field management method as described in claims 1-8.
Citation Information
Patent Citations
Method and system for adjusting position and volume of sound field of cabin, vehicle and storage medium
CN117041857A
In-vehicle sound field optimization method, audio system, electronic equipment and storage medium
CN118042398B