An automatic driving takeover prompt tone regulation method and system based on sound field monitoring

CN122343680BActive Publication Date: 2026-08-07JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2026-06-08
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

仅根据整体噪声值提高提示音音量,难以解决提示音在特定频段被音乐、人声或风噪、路噪掩蔽的问题;同时,过度提高提示音响度还可能造成驾驶员惊吓、乘员烦扰,影响接管动作的稳定性和乘坐舒适性

Benefits of technology

[0012]本发明的有益效果是:该基于声场监测的自动驾驶接管提示音调控方法及系统,

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122343680B_ABST
    Figure CN122343680B_ABST
Patent Text Reader

Abstract

The application provides an automatic driving takeover prompt tone control method and system based on sound field monitoring, applied to the technical field of automatic driving. After obtaining takeover trigger information output by an automatic driving system, the method collects real-time audio signals in a vehicle within a takeover prompt full cycle, extracts dynamic sound field features in the vehicle, and constructs an acoustic state in the vehicle containing a frequency band masking distribution and a local audibility risk of a driving position. According to the acoustic state, the frequency spectrum parameters, rhythm parameters and loudness parameters of a takeover safety alarm tone are generated, and multi-objective constraint coupling optimization is performed under the constraints of safety, comfort and recognizability to obtain optimal alarm tone control parameters. Then, the corresponding takeover safety alarm tone is generated or matched and played, and the playing parameters are dynamically adjusted according to the changes in the sound field in the vehicle. The application can reduce the masking effect of a complex sound field in the vehicle on the takeover prompt tone, improve the timeliness of the driver's awareness of the takeover request and the recognition of the prompt tone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving technology, specifically relating to a method and system for controlling autonomous driving takeover prompts based on sound field monitoring. Background Technology

[0002] With the development of advanced driver assistance systems (ADAS) and conditional automated driving (MAD) functions, vehicles need to issue takeover requests to the driver in specific operating scenarios to prompt the driver to relinquish control of the vehicle within a limited time. Whether the takeover prompt can be perceived and accurately recognized by the driver in a timely manner directly affects the safety of human-machine handover during the exit phase of the automated driving function. Especially when the takeover time window is short, the road environment is complex, or the driver's attention is distracted by cockpit activities, the perceptibility and recognizability of the takeover prompt tone are of great significance for ensuring driving safety.

[0003] Meanwhile, the audio environment in intelligent connected cockpits is becoming increasingly complex. In addition to traditional in-car music, navigation announcements, and phone calls, video conferencing, cockpit entertainment, rear-seat independent audio-visual systems, external speakers from passenger mobile devices, screen projection, and conversations among in-vehicle occupants frequently occur during vehicle operation. Consequently, the in-vehicle sound field is no longer a relatively stable background primarily composed of wind noise, road noise, and engine noise, but rather exhibits characteristics of multiple sound source superposition, complex frequency components, changing sound source locations, and continuous fluctuations in the sound field state. The propagation paths and frequency band distributions of different sound sources within the vehicle space vary, easily creating localized sound field changes near the driver's seat. This can cause the takeover safety warning sound to be masked by music, human voices, or ambient noise in certain frequency bands, resulting in the driver being unable to promptly perceive or accurately identify the takeover warning even when within audible range.

[0004] Existing takeover alert solutions typically employ a fixed-volume broadcast alert tone, or simultaneously use multimodal alerts (visual, auditory, and tactile) when a takeover request is triggered. Some solutions adjust the triggering timing, content, or intensity of the takeover alert based on driving environment risks, driver status, or autonomous driving system status; others enhance the audibility of the alert tone by detecting the overall noise level inside the vehicle and increasing the volume when ambient noise is high. While these solutions can provide some alerting effect under normal in-vehicle noise conditions, their adjustment mechanisms are largely based on overall noise levels, takeover risk levels, or driver status, failing to fully reflect the impact of the complex sound field inside the vehicle on the actual perceptibility of the alert tone.

[0005] Specifically, most existing noise-adaptive alert solutions equate changes in in-vehicle noise to changes in overall sound pressure level and compensate by increasing the volume of the alert tone. However, in the complex in-vehicle sound field, whether the alert tone can be perceived by the driver in a timely manner depends not only on the overall loudness but also on the energy distribution of background sound across different frequency bands, the local sound field state at the driver's seat, and the masking effect of human hearing. Simply increasing the alert tone volume based on the overall noise level is insufficient to address the problem of alert tones being masked by music, human voices, wind noise, or road noise in specific frequency bands. Furthermore, excessively increasing the alert tone volume may startle the driver, annoy passengers, and affect the stability of takeover actions and ride comfort.

[0006] Furthermore, while some existing solutions involve adjusting the alarm audio frequency, volume, or prompt format, they typically only control a single parameter independently, lacking a comprehensive consideration of the interplay between the audio spectrum characteristics, rhythmic characteristics, and loudness characteristics of the prompt. For example, simply changing the prompt audio frequency may affect its original recognizability, simply increasing the repetition frequency may increase tension, and simply increasing the loudness may easily cause interference and discomfort. For the safety scenario of autonomous driving takeover prompts, the prompt sound needs to be sufficiently perceptible in a complex sound field while maintaining high recognizability and appropriate comfort. Existing single-parameter, open-loop adjustment methods are insufficient to meet these requirements simultaneously.

[0007] Meanwhile, the in-vehicle sound field may continuously change during the takeover warning process, such as music track changes, occupants starting or stopping conversations, changes in audio output from mobile devices, and increased wind and road noise due to changes in vehicle speed. Existing solutions often use fixed warning parameters when takeover is triggered, or make simple adjustments based on the noise level at a specific moment, lacking a mechanism to continuously track changes in the in-vehicle sound field and dynamically adapt the warning tone parameters throughout the entire warning cycle. Therefore, in a cabin environment with multiple mixed sound sources and a continuously changing sound field, existing takeover warning tones are prone to problems such as being masked, insufficient intelligibility, increased response delay, or excessive interference.

[0008] In summary, existing technologies still struggle to achieve a takeover safety warning tone control method that balances perceptibility, recognizability, low interference, and dynamic adaptability under complex dynamic in-vehicle sound fields. Summary of the Invention

[0009] In view of the above-mentioned problems in the prior art, the purpose of the present invention is to provide an autonomous driving takeover warning tone control method based on sound field monitoring, which can combine the local sound field influence of the driver's seat and the frequency band masking characteristics to perform multi-dimensional coordinated control of the takeover safety warning tone, and continuously adapt to sound field changes during the takeover warning process.

[0010] A method for controlling the automatic driving takeover prompt tone based on sound field monitoring includes the following steps: Acquire and respond to takeover trigger information output by the vehicle's autonomous driving system, collect real-time audio signals inside the vehicle throughout the entire takeover prompt cycle, and extract dynamic sound field features inside the vehicle that are updated frame by frame. Based on the in-vehicle dynamic sound field characteristics, an in-vehicle acoustic state vector is constructed. The in-vehicle acoustic state vector includes at least the frequency band masking distribution and the risk of local audibility in the driver's seat. The spectral parameters, rhythm parameters, and loudness parameters of the takeover safety alarm sound are generated based on the in-vehicle acoustic state vector. The spectral parameters are determined according to the frequency band masking distribution so that the core energy of the takeover safety alarm sound avoids strong masking frequency bands. The rhythm parameters and loudness parameters are determined in conjunction with the spectral parameters and the local audibility risk of the driver's seat. Under preset constraints, the spectrum parameters, rhythm parameters, and loudness parameters are optimized by multi-objective constraint coupling to obtain the optimal alarm tone control parameters; The takeover safety alarm tone is generated or matched according to the optimal alarm tone control parameters and played, and the playback parameters of the takeover safety alarm tone are dynamically adjusted according to the update of the in-vehicle dynamic sound field characteristics throughout the entire takeover prompt cycle.

[0011] Another objective of this invention is to provide an autonomous driving takeover warning tone control system based on sound field monitoring, characterized in that the method for implementing the above-mentioned autonomous driving takeover warning tone control based on sound field monitoring includes: The takeover event acquisition module is used to acquire takeover trigger information output by the vehicle's autonomous driving system. The in-vehicle sound field monitoring module is used to respond to the takeover trigger information, collect real-time audio signals in the vehicle during the entire takeover prompt cycle, and extract dynamic sound field features in the vehicle that are updated frame by frame. An acoustic state construction module is communicatively connected to the in-vehicle sound field monitoring module and is used to construct the in-vehicle acoustic state based on the in-vehicle dynamic sound field characteristics. The in-vehicle acoustic state includes at least the frequency band masking distribution and the risk of local audibility in the driver's seat. The prompt tone parameter decision module is communicatively connected to the acoustic state construction module and is used to generate the spectral parameters, rhythm parameters, and loudness parameters of the takeover safety alarm tone based on the in-vehicle acoustic state. The constraint optimization module is communicatively connected to the parameter decision module and is used to perform multi-objective constraint coupling optimization on the spectrum parameters, rhythm parameters and loudness parameters under preset constraints to obtain the optimal alarm sound control parameters. The prompt tone generation and playback module is communicatively connected to the constraint optimization module. It is used to generate or match the takeover safety alarm tone according to the optimal alarm tone control parameters, and control the in-vehicle speaker system to play the takeover safety alarm tone. The in-vehicle sound field monitoring module, acoustic state construction module, prompt tone parameter decision module, constraint optimization module, and prompt tone generation and playback module are executed cyclically throughout the entire takeover prompt cycle as the dynamic sound field characteristics of the in-vehicle are updated frame by frame, so as to dynamically adjust the playback parameters of the takeover safety alarm tone.

[0012] The beneficial effects of this invention are: the autonomous driving takeover prompt tone control method and system based on sound field monitoring, After an autonomous driving takeover request is triggered, the volume of the takeover safety warning tone is no longer simply increased based on the overall noise level inside the vehicle. Instead, the system collects real-time audio signals from inside the vehicle, extracts dynamic sound field characteristics, and constructs an in-vehicle acoustic state that includes frequency band masking distribution and local audibility risk in the driver's seat. This allows the control of the takeover safety warning tone to more accurately reflect the impact of the complex sound field inside the vehicle on the driver's actual auditory perception, thereby improving the targeting of the takeover warning control.

[0013] By assessing the potential masking effect of different frequency bands in the current in-vehicle sound field on the emergency call safety warning tone through frequency band masking distribution, and accordingly determining the spectral parameters of the warning tone, the core energy of the warning tone can be directed away from strong masking frequency bands. Therefore, compared to simply increasing the overall volume of the warning tone, this invention can improve the perceptibility of the warning tone in the driver's seat area without significantly increasing the playback loudness, and reduce the masking effect of background noise such as music, human voices, wind noise, and road noise on the emergency call safety warning tone.

[0014] Further incorporating the risk of local audibility in the driver's seat, this design integrates factors such as the local signal-to-noise ratio in the driver's seat, the overlap between the alarm audio band and the masking band, and interference from in-vehicle voice activity into the takeover warning tone control process. This shifts the optimization goal of the warning tone from "overall audibility in the vehicle" to "timely detection in the driver's area." This design better adapts to takeover safety warning scenarios and reduces issues such as missed hearing, mishearing, or response delays caused by unfavorable sound fields in the driver's seat area.

[0015] The system coordinates and regulates the spectral, rhythmic, and loudness parameters of the safety alarm tone, rather than adjusting individual audio parameters independently. When the spectral parameters change according to the masking distribution, the system synchronously adjusts parameters such as pulse repetition rate, envelope rise time, loudness increment, and initial transient increment. This improves the detectability and transient recognition of the alarm tone while avoiding problems such as startling the driver, disturbing passengers, or abruptly raising the alert tone due to simply increasing the loudness, thus achieving a balance between detectability, recognizability, and comfort.

[0016] Under preset safety constraints, comfort constraints, and recognizability constraints, the alarm tone parameters are optimized through multi-objective constraint coupling. This ensures that the final output takeover safety alarm tone not only meets the priority and recognizability requirements of takeover safety prompts, but also maintains recognizability consistency with standard takeover prompt tones, and suppresses parameter jumps between adjacent audio frames, thereby improving the smoothness and stability of the prompt tone dynamic adjustment process.

[0017] Throughout the entire takeover alert cycle, the system continuously updates the dynamic sound field characteristics inside the vehicle and dynamically adjusts the playback parameters of the takeover safety alarm tone based on the updated in-vehicle acoustic conditions. Therefore, when in-vehicle music tracks change, the intensity of occupant conversations changes, audio is played or stopped from mobile devices, or wind and road noise changes due to vehicle speed variations, the system can adapt to these sound field changes in real time. This prevents the fixed-parameter alert tone from being masked by newly emerging background noise during the takeover process, improving the sustained effectiveness of the takeover alert.

[0018] In addition, an abnormal operating condition detection and degradation processing mechanism is set up. In the event of microphone failure, continuous audio frame loss, excessive computing power load on the controller, speaker failure, or abnormal audio playback link, the optimization process based on real-time sound field feedback can be skipped, and the pre-calibrated safety fallback alarm tone parameter scheme can be called to continue outputting the takeover safety prompt tone, thereby ensuring the stability and safety redundancy of the takeover prompt function under abnormal operating conditions.

[0019] In summary, this invention addresses the problems of easily masked, insufficiently perceptible, and reduced intelligibility of takeover warning sounds in complex dynamic in-vehicle sound fields, as well as the interference caused by excessively increasing the volume. It achieves dynamic, refined, and coordinated control of takeover safety warning sounds, improves the timeliness of the driver's perception of takeover requests and the effectiveness of the response, while also taking into account the auditory comfort of in-vehicle occupants and the reliability of system engineering implementation. Attached Figure Description

[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a system block diagram of the present invention. Detailed Implementation

[0021] Example 1 A method for controlling autonomous driving takeover prompts based on sound field monitoring is disclosed. This method is executed by a takeover prompt controller integrated into the vehicle's autonomous driving domain controller, intelligent cockpit domain controller, or independent audio safety controller. The takeover prompt controller is communicatively connected to the vehicle's autonomous driving system, in-vehicle microphone, in-vehicle audio playback system, and in-vehicle speaker system.

[0022] In this embodiment, the takeover safety warning tone is preferably a pulsed takeover safety alarm tone. This pulsed takeover safety alarm tone consists of one to three pulse signals with fixed or adjustable base frequencies. It does not contain complex speech content and is mainly used to quickly awaken the driver's attention in autonomous driving takeover scenarios. Because pulsed alarm tones have obvious transient characteristics, are less prone to semantic distortion after parameter adjustment, and are easier to synthesize or match in real time, they are suitable for real-time adaptive adjustment based on the dynamic sound field inside the vehicle throughout the entire takeover warning cycle.

[0023] For ease of description, this embodiment uses audio frames as the basic update unit. The current audio frame number is denoted as t. The human ear's sensitive frequency band from 100Hz to 8kHz is divided into B 1 / 3 octave bands, and the center frequency of the b-th band is denoted as t. Where b = 1, 2, ..., B. The frame length of the audio frame is preferably 10 ms to 30 ms, the frame shift is preferably 5 ms to 15 ms, and the sampling rate is preferably 44.1 kHz or 48 kHz to adapt to existing in-vehicle audio hardware platforms.

[0024] In this embodiment, the symbol This represents a limiting function that restricts the variable y to between 0 and 1; symbol This represents a very small positive number used to prevent the denominator from being zero.

[0025] like Figure 1 As shown, the method for controlling the automatic driving takeover prompt sound based on sound field monitoring includes the following steps: S1. Obtain the takeover trigger information output by the vehicle's autonomous driving system. When the vehicle's autonomous driving system detects that the vehicle has entered a scenario requiring the driver to take over control, the system outputs a takeover trigger message to the takeover prompt controller. This takeover trigger message indicates that the vehicle has entered a scenario requiring the driver to take over control.

[0026] The takeover trigger information may include one or more of the following: takeover request identifier, takeover request level, remaining takeover time, takeover trigger reason, current vehicle speed, and autonomous driving system operating status. The takeover request scenarios include, but are not limited to: the vehicle is about to exceed the functional operating design domain of the autonomous driving system; the risk level of the road scene ahead exceeds the system's processing capacity; the vehicle perception system experiences a degradeable fault; the autonomous driving planning and control module cannot continue to execute stably; the driver's attention does not meet the takeover requirements; or the vehicle safety strategy determines that driver intervention is necessary.

[0027] Upon receiving the takeover trigger information, the takeover alert controller initiates the full-cycle takeover alert control process. The full-cycle takeover alert refers to the time period from the receipt of the takeover trigger information until the driver completes the takeover, the takeover alert ends, or the system enters a degraded alert state.

[0028] S2. Acquire real-time audio signals from inside the vehicle and extract dynamic sound field characteristics within the vehicle throughout the entire takeover period. In response to the takeover trigger information, the takeover alert controller continuously collects real-time audio signals from inside the vehicle through at least one microphone deployed throughout the entire takeover alert period. Multiple automotive-grade microphones are installed inside the vehicle, and these microphones can be positioned above the driver's seat, above the passenger seat, in the front center console area, and in the rear headliner area, to cover the driver's seat and other major sound source areas within the vehicle.

[0029] The takeover prompt controller preprocesses and extracts features from the acquired real-time in-vehicle audio signals to obtain in-vehicle dynamic sound field features updated frame by frame. The preprocessing may include frame segmentation, DC component removal, pre-emphasis filtering, window function weighting, spectrum analysis, echo cancellation, and frequency band energy statistics. These processes are conventional implementation methods in in-vehicle audio signal processing, and this embodiment does not elaborate on their basic mathematical formulas.

[0030] In this embodiment, the in-vehicle dynamic sound field characteristics include at least one of the following: overall A-weighted background noise level, 1 / 3 octave band sub-band sound pressure level, short-time spectral energy distribution, voice activity detection results, energy proportion of the main interference sound source, local sound pressure level in the driver's seat, and playback status of the in-vehicle audio system.

[0031] Let the local sound pressure level in the driver's ear region of the t-th frame be in the b-th frequency band. The local sound pressure level can be estimated by combining the signal collected by the in-vehicle microphone with the pre-calibrated cabin acoustic transfer function.

[0032] Let the overall A-weighted background noise level of the driver's seat area in frame t be... It is used to characterize the overall background sound intensity near the current driver's seat.

[0033] In this embodiment, a pre-calibrated cabin acoustic transfer function is used to characterize the acoustic transfer relationship between the microphone placement location in the vehicle and the driver's ear area. This transfer function can be obtained through vehicle cabin acoustic simulation, real-vehicle frequency sweep testing, or acoustic calibration experiments, and is pre-stored in the takeover prompt controller. During actual operation, the system calls this transfer function to convert the sound field information collected by the microphone into a local sound field estimation result for the driver's ear area.

[0034] S3. Constructing an in-vehicle acoustic state vector based on real-time updated in-vehicle dynamic sound field features. The takeover warning controller constructs an in-vehicle acoustic state vector for takeover warning scenarios based on the dynamic sound field characteristics updated frame-by-frame. This in-vehicle acoustic state vector contains at least two core features: frequency band masking distribution and local audibility risk in the driver's seat.

[0035] S3.1 Calculation of frequency band masking distribution: To characterize the potential frequency band masking effect of the current in-vehicle sound field on the takeover safety alarm tone, the masking threshold of each frequency band is first calculated based on the local sound pressure level of each frequency band in the driver's ear region. The masking threshold of the b-th frequency band is defined as:

[0036] In the formula, This represents the masking threshold for the b-th frequency band in the t-th frame; This represents the local sound pressure level in the driver's ear region of the t-th frame within the b-th frequency band; This represents the pre-calibration frequency band masking correction coefficient corresponding to the b-th frequency band. This pre-calibration frequency band masking correction coefficient can be determined based on the psychoacoustic characteristics of the human ear, the acoustic characteristics of the vehicle cabin, and the calibration results of the actual vehicle.

[0037] To facilitate uniform processing under different sound field intensities, the masking thresholds of each frequency band are normalized to obtain the normalized masking intensity of the b-th frequency band:

[0038] In the formula, This represents the normalized masking strength of the b-th frequency band in the t-th frame; This represents the minimum masking threshold among all frequency bands in frame t; This represents the maximum masking threshold across all frequency bands in frame t.

[0039] The frequency band masking distribution vector is composed of the normalized masking intensity of each frequency band:

[0040] In the formula, Let represent the frequency band masking distribution vector of the t-th frame. The larger the value, the stronger the potential masking effect of the b-th frequency band on the takeover security alarm sound; The smaller the value, the more suitable the frequency band is as the area for deploying the core energy of the safety alarm sound.

[0041] Specifically, the system can determine the set of strong masking frequency bands based on the following conditions:

[0042] In the formula, Let represent the set of strong masking frequency bands in frame t; This represents the threshold for determining strong masking, with a preferred value range of 0.60 to 0.80.

[0043] S3.2, Calculation of Local Hearability Risk in the Driver's Seat To determine whether the takeover safety alarm sound is easily detected by the driver in the current in-vehicle sound field, the risk of local audibility in the driver's seat is further calculated.

[0044] Let the normalized energy percentage of the candidate takeover security alarm tone in the b-th frequency sub-band be . And satisfy:

[0045] The overlap between the takeover security alarm audio band and the current frequency band masking distribution is defined as follows:

[0046] In the formula, This indicates the degree of overlap between the energy distribution of the takeover security alarm sound in frame t and the current frequency band masking distribution; This represents the normalized energy percentage of the alarm tone within the b-th frequency sub-band. This represents the normalized masking strength of the b-th frequency band. The higher the value, the more likely the alarm sound is to fall into a strong masking frequency band, and the higher the risk of the driver missing it.

[0047] The signal-to-noise ratio for local estimation of the driver's seat is defined as:

[0048] In the formula, This represents the local signal-to-noise ratio estimated from the driver's seat in frame t; This indicates the predicted A-weighted sound pressure level of the current candidate takeover safety alarm sound in the driver's ear region; This represents the overall A-weighted background noise level in the driver's seat area of ​​frame t.

[0049] The risk of partial audibility in the driver's seat is defined as:

[0050] In the formula, This indicates the risk of local audibility in the driver's seat within frame t; Indicates the preset signal-to-noise ratio threshold; This indicates the degree of overlap between the alarm audio band and the current frequency band masking distribution; This represents the intensity of in-vehicle voice activity in frame t, with a value ranging from 0 to 1; , , represent the weighting coefficients of insufficient signal-to-noise ratio, bandwidth masking overlap, and speech activity interference on audibility risk, respectively. .

[0051] By adopting the above-mentioned formula for defining the risk of local audibility in the driver's seat, we can avoid judging the audibility of the warning sound solely based on the overall noise level inside the vehicle. It also considers the local signal-to-noise ratio in the driver's seat, the degree of overlap between the warning sound band and the current strong masking frequency band, and the interference from voice activities inside the vehicle. This allows us to more accurately reflect the actual perceptible risk of the driver to the takeover safety alarm sound in autonomous driving takeover scenarios.

[0052] S3.3 Construction of in-vehicle acoustic state vectors Based on the above characteristics, construct the in-vehicle acoustic state vector:

[0053] In the formula, This represents the acoustic state vector inside the vehicle in frame t. This indicates the overall A-weighted background noise level in the driver's seat area; Represents the frequency band masking distribution vector; Indicates the intensity of voice activity inside the vehicle; This indicates a risk of localized hearing impairment in the driver's seat. This indicates the playback status of the vehicle's audio system.

[0054] Among them, the playback status of the in-vehicle audio system This includes one or more of the following: in-vehicle audio playback identifier, audio type, current volume level, spectrum status of in-vehicle music or navigation broadcast, and whether there is audio from a conference call, rear-seat entertainment, or external audio from a passenger's mobile device.

[0055] By constructing the aforementioned in-vehicle acoustic state vector, the system transforms the influence of the complex dynamic sound field inside the vehicle on the takeover safety alarm sound into structured state data that can be used for subsequent parameter decisions.

[0056] S4. Generate a set of target parameters for the takeover safety warning tone based on the real-time updated in-vehicle acoustic state vector. The takeover prompt controller is based on the in-vehicle acoustic state vector of frame t. This generates a set of target parameters for the takeover safety alarm sound. This set of target parameters includes three categories of coordinated control parameters: loudness parameters, spectral parameters, and rhythm parameters, denoted as:

[0057] In the formula, This represents the set of target parameters for taking over the security alarm sound in frame t; Represents the set of loudness parameters; Represents a set of spectral parameters; This represents the set of rhythm parameters.

[0058] The set of loudness parameters is as follows:

[0059] In the formula, Indicates the absolute playback level of the alarm tone; This indicates the loudness increment of the alarm sound relative to the background noise level; This indicates the transient increment at the start of the alarm tone.

[0060] The set of spectral parameters is as follows:

[0061] In the formula, Indicates the base frequency of the alarm tone; Indicates the width of the alarm audio band; This represents the set of weighted harmonic components of the alarm tone.

[0062] The rhythm parameter set is as follows:

[0063] In the formula, Indicates the pulse repetition rate; Indicates the duration of a single pulse; Indicates the rising time of the envelope; Indicates the time of the falling edge of the envelope; Indicates the number of times the video will be played repeatedly.

[0064] S4.1 Determination of Spectrum Parameters When determining spectral parameters, the system prioritizes the frequency band masking distribution. Select a frequency band with a lower degree of masking as the core energy distribution area for the alarm sound, so that the alarm sound avoids the strong masking frequency band in the current in-vehicle sound field.

[0065] In one specific implementation, the candidate frequency band set is defined as follows:

[0066] In the formula, Represents the set of candidate frequency bands for frame t; This represents the center frequency of the b-th frequency band; and These represent the lower and upper limits of the permissible range of the base frequency for the takeover security alarm tone, respectively. This indicates the threshold for determining strong masking.

[0067] When the candidate frequency band set When not empty, the frequency band with the lowest masking strength or two adjacent frequency bands are selected as the target frequency band. This embodiment takes the selection of a single target frequency band as an example, and the target frequency band satisfies:

[0068] And the alarm tone fundamental frequency was determined to be: ; In the formula, This represents the target frequency band selected in frame t; This indicates the base frequency of the security alarm tone taken over in frame t; This indicates the center frequency of the target frequency band.

[0069] When the candidate frequency band set When empty, the system can select a sub-band with relatively low masking strength as the target frequency band within the allowed baseband range, or call a preset safe baseband template to ensure that the takeover safety alarm sound can be output stably.

[0070] S4.2 Determination of rhythm parameters After determining the spectral parameters, the system considers the risk of local audibility from the driver's seat. and alarm tone base frequency Determine the rhythm parameters.

[0071] Local audibility risk in the driver's seat When the pulse repetition rate increases, the system increases the pulse repetition rate. Or increase the number of times to repeat playback To improve the recognizability of the alarm tone in the time dimension. When the alarm tone fundamental frequency When shifting to a higher frequency band, the system synchronously shortens the envelope rise time. To enhance the transient characteristics of the alarm sound and compensate for the potential decrease in perceived intensity after the frequency band shifts upwards, the pulse repetition rate, single pulse duration, envelope rise time, and envelope fall time are all limited to a feasible range of preset rhythm parameters to avoid the alarm sound being too abrupt or sudden.

[0072] S4.3 Determination of loudness parameters After determining the spectral and tempo parameters, the system adjusts the overall background noise level based on the driver's seat. Risk of partial hearing impairment in the driver's seat Alarm tone frequency and envelope rising time The loudness parameters are determined by linkage.

[0073] When the background noise level Risk of increased or localized hearing loss in the driver's seat When the noise level increases, the system appropriately increases the loudness increment of the alarm sound relative to the background noise level. When the alarm tone's fundamental frequency is within a range sensitive to human hearing, the system can correspondingly reduce the loudness increment to avoid disturbing passengers by simply increasing the playback level. When the envelope rise time... When shortened to enhance transient recognition, the system correspondingly limits the initial transient increment. This is to avoid startling the driver due to sudden changes in volume.

[0074] Therefore, the spectral parameters are determined based on the frequency band masking distribution, the rhythm parameters are determined based on the spectral parameters and audibility risk, and the loudness parameters are determined based on the spectral parameters, rhythm parameters, and audibility risk linkage, thus forming a coupled and coordinated control relationship among the three types of parameters: loudness, spectral, and rhythm.

[0075] S5. Perform multi-objective constraint optimization on the target parameter set under preset constraints to obtain the optimal takeover prompt tone control parameters. In generating the target parameter set Then, the takeover prompt controller performs multi-objective constraint optimization on the target parameter set under preset safety constraints, comfort constraints, and prompt tone recognizability constraints to obtain the optimal alarm tone control parameters for the current audio frame.

[0076] The safety constraints include: the maximum playback level of the alarm tone does not exceed a preset safety limit; the response delay of the takeover warning does not exceed a preset time threshold; and the perceptibility of the alarm tone in the driver's seat area is not lower than a preset requirement. Comfort constraints include: the changes in loudness, fundamental frequency, and pulse repetition rate between adjacent audio frames do not exceed corresponding thresholds to avoid abrupt auditory changes caused by sudden parameter shifts. The perceptibility constraints include: the fundamental frequency, rhythm, and harmonic structure of the alarm tone remain within the preset design range for the takeover alarm tone, enabling the driver to recognize it as a takeover safety warning tone.

[0077] Let Ω be the feasible region of parameters that satisfy the above constraints, then the optimal alarm tone control parameters for the current frame are:

[0078] In the formula, Ω represents the optimal alarm tone control parameters for frame t; Ω represents the feasible region of parameters that satisfy safety constraints, comfort constraints, and identifiability constraints. Represents the current acoustic state vector inside the vehicle. Below is the candidate alarm tone parameter set. The takeover prompts the validity objective function.

[0079] The objective function for the takeover notification validity is defined as follows:

[0080] In the formula, The indicator represents the perceptibility of the alarm sound in the driver's seat area, which is related to the estimated signal-to-noise ratio in the driver's seat area and the degree of frequency band masking avoidance. The indicator of alarm tone recognizability is related to the similarity of the fundamental frequency, rhythm, and harmonic characteristics between the alarm tone and the preset standard takeover alarm tone. This indicates the annoyance level of the alarm sound, which is related to the alarm sound's volume, sharpness, and duration. This represents the parameter jump penalty term, which is related to the amount of change in the current frame parameter relative to the previous frame parameter; , , , These represent the weighting coefficients of the above indicators, and .

[0081] In this embodiment, the optimization solution can be achieved by combining a pre-set lookup table with local iterative optimization. Specifically, the system pre-establishes a parameter lookup table between the in-vehicle acoustic state vector and the alarm tone parameters based on typical in-vehicle sound field conditions. During operation, the takeover prompt controller determines the current in-vehicle acoustic state vector based on the alarm tone parameters. The initial parameters are obtained from the parameter lookup table, and then a finite number of local searches or gradient iterations are performed within the neighborhood of these initial parameters to obtain the optimal alarm tone control parameters that meet the real-time requirements of the vehicle. .

[0082] To avoid abrupt changes in alarm tone parameters between adjacent audio frames, the system smooths the optimal alarm tone control parameters, resulting in smoothed alarm tone control parameters:

[0083] In the formula, This represents the set of alarm tone control parameters after smoothing in frame t; This represents the set of optimal alarm tone control parameters obtained by optimization in frame t; This represents the set of alarm tone control parameters after smoothing the previous frame; Represents the smoothing coefficient. ,in, Indicates audio frame shift, This represents the time constant for parameter smoothing.

[0084] By using the optimal alarm tone control parameters for the current frame, the objective function for the effectiveness of takeover prompts, and the smoothed alarm tone control parameters, the system can coordinate between driver visibility, alarm tone recognizability, passenger comfort, and parameter smoothness, so that the takeover safety alarm tone can effectively wake up the driver without causing significant interference due to excessive loudness or frequent parameter jumps.

[0085] S6. Generate or match a takeover safety alert tone based on the optimal takeover alert tone control parameters and play it in real time. The takeover prompt controller controls parameters based on the smoothed alarm tone. It generates or matches the corresponding takeover safety alarm sound and plays it in real time through the vehicle's speaker system.

[0086] In one specific embodiment, the takeover alert controller matches the smoothed alarm tone control parameters from a preset parametric pulse alarm tone template library. The corresponding target template is then used, and based on the control parameters, the target template undergoes fundamental frequency adjustment, spectrum shaping, loudness shaping, and rhythm shaping.

[0087] In another specific embodiment, the takeover prompt controller controls parameters based on the smoothed alarm tone. The basic pulse alarm tone waveform is then parametrically synthesized in real time. This real-time parametric synthesis includes at least: based on the fundamental frequency... and bandwidth Determine the core frequency band of the alarm tone; based on the harmonic component weight set Determine the low-amplitude harmonic structure; based on the pulse repetition rate Single pulse duration envelope rising time and envelope falling edge time Determine the pulse rhythm; based on the absolute playback level. Relative background noise increment and initial transient increment Determine the playback volume.

[0088] When the takeover safety alarm tone is played, the takeover alert controller activates the vehicle's speaker system to play the alarm tone with the highest safety audio priority. During playback, the takeover alert controller can simultaneously perform volume ducking on non-safety audio streams such as in-vehicle music, navigation announcements, rear-seat entertainment audio, and conference call audio.

[0089] Let the playback level of the non-secure audio stream before ducking in frame t be... The playback level after dodging is ,but:

[0090] In the formula, This represents the playback level of the non-safe audio stream after ducking in frame t; This represents the playback level of the non-safe audio stream before the ducking in frame t; This indicates the volume ducking level.

[0091] By using volume ducking, the masking of takeover security alarm sounds by non-security audio streams can be reduced, while avoiding the superposition of distortion from multiple audio streams in the playback chain.

[0092] During the takeover of the safety alarm sound playback, the system continuously executes processes S2 to S6, reconstructing the in-vehicle acoustic state vector in real time as the dynamic sound field characteristics inside the vehicle are updated frame by frame. The system dynamically updates the alarm tone control parameters. Thus, when the in-vehicle music spectrum changes, passenger conversations increase or decrease, mobile terminal audio playback is turned on or off, or wind and road noise change with vehicle speed, the system can adjust the alarm tone spectrum, loudness, and rhythm accordingly, ensuring that the takeover safety alarm tone always matches the current in-vehicle sound field.

[0093] S7. Perform echo cancellation, audio stream management, and abnormal condition degradation handling. To ensure the authenticity of the sound field feature acquisition, in this embodiment, before extracting the dynamic sound field features inside the vehicle, the system preferably performs acoustic echo cancellation processing to eliminate the influence of alarm sounds played by the system itself on the microphone acquisition signal. The acoustic echo cancellation can employ a normalized least mean square adaptive filtering algorithm or other echo cancellation algorithms implementable by the vehicle audio system. The echo-cancelled audio signal is used for subsequent frequency band masking distribution and driver's seat local audibility risk calculation.

[0094] Meanwhile, the system monitors microphone connection status, audio frame integrity, controller computing load, speaker connection status, and in-vehicle audio playback link status in real time. When microphone failure, continuous audio frame loss, controller computing load exceeding threshold, speaker failure, or audio playback link abnormality is detected, the system triggers abnormal condition degradation processing.

[0095] Let the number of consecutive frame drops be The frame dropping threshold is The controller's computing power load is The computing power load threshold is The system will enter a downgrade warning state when any of the following conditions are met: or

[0096] In the degraded state, the takeover warning controller no longer relies on real-time microphone sound field feedback, but instead calls a pre-calibrated safety fallback alarm tone parameter scheme. This safety fallback scheme can determine the takeover safety alarm tone parameters based on the vehicle's current speed, the playback status of the in-vehicle audio system, and the pre-calibrated wind and road noise model.

[0097] The set of safety alarm tone parameters in the degraded state can be represented as:

[0098] In the formula, This represents the set of safety fallback alarm tone parameters under degraded status; This represents the pre-calibrated safety alarm tone parameter mapping function; Indicates the vehicle's current speed; Indicates the playback status of the vehicle's audio system; This represents the background noise level estimated based on vehicle speed, window status, and wind and road noise models.

[0099] Through the aforementioned degradation process, even if the real-time sound field monitoring link or the controller's real-time optimization link malfunctions, the system can still output a stable and clear takeover safety alarm tone, preventing the takeover request from failing to be effectively conveyed to the driver due to abnormal tone control.

[0100] Example 2 This embodiment is a specific implementation example in a multi-source mixed sound source superposition scenario, used to illustrate the application process of the method described in Embodiment 1 in an autonomous driving takeover prompt scenario on urban roads.

[0101] The vehicle is traveling at 50 km / h on a city road. The in-vehicle audio system is playing popular music, the front passenger is using a mobile device to play video and audio aloud, and the rear passengers are conversing. At this time, the vehicle is in autonomous driving mode. When the autonomous driving system detects that the traffic environment ahead, road boundary conditions, or driver takeover conditions meet the preset takeover trigger rules, it outputs takeover trigger information to the takeover prompt controller. Upon receiving the takeover trigger information, the takeover prompt controller initiates the full-cycle takeover prompt control process.

[0102] At the start of the takeover notification cycle, the takeover notification controller collects real-time audio signals from multiple microphones inside the vehicle. Combining this with a pre-calibrated cabin acoustic transfer function, the system converts the sound field information collected by the microphones into a local sound field estimation result for the driver's ear area. After frame processing, spectrum analysis, frequency band energy statistics, and voice activity detection, the system identifies that the current in-vehicle sound field is formed by the superposition of in-vehicle music, audio from the passenger's mobile terminal, conversations in the rear seats, and vehicle wind noise and road noise. The mid-frequency band from 300Hz to 2kHz has consistently high energy and is identified as a strong masking band. The overall A-weighted background noise level in the driver's seat area is 72dB(A), indicating a high level of in-vehicle voice activity. The sound field continuously changes with the music rhythm, voice volume, and content played from the mobile terminal.

[0103] The takeover warning controller constructs the in-vehicle acoustic state vector for frame t based on the aforementioned dynamic sound field characteristics. This in-vehicle acoustic state vector includes the overall A-weighted background noise level in the driver's seat area, frequency band masking distribution, in-vehicle voice activity intensity, local audibility risk in the driver's seat area, and the playback status of the in-vehicle audio system. Calculations show that the current in-vehicle sound field forms a continuous strong masking region within the range of 300Hz to 2kHz. If the takeover safety warning sound still uses a conventional fixed mid-frequency fundamental frequency, it is prone to frequency band overlap with music, human voices, and external audio, increasing the risk of the driver missing it. In this embodiment, the system calculates the local audibility risk in the driver's seat area to be 0.68, indicating that the current takeover warning sound has a high risk of being masked in the driver's seat area, requiring adaptive adjustment of the takeover safety warning sound.

[0104] The takeover alert controller generates a target parameter set for the takeover safety alarm tone based on the frequency band masking distribution and the local audibility risk in the driver's seat. First, in terms of frequency spectrum parameters, the system avoids the strong masking frequency band from 300Hz to 2kHz, adjusts the fundamental frequency of the pulsed takeover safety alarm tone to 2800Hz, and narrows the alarm audio bandwidth to 400Hz, so that the core energy of the alarm tone is concentrated in the frequency band with lower current masking intensity, thereby reducing the frequency band overlap between the alarm tone and popular music, human voices, and external audio from mobile terminals.

[0105] After determining the spectral parameters, the system further determines the rhythm parameters. Since the local audibility risk from the driver's seat is 0.68, the system sets the pulse repetition rate to 3Hz, the duration of a single pulse to 150ms, and the envelope rise time to 20ms to enhance the transient recognition of the takeover safety alarm tone. By increasing the pulse repetition rate and shortening the envelope rise time, even with a certain change in perceived loudness after the alarm tone's fundamental frequency shifts to 2800Hz, the system can still improve the driver's probability of perceiving the warning tone through the obvious pulse characteristics in the time dimension.

[0106] Regarding loudness parameters, the system, based on a background noise level of 72 dB(A) in the driver's seat area, a local audibility risk of 0.68 in the driver's seat area, a fundamental frequency of 2800 Hz for the alarm tone, and an envelope rise time of 20 ms, determines that the loudness increment of the takeover safety alarm tone relative to the background noise level is 8 dB, and the initial transient increment is 5 dB. Since the core frequency band of the alarm tone has already avoided major strong masking frequency bands, the system does not need to simply increase the overall playback level significantly to improve the perceptibility of the takeover safety alarm tone in the driver's seat area; at the same time, by limiting the initial transient increment, the system avoids the alarm tone from starting too abruptly and causing startling to the driver or significant discomfort to the passengers.

[0107] Subsequently, under safety, comfort, and identifiability constraints, the takeover alert controller performs multi-objective constraint optimization on the aforementioned spectral, rhythm, and loudness parameters, and smooths the parameter changes between adjacent audio frames to obtain the smoothed alarm tone control parameters for the current frame. Based on these smoothed alarm tone control parameters, the takeover alert controller matches a corresponding template from a parametric pulse alarm tone template library, or performs real-time parametric synthesis of the basic pulse alarm tone waveform to generate a takeover safety alarm tone with a fundamental frequency of 2800Hz, a bandwidth of 400Hz, a pulse repetition rate of 3Hz, a single pulse duration of 150ms, an envelope rise time of 20ms, a loudness increment relative to the background noise level of 8dB, and an initial transient increment of 5dB. This alarm tone is then played through the vehicle's speaker system with safety alert priority.

[0108] While playing the takeover safety alarm tone, the takeover alert controller performs volume ducking on non-safety audio streams. In this embodiment, the playback level of the in-vehicle music is reduced by 18dB to reduce the masking of the takeover safety alarm tone by the in-vehicle music; for audio played from the passenger's mobile terminal and rear-seat entertainment audio, the system can perform corresponding volume reduction, channel suppression, or alert priority enhancement processing based on the controllable state of the in-vehicle audio link. Through the above audio stream management, the takeover safety alarm tone can form a clearer auditory alert in the driver's seat area without relying on excessively high overall loudness.

[0109] Throughout the entire takeover alert cycle, the takeover alert controller continuously collects real-time audio signals from inside the vehicle and updates the in-vehicle acoustic state vector frame by frame. When the music on the in-vehicle system switches from pop music with strong vocals and melodies to music dominated by low-frequency percussion, the system detects that the strong masking frequency band gradually shifts down from the original 300Hz to 2kHz range to below 1kHz. At this point, although the original 2800Hz alarm tone fundamental frequency still avoids the main low-frequency masking area, considering the overall balance between the risk of local audibility in the driver's seat, the recognizability of the alarm tone, and the comfort of the occupants, the system smoothly adjusts the alarm tone pulse fundamental frequency from 2800Hz to 2200Hz, while maintaining the core energy of the alarm tone away from the current strong masking frequency band below 1kHz. This adjustment process is completed through parameter smoothing to avoid significant frequency jumps in the alarm tone during the takeover alert process.

[0110] When the front passenger stops playing video and audio from their mobile device, the number of mixed sound sources inside the vehicle decreases, the background noise level in the driver's area drops, and the masking intensity of mid-to-high frequencies in the frequency band masking distribution decreases simultaneously. Based on the updated acoustic state vector, the takeover alert controller recalculates the local audibility risk in the driver's area and, while ensuring safety and detectability, smoothly adjusts the loudness parameters of the alarm tone. For example, it reduces the relative background noise increment or the initial transient increment to avoid maintaining an excessively high playback level after the sound field interference weakens. Thus, the system can continuously adjust the takeover safety alarm tone according to changes in the in-vehicle sound field, maintaining high detectability and recognizability throughout the entire takeover alert cycle while minimizing excessive interference to the occupants.

[0111] As can be seen from this embodiment, in the scenario of multiple mixed sound sources superimposed on urban roads, this application does not simply increase the volume of the takeover warning tone based on the background noise level. Instead, it first identifies strong masking frequency bands and audibility risks in the local sound field of the driver's seat, and then couples and coordinates the spectral parameters, rhythm parameters, and loudness parameters of the takeover safety alarm tone, and combines audio stream ducking and smoothing processing to achieve real-time playback. Therefore, it can improve the driver's ability to promptly perceive the autonomous driving takeover warning tone in a complex sound field where in-vehicle music, mobile terminal speakers, human conversations, and wind and road noise coexist, while taking into account the safety, recognizability, and ride comfort of the warning tone.

[0112] Example 3 An autonomous driving takeover prompt tone control system based on sound field monitoring is provided to implement the method described in Embodiment 1. The system can be integrated into the vehicle's autonomous driving domain controller, intelligent cockpit domain controller, or independent audio safety controller, and communicates with the vehicle's autonomous driving system, in-vehicle microphone, in-vehicle audio system, and in-vehicle speaker system.

[0113] like Figure 2 As shown, the system includes a takeover event acquisition module, an in-vehicle sound field monitoring module, an acoustic state construction module, a prompt tone parameter decision module, a constraint optimization module, a prompt tone generation and playback module, and an anomaly degradation module.

[0114] The takeover event acquisition module is used to acquire the takeover trigger information output by the vehicle's autonomous driving system, and initiates the full-cycle control process of takeover prompt upon receiving the takeover trigger information.

[0115] The in-vehicle sound field monitoring module is used to collect real-time audio signals within the vehicle throughout the entire period of the takeover prompt and extract dynamic sound field characteristics. These dynamic sound field characteristics include one or more of the following: background noise level in the driver's seat area, cross-frequency sound pressure level, voice activity status, local sound field status in the driver's seat area, and playback status of the in-vehicle audio system. Specifically, the in-vehicle sound field monitoring module calls a pre-calibrated cabin acoustic transfer function to map the sound field information collected by the microphone to the driver's ear area and performs echo cancellation processing during the prompt playback.

[0116] The acoustic state construction module is used to construct an in-vehicle acoustic state vector based on the in-vehicle dynamic sound field characteristics. The in-vehicle acoustic state vector includes at least a frequency band masking distribution and a driver's seat local audibility risk; wherein, the frequency band masking distribution is used to characterize the degree of masking of the takeover safety alarm sound by different frequency bands, and the driver's seat local audibility risk is used to characterize the risk that the takeover safety alarm sound will be detected in the driver's seat area in a timely manner.

[0117] The prompt tone parameter decision module is used to generate the spectral parameters, rhythm parameters, and loudness parameters of the takeover safety alarm tone based on the in-vehicle acoustic state vector. The spectral parameters are used to ensure that the core energy of the alarm tone avoids strong masking frequency bands, the rhythm parameters are used to adjust the pulse repetition rate, duration, and envelope characteristics of the alarm tone, and the loudness parameters are used to adjust the playback level, loudness increment, and initial transient increment of the alarm tone.

[0118] The constraint optimization module is used to optimize the spectrum parameters, rhythm parameters, and loudness parameters under safety constraints, comfort constraints, and identifiability constraints to obtain the optimal takeover prompt tone control parameters, and to smooth the control parameters between adjacent audio frames.

[0119] The alert tone generation and playback module is used to generate or match a takeover safety alarm tone based on the smoothed takeover alert tone control parameters, and to control the in-vehicle speaker system to play the takeover safety alarm tone with safety alert priority. Furthermore, the alert tone generation and playback module is also used to perform volume ducking processing on non-safety audio streams such as in-vehicle music, navigation announcements, rear-seat entertainment audio, or conference call audio when playing the takeover safety alarm tone.

[0120] The anomaly degradation module is used to detect microphone failure, continuous audio frame loss, controller computing power load exceeding the threshold, speaker failure, or audio playback link anomaly. When an anomaly is detected, it calls the pre-calibrated safety fallback alarm tone parameter scheme, so that the system can still output takeover safety prompt tone when real-time sound field monitoring or real-time optimization is unavailable.

[0121] During system operation, after the takeover event acquisition module receives the takeover trigger information, the in-vehicle sound field monitoring module continuously acquires the dynamic sound field characteristics in the vehicle, the acoustic state construction module forms the in-vehicle acoustic state vector, the prompt tone parameter decision module generates spectrum, rhythm and loudness parameters, the constraint optimization module obtains the smoothed takeover prompt tone control parameters, and the prompt tone generation and playback module generates or matches the takeover safety alarm tone and plays it.

[0122] The above process is executed cyclically throughout the entire takeover prompt cycle until the driver completes the takeover, the takeover prompt ends, or the system enters a degrade prompt state.

[0123] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for controlling the automatic driving takeover prompt tone based on sound field monitoring, characterized in that, Includes the following steps: Acquire and respond to takeover trigger information output by the vehicle's autonomous driving system, collect real-time audio signals inside the vehicle throughout the entire takeover prompt cycle, and extract dynamic sound field features inside the vehicle that are updated frame by frame. Based on the in-vehicle dynamic sound field characteristics, an in-vehicle acoustic state vector is constructed. The in-vehicle acoustic state vector includes at least the frequency band masking distribution and the risk of local audibility in the driver's seat. The spectral parameters, rhythm parameters, and loudness parameters of the takeover safety alarm sound are generated based on the in-vehicle acoustic state vector. The spectral parameters are determined according to the frequency band masking distribution so that the core energy of the takeover safety alarm sound avoids strong masking frequency bands. The rhythm parameters and loudness parameters are determined in conjunction with the spectral parameters and the local audibility risk of the driver's seat. Under preset constraints, the spectrum parameters, rhythm parameters, and loudness parameters are optimized by multi-objective constraint coupling to obtain the optimal alarm tone control parameters; The takeover safety alarm tone is generated or matched according to the optimal alarm tone control parameters and played, and the playback parameters of the takeover safety alarm tone are dynamically adjusted according to the update of the in-vehicle dynamic sound field characteristics throughout the entire takeover prompt cycle. The frequency band masking distribution is determined as follows: If the frequency band sensitive to human hearing is divided into B frequency bands, then the masking threshold corresponding to the b-th frequency band in the t-th frame is: In the formula, This represents the masking threshold for the b-th frequency band in the t-th frame; This represents the local sound pressure level in the driver's ear region of the t-th frame within the b-th frequency band; This represents the pre-calibrated frequency band masking correction coefficient corresponding to the b-th frequency band. Normalize the masking threshold of each frequency band to obtain the normalized masking strength of the b-th frequency band: In the formula, This represents the normalized masking strength of the b-th frequency band in the t-th frame; This represents the minimum masking threshold among all frequency bands in frame t; ε represents the maximum masking threshold in all frequency bands of the t-th frame, and ε represents the smallest positive number that prevents the denominator from being zero. The frequency band masking distribution vector is composed of the normalized masking intensity of each frequency band: In the formula, This represents the frequency band masking distribution of the t-th frame; The risk of partial hearing in the driver's seat was determined as follows: Let the normalized energy percentage of the current candidate takeover safety alarm tone in the b-th frequency band be . And satisfy: The overlap between the takeover security alarm audio band and the current frequency band masking distribution is: In the formula, This represents the normalized masking strength of the b-th frequency band; Let the predicted A-weighted sound pressure level of the current candidate takeover safety alarm sound in the driver's ear region be . The overall A-weighted background noise level in the driver's seat area is [value missing]. The signal-to-noise ratio for the driver's seat local estimation is: The risk of localized audibility in the driver's seat is: In the formula, This indicates the risk of local audibility from the driver's seat in frame t. This indicates the preset signal-to-noise ratio threshold. This represents the intensity of in-vehicle voice activity in frame t. , , represent the weighting coefficients of insufficient signal-to-noise ratio, bandwidth masking overlap, and speech activity interference on audibility risk, respectively. .

2. The method for controlling the automatic driving takeover prompt tone based on sound field monitoring according to claim 1, characterized in that, The in-vehicle dynamic sound field characteristics include the overall A-weighted background noise level, frequency band sound pressure level, short-time spectrum energy distribution, voice activity detection results, energy proportion of the main interference sound source, local sound pressure level in the driver's seat area, and the playback status of the in-vehicle audio system. The local sound pressure level in the driver's seat is estimated by combining the signals collected by the in-vehicle microphones with a pre-calibrated cabin acoustic transfer function. The cabin acoustic transfer function is used to characterize the acoustic transfer relationship between the location of the in-vehicle microphones and the ear area of ​​the driver's seat.

3. The method for controlling the automatic driving takeover prompt tone based on sound field monitoring according to claim 1, characterized in that, The in-vehicle acoustic state vector is represented as follows: In the formula, This represents the acoustic state vector inside the vehicle in frame t. This indicates the overall A-weighted background noise level in the driver's seat area; Represents the frequency band masking distribution vector; Indicates the intensity of voice activity inside the vehicle; This indicates a risk of localized hearing impairment in the driver's seat. Indicates the playback status of the vehicle's audio system; The playback status of the vehicle audio system includes at least one of the following: vehicle audio playback identifier, audio type, current volume level, spectrum status of vehicle music or navigation broadcast, audio status of telephone conference, audio status of rear-seat entertainment, and audio status of passenger mobile terminal external speaker.

4. The method for controlling the automatic driving takeover prompt tone based on sound field monitoring according to claim 1, characterized in that, The spectral parameters, rhythm parameters, and loudness parameters constitute the target parameter set: In the formula, This represents the set of target parameters for taking over the security alarm sound in frame t; Represents the set of loudness parameters; Represents a set of spectral parameters; Represents the set of rhythm parameters; The loudness parameter set includes the absolute playback level of the alarm tone, the loudness increment relative to the background noise level, and the initial transient increment; the spectrum parameter set includes the fundamental frequency of the alarm tone, the bandwidth, and the set of harmonic component weights; the rhythm parameter set includes the pulse repetition rate, the duration of a single pulse, the envelope rising edge time, the envelope falling edge time, and the number of repetitions.

5. The method for controlling the automatic driving takeover prompt tone based on sound field monitoring according to claim 4, characterized in that, The set of spectral parameters, the set of rhythm parameters, and the set of loudness parameters are determined as follows: When determining the set of spectral parameters, the frequency sub-bands in the frequency band masking distribution with normalized masking intensity lower than the strong masking judgment threshold and center frequency within the allowable range of the alarm tone fundamental frequency are identified as candidate frequency bands. The frequency sub-band with the lowest normalized masking intensity or two adjacent frequency sub-bands are selected from the candidate frequency bands as the target frequency band. The center frequency of the target frequency band is determined as the alarm tone fundamental frequency. The alarm tone bandwidth and harmonic component weight set are determined based on the target frequency band. When determining the set of rhythm parameters, the pulse repetition rate and the number of repetitions are determined based on the local audibility risk of the driver's seat. When the local audibility risk of the driver's seat increases, the pulse repetition rate is increased or the number of repetitions is increased. The envelope rise time is determined based on the fundamental frequency of the alarm tone. When the fundamental frequency of the alarm tone is higher than a preset high-frequency judgment threshold, the envelope rise time is shortened. When determining the set of loudness parameters, the loudness increment of the alarm sound relative to the background noise level is determined based on the overall A-weighted background noise level of the driver's seat area and the local audibility risk of the driver's seat, and the absolute playback level of the alarm sound is determined based on the loudness increment; when the fundamental frequency of the alarm sound is in the frequency band sensitive to human hearing, the loudness increment is limited; when the envelope rise time is less than a preset transient judgment threshold, the initial transient increment is limited.

6. The method for controlling the automatic driving takeover prompt tone based on sound field monitoring according to claim 1, characterized in that, The multi-objective constraint coupling optimization is performed as follows: Let Ω be the feasible region of parameters satisfying safety, comfort, and identifiability constraints. The optimal alarm tone control parameters for the current frame are: In the formula, This represents the optimal alarm tone control parameters for frame t. Represents the current acoustic state vector inside the vehicle. Below is the candidate alarm tone parameter set. The takeover prompts the effectiveness objective function; The objective function for the validity of the takeover notification is: In the formula, This indicates the visibility of the alarm sound in the driver's seat area. This indicates the distinguishability of the alarm tone. This indicates the level of annoyance of the alarm sound. This indicates a penalty term for parameter jumps. , , , These represent the weight coefficients of the corresponding indicators, and ; Furthermore, the optimal alarm tone control parameters are smoothed to obtain smoothed alarm tone control parameters: In the formula, This represents the set of alarm tone control parameters after smoothing in frame t; This represents the set of optimal alarm tone control parameters obtained by optimization in frame t; This represents the set of alarm tone control parameters after smoothing the previous frame; Represents the smoothing coefficient. ,in, Indicates audio frame shift, This represents the time constant for parameter smoothing.

7. The method for controlling the automatic driving takeover prompt tone based on sound field monitoring according to claim 1, characterized in that, During the process of acquiring real-time audio signals inside the vehicle and extracting dynamic sound field features inside the vehicle, abnormal operating condition detection and degradation processing are performed simultaneously. The abnormal operating condition detection includes detecting at least one of the following: microphone connection status, audio frame integrity, controller computing load, speaker connection status, and vehicle audio playback link status. When microphone failure, continuous audio frame loss, controller computing power load exceeding threshold, speaker failure, or audio playback link abnormality is detected, the parameter optimization process based on real-time in-vehicle dynamic sound field characteristics is stopped or skipped, and the pre-calibrated safety fallback alarm tone parameter scheme is called to generate the takeover safety alarm tone. The safety fallback alarm tone parameter scheme is determined based on one or more of the vehicle's current driving speed, the playback status of the in-vehicle audio system, and the wind noise and road noise pre-calibration model, and the corresponding downgrade alarm tone parameters are input into the generation or playback step of the takeover safety alarm tone.

8. An automated driving takeover prompt tone control system based on sound field monitoring, characterized in that, The method for controlling the automatic driving takeover prompt tone based on sound field monitoring as described in any one of claims 1-7 includes: The takeover event acquisition module is used to acquire takeover trigger information output by the vehicle's autonomous driving system. The in-vehicle sound field monitoring module is used to respond to the takeover trigger information, collect real-time audio signals in the vehicle during the entire takeover prompt cycle, and extract dynamic sound field features in the vehicle that are updated frame by frame. An acoustic state construction module is communicatively connected to the in-vehicle sound field monitoring module and is used to construct the in-vehicle acoustic state based on the in-vehicle dynamic sound field characteristics. The in-vehicle acoustic state includes at least the frequency band masking distribution and the risk of local audibility in the driver's seat. The prompt tone parameter decision module is communicatively connected to the acoustic state construction module and is used to generate the spectral parameters, rhythm parameters, and loudness parameters of the takeover safety alarm tone based on the in-vehicle acoustic state. The constraint optimization module is communicatively connected to the parameter decision module and is used to perform multi-objective constraint coupling optimization on the spectrum parameters, rhythm parameters and loudness parameters under preset constraints to obtain the optimal alarm sound control parameters. The prompt tone generation and playback module is communicatively connected to the constraint optimization module. It is used to generate or match the takeover safety alarm tone according to the optimal alarm tone control parameters, and control the in-vehicle speaker system to play the takeover safety alarm tone. The in-vehicle sound field monitoring module, acoustic state construction module, prompt tone parameter decision module, constraint optimization module, and prompt tone generation and playback module are executed cyclically throughout the entire takeover prompt cycle as the dynamic sound field characteristics of the in-vehicle are updated frame by frame, so as to dynamically adjust the playback parameters of the takeover safety alarm tone.

Citation Information

Patent Citations

  • Audio processing method, electronic equipment and storage medium

    CN119946506A

  • Vehicle warning tone dynamic adjustment method and system based on multi-dimensional perception

    CN120564711A