An adaptive audio coding method and system for autonomous driving state auditory cues
By using an adaptive audio coding method to dynamically generate auditory cues using global and local audio feature parameters, the problem of insufficient and redundant auditory interaction information in existing technologies is solved. This enables clear communication of the autonomous driving system's status and improves the driver's situational awareness, thereby enhancing the system's interpretability and safety.
Patent Information
- Application Number
- CN202610030267.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-12
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2046-01-12
AI Technical Summary
Existing human-machine interaction technologies for autonomous vehicles suffer from limitations in conveying system operating status and decision-making information through auditory means. These limitations include limited information volume, inability to maintain clarity and continuity of information, inability to maintain consistency with the decision-making process of the autonomous driving system, and inability to serve as an alternative interaction method when the driver's visual passage is occupied.
An adaptive audio coding method is adopted to collect internal operating status and environmental target data of the autonomous driving system in the vehicle, filter decision-related targets, construct global audio features and local audio attribute parameters, and generate dynamic auditory cues. This ensures that the auditory cues are consistent with the decision-making behavior of the autonomous driving system, reduces redundancy, and improves information clarity and continuity.
This technology enables the clear communication of the autonomous driving system's state changes and decision-making processes through auditory means without increasing the driver's visual load. This enhances the driver's situational awareness, improves the system's interpretability and trustworthiness, reduces interference and abrupt changes in auditory cues, and strengthens the driver's safety and efficiency in takeover scenarios.
Smart Images

Figure CN121516022B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of intelligent driving, and particularly relates to an automatic driving state auditory prompting method and system based on adaptive audio coding. BACKGROUND
[0002] With the continuous improvement of vehicle automation and intelligence level, automatic driving systems are gradually applied to various driving scenarios to complete functions such as environment perception, decision planning, and vehicle control. In the automatic driving mode, the driver no longer directly controls the vehicle continuously, but mainly undertakes the task of supervising the running state of the automatic driving system and taking over the vehicle when necessary. Under this background, how the automatic driving system conveys its running state and external environment information to the driver has become one of the important research directions in the field of vehicle human-machine interaction.
[0003] Currently, the prompting method of the automatic driving vehicle for the running state and the environment information mainly relies on visual interaction, such as displaying the automatic driving mode, surrounding targets, or risk prompt information through the instrument panel, central control screen, or head-up display device. The above prompting method essentially realizes information transmission by occupying the visual perception ability of the driver. However, in the actual driving process, the visual perception ability of the driver often needs to be used to observe the road environment or perform visual tasks unrelated to driving. When the visual perception ability is occupied or dispersed, the driver will have difficulty in obtaining the running state or decision-related information presented by the automatic driving system in a timely manner, thereby causing interruption or delay in information acquisition. In the case where the automatic driving system needs to be taken over by the driver, the system state changes, or there is a potential risk, if the driver cannot continuously perceive the system running state through visual means, the safety and reliability of human-machine cooperation will be directly affected. Therefore, relying only on visual means as the state prompting method of the automatic driving system cannot meet the requirements of safety and redundancy in complex driving scenarios.
[0004] Based on the above problems, it is necessary to introduce an information acquisition channel different from the visual method, so that the driver can continuously obtain the running state of the automatic driving system and its decision-related information without relying on visual perception ability. By distributing the automatic driving system state information to different information transmission channels, the continuity of information transmission can be maintained when a single channel is limited, thereby improving the safety and reliability of human-machine interaction of the automatic driving system.
[0005] Hearing as an independent interaction mode from vision, is not easily affected by activities that occupy visual attention, and can continue to play the role of information prompt without the driver needing to shift his gaze, and is suitable for reminding the driver to pay attention to changes in vehicle operating state or potential risk information. Existing auditory prompts mainly include voice and non-voice. Voice prompts use voice broadcasting as an interaction method, which requires the driver to listen to the complete sentence broadcast when expressing complex information, has a slower trigger time, and causes higher cognitive load and annoyance. Therefore, short voice broadcasts, although fast, contain less information. Therefore, voice prompts are not conducive to quickly restoring the driver's situational awareness in automatic driving takeover scenarios, and are usually only used to arouse the driver's attention in takeover scenarios. Non-voice prompts are triggered more quickly and have the ability to convey complex information, making them more suitable as real-time state prompting methods for auditory interaction systems in autonomous vehicles.
[0006] Existing technologies related to non-voice prompts in vehicle cabins include:
[0007] (1) Chinese invention patent with publication number CN111308999A discloses an information providing device and a vehicle-mounted device, which uses an automatic driving guidance technology based on sound effects.
[0008] Defects and deficiencies: This technology uses simple, unstructured "sound effects" (such as "ding ding" sound) to prompt the actions of the autonomous driving system (such as lane changing), the main purpose is to distinguish the direction of travel of the autonomous driving system from the voice direction prompt of the navigation. Its main defect is that the information dimension is single, only the "rough direction of travel of the autonomous driving system" can be expressed, and the complex internal state of the system such as the current real-time "decision state" (accelerate to overtake the front vehicle), "decision intent" (why to change lanes), "system confidence" (how much confidence) and the precise spatial distribution of related external targets cannot be conveyed. This prompt sound has low information content, cannot make the driving operation of the autonomous driving system transparent, and cannot selectively prompt information, therefore, cannot effectively establish the driver's deep trust in the system, and cannot provide more effective situational awareness for the driver of the autonomous driving vehicle in the takeover scenario.
[0009] (2) Chinese invention patent with publication number CN120482082A discloses a vehicle audio control method and device, vehicle, and storage medium, which uses an environmental prompting technology based on simulated audio.
[0010] Defects and shortcomings: this technology prompts external targets by simulating real-world sounds (such as using horse hooves to represent cars). The main problem is that this technology is not for autonomous vehicles, the simulated sound has no relevance to the internal decision-making logic of the autonomous driving system, and is more used to create a unique sound landscape inside the cabin. The driver cannot understand what the system "thinks" through sound. Simulating multiple real sounds can cause a mix of sounds in the cabin, increase cognitive load, and may cause ambiguity in different cultural backgrounds.
[0011] In summary, the existing human-computer interaction technology of autonomous vehicles still has some shortcomings in conveying system running status and decision-related information through auditory means. The amount of information provided is small, the system state that can be reflected is limited, it cannot be used as an alternative interaction method when the driver's visual channel is occupied, and there is still a lack of an auditory prompt mechanism that can ensure information clarity and continuity while maintaining consistency with the decision-making process of the autonomous driving system. SUMMARY
[0012] In view of the above problems in the prior art, the purpose of the present application is to provide an adaptive audio coding autonomous driving state auditory prompt method, which converts the running state of the autonomous driving system and its decision basis into real-time and interpretable auditory prompts, thereby improving the driver's situational awareness of the autonomous driving system without increasing the driver's visual load.
[0013] An adaptive audio coding autonomous driving state auditory prompt method applied to a vehicle with autonomous driving function, comprising the following steps:
[0014] During vehicle operation, internal running state data of the autonomous driving system and environmental target data are collected at a preset update period. The internal running state data at least includes the current autonomous driving mode level, the current decision behavior type, the decision confidence, and the system running stability indication information. The environmental target data includes the categories, relative positions, relative motion states, collision times, and risk assessment indicators of multiple detected targets.
[0015] Based on the decision result of the current autonomous driving system, the environmental target data is filtered, and only the targets that directly affect the current decision process are retained as decision-related targets, and the remaining targets do not generate corresponding auditory prompt information.
[0016] The internal running state data is constructed as internal running state time sequence input and time sequence modeling is performed to generate a global state representation for describing the overall running situation and change trend of the autonomous driving system, and based on the global state representation, global audio feature parameters reflecting the auditory background are generated.
[0017] The target state data corresponding to the decision-related target is constructed as a target time sequence input, and a corresponding local audio attribute parameter is generated based on a target importance evaluation mechanism.
[0018] The global audio feature parameter and the local audio attribute parameter are uniformly scheduled, so that the global audio feature parameter is output as a persistent auditory background, and the local audio attribute parameter is dynamically output as an auditory foreground superimposed on the auditory background.
[0019] The scheduled global audio feature parameter and local audio attribute parameter are converted into audio control instructions and audio synthesis and playback are performed to form auditory cues that dynamically change with the running state and decision-making process of the autonomous driving system.
[0020] Preferably, the selection of the decision-related target is based on a set of decision-making targets output by the planning module of the autonomous driving system, or is determined based on a comprehensive target risk evaluation index, lane correlation relationship and target motion trend.
[0021] Preferably, the global audio feature parameter at least includes control parameters for controlling the overall rhythm of the auditory cue, the degree of harmonic stability, and the timbre tension or brightness.
[0022] Preferably, the local audio attribute parameter at least includes parameters for controlling the volume intensity, pitch variation range and timbre modification degree of the target voice part.
[0023] Preferably, the target importance evaluation mechanism includes an attention mechanism that assigns auditory weights to different decision-related targets based on a target risk evaluation index, and the auditory weights are positively correlated with the corresponding local audio attribute parameters in terms of auditory saliency.
[0024] Preferably, the local audio attribute parameter further includes a spatial sound image control parameter, which is set according to the spatial orientation of the decision-related target relative to the vehicle, so that the auditory cue has a direction indicating property.
[0025] Preferably, different update strategies are adopted for the global audio feature parameter and the local audio attribute parameter, wherein the update frequency of the global audio feature parameter is lower than that of the local audio attribute parameter.
[0026] Preferably, before the audio control instructions are subjected to audio synthesis and playback, the generated audio events are time-scheduled and written into an event buffer to absorb scheduling jitter and ensure the real-time performance and stability of audio output.
[0027] Preferably, during the audio synthesis and playback process, the auditory cue is mixed with other audio signals in the vehicle, and the auditory cue is set with a priority or a loudness limit.
[0028] Another object of the present application is to provide an automatic driving state auditory cue system for adaptive audio coding, comprising:
[0029] a data acquisition module configured to acquire internal operation state data and environmental target data of the automatic driving system at a preset update period during vehicle operation, wherein the internal operation state data at least includes a current automatic driving mode level, a current decision behavior type, a decision confidence, and system operation stability indication information, and the environmental target data includes categories, relative positions, relative motion states, collision times, and risk assessment indicators of a plurality of detected targets;
[0030] a decision-related target screening module configured to screen the environmental target data based on a current decision result of the automatic driving system, and only keep targets that have a direct impact on the current decision process as decision-related targets;
[0031] a global audio feature generation module configured to construct the internal operation state data into internal operation state time sequence input and perform time sequence modeling to generate a global state representation for describing overall operation situation and change trend of the automatic driving system, and generate global audio feature parameters reflecting an auditory background based on the global state representation;
[0032] a local audio attribute generation module configured to construct target state data corresponding to the decision-related targets into target time sequence input, and generate corresponding local audio attribute parameters based on a target importance evaluation mechanism, wherein the local audio attribute parameters are used to reflect the saliency of each decision-related target in the auditory sense;
[0033] a parameter scheduling module configured to uniformly schedule the global audio feature parameters and the local audio attribute parameters, so that the global audio feature parameters are output as a persistent auditory background, and the local audio attribute parameters are output as an auditory foreground dynamically superimposed on the auditory background;
[0034] an audio synthesis and output module configured to convert the scheduled global audio feature parameters and local audio attribute parameters into audio control instructions, and perform audio synthesis and playback to form an auditory cue dynamically changing with the operation state and decision process of the automatic driving system.
[0035] The self-adaptive audio coding automatic driving state auditory prompt method and system can convey the running state of the automatic driving system and its decision basis to the driver in a continuous and intuitive auditory manner, thereby improving the situational awareness of the driver to the automatic driving system without increasing the visual load. By structurally processing the internal running state of the automatic driving system and the environmental perception information, and appropriately encoding the auditory prompt as an information output channel, the driver can perceive the system state change through auditory perception in a complex driving scene, and can maintain a certain situational awareness even if the visual channel is occupied by other factors such as irrelevant driving behaviors, thereby improving the problem of limited information acquisition efficiency caused by excessive reliance on visual prompts in the prior art.
[0036] By filtering decision-related targets based on the decision results of the automatic driving system, only the key targets involved in the current decision process are generated to generate auditory prompts, and the environmental targets not involved in the decision are kept silent, thereby fundamentally reducing auditory information redundancy and avoiding the interference problem caused by frequent triggering or multi-target superposition in the prior art. The mechanism makes the auditory prompt consistent with the real decision behavior of the automatic driving system, improves the explainability of the automatic driving system, helps the driver to understand why the system makes a specific decision, and thereby improves the predictability and trustworthiness of the behavior of the automatic driving system.
[0037] By constructing a hierarchical auditory coding structure in which global audio features are separated from local audio attributes, the overall running situation of the automatic driving system is mapped to a continuously changing auditory background, and the decision-related targets are mapped to an auditory foreground superimposed on the background, thereby realizing the unity of continuous expression of system state and highlighting of key causal targets. Compared with the prior art which uses a single prompt tone or fixed sound effect, the hierarchical structure can clearly express the trend of system state change and the decision causal relationship while ensuring the stability and comfort of the auditory prompt.
[0038] The recurrent neural network or long short-term memory network is used to model the internal state of the automatic driving system in time sequence, so that the auditory prompt can reflect the trend characteristics of the change of the system state with time, rather than only responding to the instantaneous state, thereby avoiding the mutation and jump problem of the auditory prompt. By imposing smoothing and change rate limitation on the audio parameter update process, the auditory prompt presents a natural gradual change effect in the time dimension, thereby further improving the acceptance and use comfort of the driver.
[0039] By evaluating the importance of the decision-related targets and mapping the evaluation results to the saliency differences of the local audio attributes, the prominence of different targets in the hearing corresponds to the degree of influence of the current decision, so that the driver can quickly distinguish the key targets from the secondary targets through hearing. Combined with the spatial sound image mapping technology, the auditory cues can also reflect the spatial orientation of the target relative to the vehicle, which helps the driver to understand the key risk sources in the surrounding environment without looking at the display interface, thereby improving the performance of the driver when taking over. BRIEF DESCRIPTION OF DRAWINGS
[0040] The accompanying drawings are included to provide a further understanding of the application, and constitute a part of the specification, illustrate the application, and are included to explain the application, and do not constitute a limitation on the application. In the drawings:
[0041] Figure 1 is a system block diagram of the application. DETAILED DESCRIPTION
[0042] Embodiment one
[0043] As shown in Figure 1 An adaptive audio coding method for hearing prompt of autonomous driving state, which is run in a vehicle with autonomous driving function, is used to convert the running state of the autonomous driving system and its decision basis into real-time and interpretable hearing prompts, so as to improve the situational awareness ability of the autonomous driving system without increasing the visual load of the driver, and thus improve the trust of the driver in the autonomous driving system and the performance in the takeover scenario. "Takeover performance" means: under certain conditions, when the vehicle system issues a takeover request, the efficiency and quality of the driver taking over the control of the vehicle from the autonomous driving system.
[0044] The adaptive audio coding method for hearing prompt of autonomous driving state, on the one hand, improves the efficiency and quality of the driver when taking over, thereby improving the safety of the autonomous driving vehicle, preventing the driver from being unable to monitor the state of the autonomous driving system in time due to the occupation of the visual channel, and causing danger because the driver cannot quickly recover the situational awareness in the scenario where the driver needs to take over. On the other hand, by transmitting the internal decision state of the autonomous driving system to the driver in real time, the decision transparency of the autonomous driving system is improved, thereby improving the trust of the driver in the autonomous driving system and the acceptance of the autonomous driving vehicle.
[0045] The specific process of the method is as follows:
[0046] During the vehicle operation, the system performs the auditory cue generation process with a fixed update period, which can be set to 20-50 milliseconds according to the real-time requirements of the vehicle control system. At the beginning of each update period, the system first obtains the internal operating state data from the autonomous driving domain controller, which at least includes the current autonomous driving mode level, the current decision behavior type (acceleration, deceleration, lane change, etc.), the decision confidence, and the system operation stability indication information. At the same time, the system obtains the environmental target information from the perception module, which includes the categories, relative positions, relative motion states, time-to-collision (TTC), and risk assessment indicators of multiple detected targets. The system aligns and normalizes the timestamps of the above data from different modules, and writes the processed data into the state buffer area for subsequent analysis of the state change trend.
[0047] After completing data acquisition and synchronization, the system filters the environmental target information based on the current autonomous driving decision result, and only retains the targets that have a direct impact on the current decision process as decision-related targets. This filtering process can be directly completed by the set of decision-involved targets provided by the autonomous driving system planning module, or by combining target risk assessment results, lane association relationships, and target motion trend judgments to obtain an equivalent set of decision-related targets. To ensure the clarity and controllability of the auditory cues, the system can set an upper limit on the number of targets entering the auditory cue process. When the number of candidate targets exceeds the upper limit, only the targets with higher risk levels or decision contribution are retained, and the remaining targets do not generate corresponding auditory foreground information. Through this process, the audio information that needs to be conveyed to the driver is selected, avoiding unnecessary redundancy.
[0048] After completing the decision-related target filtering, the system enters the audio mapping phase. The audio mapping phase uses a set of hierarchical time sequence models to process the internal state information and decision-related target information respectively.
[0049] First, the system constructs the internal state data in the recent period of time into an internal state sequence input according to the time sequence, which is used to represent the change trend of the autonomous driving system operating state, and inputs the sequence into the global audio feature generation model. The global audio feature generation model uses a recurrent neural network (RNN) or long short-term memory network (LSTM) structure, which extracts the global state representation reflecting the overall operating situation and change trend of the autonomous driving system by modeling the internal state sequence over time. During the model inference process, the network not only considers the internal state value at the current time, but also combines the state change at the historical time, so that the output result can reflect the continuity and trend of the system state.
[0050] Based on the global state representation, the global audio feature generation model outputs a set of global audio feature parameters for describing the overall auditory background, reflecting the overall running situation of the system, and ensuring that the driver has clear auditory feedback on the state change of the system. The global audio feature parameters at least include control quantities for controlling the overall rhythm speed, the degree of harmonic stability, and the timbre tension or brightness. These parameters jointly determine the overall style and emotional tone of the auditory prompt, which is used to express whether the current running of the autonomous driving system is calm, stable or tends to be nervous.
[0051] Specifically, the time memory characteristics of the LSTM network enable the audio to reflect the trend of state change. For example, when the system confidence continues to decline, the network automatically increases the harmonic tension, and the output rhythm becomes more nervous; when the external threat gradually disappears, the network automatically reduces the tone intensity and restores the stable rhythm.
[0052] To avoid sudden changes in sound due to state fluctuations, the system imposes smoothing and change rate limits on the parameter update process after generating the global audio feature parameters, so that the global audio features present a continuous and gradual evolution effect in the time dimension, thereby being stably output as a persistent auditory background.
[0053] While generating the global audio feature parameters, the system generates corresponding local audio attribute parameters for each decision-related target. To this end, the system constructs the state information of each decision-related target in the recent several update cycles as a target time sequence input, and inputs it into the local audio attribute generation model. The local audio attribute generation model can adopt a lightweight recurrent neural network or a feedforward neural network structure, which is used to extract the motion characteristics and risk change features of the target, and further generate local control parameters reflecting the auditory saliency of the target. At the same time, the system evaluates the importance of different targets through the attention mechanism, and assigns different auditory weights to each target according to its influence on the current decision.
[0054] Among them, the auditory weight is obtained through offline training in the system development stage. The offline training is based on the historical running log data of the autonomous driving system and the artificially designed audio reference sequence, and is used to establish the mapping relationship between the autonomous driving state and the audio parameters. The training process aims to minimize the error between the audio parameters output by the model and the corresponding parameters of the audio reference sequence, thereby determining the model parameters for auditory prompt generation.
[0055] In the actual running process of the system, the model parameters can also be adjusted based on the historical running data of the vehicle or the interactive behavior of the driver, without affecting the safety and stability of the autonomous driving system, so as to adapt to the use preferences of different drivers, and gradually make the output mode of the auditory prompt consistent with the driver's understanding habit of the state of the autonomous driving system.
[0056] Based on the above target importance evaluation results, the local audio attribute generation model outputs a set of local audio attribute parameters for each decision-related target, which are used to highlight the key targets affecting the current decision and enhance the driver's perception of the environmental target state affecting the decision state. Local audio attribute parameters include parameters for controlling the volume intensity, pitch variation range, and timbre modification degree of the target voice part. The local audio attribute corresponding to the target with higher importance has higher presence in the auditory sense, while the target with lower importance is presented in a weakened form or does not generate an independent local voice part when the conditions are met. At the same time, the system combines the spatial orientation information of the target relative to the vehicle to apply spatial sound image mapping to the local audio attribute, so that the corresponding target voice part has a clear directionality when played, enabling the driver to perceive the approximate orientation of the target through auditory perception.
[0057] For example, if there is a fast-approaching motorcycle in the right rear, an auditory cue with higher volume, sharper timbre, and right sound image position will be output, so that the driver can directly hear why the intelligent driving system is "hesitating to change lanes" at the moment.
[0058] After generating the global audio feature parameters and local audio attribute parameters, the system uniformly manages and schedules the two types of parameters. The global audio feature parameters, as the background layer of the auditory cue, persistently exist, and their update rhythm is relatively slow; the local audio attribute parameters, as the foreground layer of the auditory cue, are dynamically superimposed or removed according to the appearance, change, or disappearance of the decision-related target. By adopting different update strategies for the background layer and the foreground layer, the system avoids the dramatic disturbance of the transient changes of local targets to the overall auditory background, while ensuring that key targets can be expressed clearly and timely through auditory highlighting.
[0059] This step completes the adaptive audio encoding based on the automatic driving decision state and related environmental target information, and encodes the automatic driving decision state and related environmental target information into audio parameters of the prompt tone in real time, and ensures dynamic, continuous, and hierarchical audio prompts.
[0060] Subsequently, the system converts the global audio feature parameters and local audio attribute parameters after smoothing and constraint processing into executable audio control instructions, and generates the corresponding audio event sequence. The audio event sequence includes trigger information, duration, timbre control information, and sound image control information of the background voice part and target voice part. The system schedules the generated audio events in time and writes the events in the event buffer for a future time window to absorb system scheduling jitter and ensure the real-time and stability of audio output.
[0061] Finally, the audio synthesis and playback module generates the auditory cue audio in real time according to the audio events in the event buffer, and outputs it through the car audio system after mixing it with other audio signals in the vehicle. During the mixing process, the system can set priority and loudness limit for the auditory cue sound part, so that the auditory cue can be clearly perceived while not interfering with other necessary audio information in the vehicle.
[0062] Through the above process, the final presentation of adaptive audio coding is realized, and the system state change is transmitted to the driver in real time through dynamic audio signals, so that the driver can use hearing to perceive and understand the decision-making process of the automatic driving system in real time, thereby effectively maintaining the situational awareness of the driver when the visual attention is affected, and improving the safety of the automatic driving system.
[0063] Embodiment two
[0064] Typical scenario example: the automatic driving system cruises at 60 km / h and plans to perform right lane change.
[0065] At the initial time T0, the decision confidence is 0.92, and the corresponding auditory cue is output, which is a stable G major scale trill with uniform rhythm and bright timbre. At this time, the driver can perceive "the automatic driving system is confident to change lanes" according to the auditory cue.
[0066] At time T1, it is detected that a motorcycle is approaching quickly from the right rear, with a time to collision TTC of 2.0 seconds, and the target importance evaluation mechanism assigns a high weight of 0.82, and the corresponding auditory cue is output, which is a high-frequency short sound with a sound image biased to the right rear, and a slight dissonant interval is superimposed. At this time, the driver can immediately perceive "the automatic driving system hesitates to change lanes due to the target" according to the auditory cue of a sharp and short sound appearing from the right rear.
[0067] At time T2, the decision confidence decreases to 0.55, and the corresponding auditory cue is output, which increases the harmonic tension and more divided rhythm and tense chords appear in the audio. At this time, the driver can perceive "the automatic driving system is no longer sure whether to change lanes" according to the auditory cue.
[0068] At time T3, the motorcycle moves away, the high weight assigned by the target importance evaluation mechanism decreases, and the audio rhythm returns to a steady state, and the driver perceives "the automatic driving system regains confidence and continues to change lanes" from the audio changes of the auditory cue.
[0069] Embodiment three
[0070] An adaptive audio coding automatic driving state auditory cue system is proposed, comprising:
[0071] The data acquisition module is configured to acquire internal operation state data and environment target data of the automatic driving system at a preset update period during operation of the vehicle. The internal operation state data at least includes a current automatic driving mode level, a current decision behavior type, a decision confidence, and system operation stability indication information. The environment target data includes categories, relative positions, relative motion states, collision times, and risk assessment indexes of a plurality of detected targets.
[0072] The decision-related target screening module is configured to screen the environment target data based on a current decision result of the automatic driving system, and only keep targets that have a direct impact on the current decision process as decision-related targets. Specifically, the decision-related targets are determined based on a set of decision-involved targets output by the automatic driving system planning module, or based on a comprehensive determination of target risk assessment indexes, lane association relationships, and target motion trends.
[0073] The global audio feature generation module is configured to construct the internal operation state data into internal operation state time sequence input and perform time sequence modeling, generate a global state representation for describing overall operation situation and change trend of the automatic driving system, and generate global audio feature parameters reflecting the auditory background based on the global state representation. The global audio feature parameters at least include control parameters for controlling overall rhythm speed, harmonic stability, and timbre tension or brightness of the auditory cues.
[0074] Further, the global audio feature generation module further includes a parameter smoothing unit configured to perform smoothing processing and change rate limiting on the generated global audio feature parameters, so that the auditory background presents a continuous and gradual change effect in the time dimension.
[0075] The local audio attribute generation module is configured to construct target state data corresponding to the decision-related targets into target time sequence input, and generate corresponding local audio attribute parameters based on a target importance evaluation mechanism. The local audio attribute parameters are used to reflect the saliency of each decision-related target in the auditory sense. The local audio attribute parameters at least include parameters for controlling target voice volume intensity, pitch variation range, and timbre modification degree.
[0076] The local audio attribute generation module includes a target importance evaluation unit configured to assign auditory weights to different decision-related targets based on target risk assessment indexes, and the higher the auditory weight, the stronger the auditory saliency of the corresponding target voice part.
[0077] Further, the local audio attribute generation module further includes a spatial sound image mapping unit configured to generate spatial sound image control parameters according to spatial orientations of the decision-related targets relative to the vehicle.
[0078] The parameter scheduling module is configured to uniformly schedule the global audio feature parameters and the local audio attribute parameters, so that the global audio feature parameters are output as a persistent auditory background, and the local audio attribute parameters are output as an auditory foreground dynamically superimposed on the auditory background. The parameter scheduling module adopts different update strategies for the global audio feature parameters and the local audio attribute parameters, wherein the update frequency of the global audio feature parameters is lower than that of the local audio attribute parameters.
[0079] The audio synthesis and output module is configured to convert the scheduled global audio feature parameters and the local audio attribute parameters into audio control instructions, and perform audio synthesis and playing, so as to form auditory prompts dynamically changing with the running state and decision-making process of the automatic driving system. The audio synthesis and output module can perform audio mixing processing on the auditory prompts and other audio signals in the vehicle, and set a priority or a loudness limit for the auditory prompts.
[0080] In addition, the audio synthesis and output module performs time scheduling on the generated audio events and writes the audio events into an event buffer before performing audio synthesis and playing, so as to absorb scheduling jitter and ensure the real-time performance and stability of audio output.
[0081] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can modify the technical solutions described in the foregoing embodiments or replace some of the technical features with equivalent ones without departing from the spirit and principles of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. An adaptive audio coding method for auditory cues of autonomous driving status, applied in vehicles with autonomous driving capabilities, characterized in that, Includes the following steps: During vehicle operation, internal operating status data and environmental target data of the autonomous driving system are collected at a preset update cycle. The internal operating status data includes at least the current autonomous driving mode level, current decision behavior type, decision confidence level, and system operation stability indication information. The environmental target data includes the category, relative position, relative motion state, collision time, and risk assessment indicators of multiple detected targets. Based on the decision results of the current autonomous driving system, the environmental target data is filtered, and only targets that have a direct impact on the current decision-making process are retained as decision-related targets, while no corresponding auditory cues are generated for the remaining targets; The internal operating state data is constructed as an internal operating state time series input, and time series modeling is performed to generate a global state representation that describes the overall operating status and changing trend of the autonomous driving system. Based on the global state representation, global audio feature parameters reflecting the auditory background are generated. The target state data corresponding to the decision-related targets are constructed as target time series inputs, and corresponding local audio attribute parameters are generated based on the target importance assessment mechanism; The global audio feature parameters and local audio attribute parameters are uniformly scheduled so that the global audio feature parameters are output as a continuous auditory background, while the local audio attribute parameters are output as a dynamic auditory foreground superimposed on the auditory background. The scheduled global audio feature parameters and local audio attribute parameters are converted into audio control commands and then synthesized and played to form auditory cues that dynamically change with the operating status and decision-making process of the autonomous driving system.
2. The adaptive audio coding method for auditory cues of autonomous driving status according to claim 1, characterized in that, The selection of decision-related targets is based on the set of decision-making targets output by the autonomous driving system planning module, or on comprehensive target risk assessment indicators, lane correlation and target movement trend determination.
3. The adaptive audio coding method for auditory cues of autonomous driving status according to claim 1, characterized in that, The global audio feature parameters include at least control parameters for controlling the overall tempo of the auditory cues, the degree of harmonic stability, and the tension or brightness of the timbre.
4. The adaptive audio coding method for auditory cues of autonomous driving status according to claim 1, characterized in that, The local audio attribute parameters include at least the parameters used to control the volume intensity, pitch variation range, and timbre modification degree of the target voice.
5. The adaptive audio coding method for auditory cues of autonomous driving status according to claim 4, characterized in that, The target importance assessment mechanism includes an attention mechanism that assigns auditory weights to different decision-related targets based on target risk assessment indicators. The auditory weights are positively correlated with the corresponding local audio attribute parameters in terms of auditory saliency.
6. The adaptive audio coding method for auditory cues of autonomous driving status according to claim 4 or 5, characterized in that, The local audio attribute parameters also include spatial sound image control parameters, which are set according to the spatial orientation of the decision-related target relative to the vehicle, so that the auditory cues have directional indication.
7. The adaptive audio coding method for auditory cues of autonomous driving status according to claim 1, characterized in that, Different update strategies are used for the global audio feature parameters and the local audio attribute parameters, wherein the update frequency of the global audio feature parameters is lower than that of the local audio attribute parameters.
8. The adaptive audio coding method for auditory cues of autonomous driving status according to claim 1, characterized in that, Before the audio control commands are synthesized and played, the generated audio events are time-scheduled and written into the event buffer to absorb scheduling jitter and ensure the real-time performance and stability of the audio output.
9. The adaptive audio coding method for auditory cues of autonomous driving status according to claim 1, characterized in that, During the audio synthesis and playback process, the auditory cues are mixed with other audio signals in the vehicle, and priority or loudness limits are set for the auditory cues.
10. An adaptive audio coding-based auditory cues system for autonomous driving status, characterized in that, include: The data acquisition module is used to collect internal operating status data and environmental target data of the autonomous driving system at a preset update cycle during vehicle operation. The internal operating status data includes at least the current autonomous driving mode level, current decision behavior type, decision confidence level, and system operation stability indication information. The environmental target data includes the category, relative position, relative motion state, collision time, and risk assessment indicators of multiple detected targets. The decision-related target filtering module is used to filter the environmental target data based on the current decision results of the autonomous driving system, and retain only the targets that have a direct impact on the current decision-making process as decision-related targets; The global audio feature generation module is used to construct the internal operating state data into an internal operating state time sequence input and perform time sequence modeling to generate a global state representation that describes the overall operating status and changing trend of the autonomous driving system, and generate global audio feature parameters that reflect the auditory background based on the global state representation. The local audio attribute generation module is used to construct the target state data corresponding to the decision-related targets as target time series inputs, and generate corresponding local audio attribute parameters based on the target importance assessment mechanism. The local audio attribute parameters are used to reflect the auditory salience of each decision-related target. The parameter scheduling module is used to uniformly schedule the global audio feature parameters and local audio attribute parameters, so that the global audio feature parameters are output as a continuous auditory background, while the local audio attribute parameters are output as a dynamic auditory foreground superimposed on the auditory background. The audio synthesis and output module is used to convert the scheduled global audio feature parameters and local audio attribute parameters into audio control commands, and to synthesize and play audio to form auditory cues that dynamically change with the operating status and decision-making process of the autonomous driving system.
Citation Information
Patent Citations
Information providing device and in-vehicle device
CN111308999A
Vehicle audio control method and device, vehicle and storage medium
CN120482082A
Method and system for testing driving alertness level under influence of auditory stimulation
CN110448276A
Automatic driving manual takeover multi-mode stimulation adjusting method and system
CN114771566A