Vehicle-mounted sound effect system based on deep learning
Through the deep learning in-vehicle sound system, the sound parameters can be adjusted in real time, solving the problem that the existing technology cannot make intelligent adjustments based on the driver's cognitive state and scenario, thereby improving driving safety and sound experience.
Patent Information
- Application Number
- CN202510783232.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-19
AI Technical Summary
Existing in-car audio systems are unable to make real-time intelligent adjustments based on the driver's cognitive state and driving scenarios, resulting in insufficient driving safety and sound experience.
A deep learning-based in-vehicle sound system is used to adjust sound parameters in real time through a multi-source physiological signal acquisition module, a cognitive load assessment module, a scene recognition and expert system activation module, and a dynamic parameter adjustment module, and is continuously optimized in combination with feedback learning and optimization modules.
It achieves real-time and accurate assessment of the driver's cognitive load status, reduces driving distraction incidents, provides customized sound experience, improves driving safety and user satisfaction, and has the ability to continuously learn.
Smart Images

Figure CN120669950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle-mounted sound effect control, and more specifically, to a vehicle-mounted sound effect system based on deep learning. Background Art
[0002] With the rapid development of in-vehicle infotainment systems, drivers can enjoy increasingly rich audio content and interactive experiences while driving. However, if this audio information is not properly processed, it can increase the driver's cognitive load, distract their attention, and thus affect driving safety.
[0003] Existing in-car audio systems primarily rely on fixed sound control schemes or simple scenario mode switching, failing to precisely adjust to the driver's real-time state and complex, ever-changing driving scenarios. Some high-end models have introduced simple adaptive adjustment features based on road type or vehicle speed, but these still lack the ability to perceive the driver's cognitive state. In recent years, some studies have attempted to integrate human-computer interaction technology with in-car audio systems, but these have primarily relied on explicit interaction methods such as voice or touch, failing to capture the driver's underlying cognitive state changes. Furthermore, although some early studies have explored the application of traditional machine learning algorithms to optimize the in-car experience, these approaches still face significant challenges in real-time integration and analysis of the complex dynamic associations of multi-source heterogeneous data (such as physiological signals, driving behavior, and environmental information), thereby enabling accurate and continuous assessment of the driver's cognitive load and refined adaptive control of sound effects.
[0004] The above-mentioned existing technologies have the following major problems: lack of real-time and accurate assessment of the driver's cognitive load status, unable to determine whether the current audio information constitutes an interference with driving safety; use of a single sound processing model, difficult to provide the optimal sound experience for diverse driving scenarios; lack of an effective physiological feedback mechanism, unable to accurately adjust the sound parameters according to the driver's actual state; the optimal sound parameters in different driving scenarios vary greatly, making it difficult to meet the needs of all driving scenarios through a preset fixed solution.
[0005] Therefore, there is an urgent need for a method that can intelligently adjust the in-vehicle sound parameters in real time according to the driver's cognitive load status and driving scenario characteristics, so as to improve the audio experience while ensuring driving safety. Summary of the Invention
[0006] The present invention provides a deep learning-based in-vehicle sound system to solve the technical problem in the related art that existing in-vehicle audio systems cannot perform real-time intelligent adjustments based on the driver's cognitive state and driving scenario.
[0007] The present invention provides a deep learning-based in-vehicle sound system, comprising: Physiological signal acquisition module, used to collect multi-source physiological signal data; The cognitive load assessment module establishes an individualized baseline model based on multi-source physiological signal data to assess the driver's current cognitive load deviation; The scene recognition and expert system activation module identifies the current driving scene and the type of driving scene in real time based on multi-source physiological signal data and cognitive load deviation, and activates the multi-expert network system to generate a combination of sound effect parameters; A dynamic parameter adjustment module adaptively adjusts sound parameters, including volume, frequency response, dynamic range, and sound field parameters, based on cognitive load deviation and sound parameter combinations, and outputs dynamically adjusted sound parameters. The feedback learning and optimization module combines dynamically adjusted sound parameters and multi-source physiological signal data, monitors changes in driving behavior and physiological indicators, evaluates the effectiveness of sound adjustment, and continuously optimizes the control strategy through reinforcement learning algorithms.
[0008] Furthermore, the multi-source physiological signal data includes the driver's real-time electrocardiogram signal, the driver's sitting posture pressure distribution map, the driver's facial thermal distribution data and environmental data.
[0009] Furthermore, the cognitive load assessment module uses a multimodal CNN structure to receive time-synchronized physiological signal data, outputs a normalized cognitive load value, and calculates the cognitive load deviation by comparing it with an individualized baseline model.
[0010] Furthermore, the scene recognition adopts a ResNet structure, specifically including: Input layer: receives the driving environment feature sequence within the time window; Feature extraction layer: contains 5 residual blocks, each residual block contains two convolutional layers and skip connections, the convolutional layer kernel size is 3, and the number of channels is 64, 128, 256, 512, and 512 respectively; Feature aggregation layer: converts the feature map into a feature vector of fixed dimension through global average pooling; Classification layer: contains two fully connected layers, the number of hidden layer neurons is 256, and the number of output layer neurons is equal to the number of predefined scene categories , use the Softmax activation function to output the probability of each scene category.
[0011] Furthermore, the multi-expert network system includes four types of expert networks: High-speed driving expert network, targeting highway and long-distance driving scenarios; Urban complex expert network, targeting urban congestion and complex traffic scenarios; Urban Smoothness Expert Network, targeting smooth driving scenarios on urban roads; Stationary waiting expert network, for traffic light waiting and parking scenarios.
[0012] Furthermore, the generated sound effect parameter combination is calculated based on the output results of each expert network and the corresponding weights, and the calculation formula is: ; in Represents the final sound effect parameter combination vector, Indicates that all The output of each expert network is weighted summed. Indicates the The weight of the expert network, The feature vector representing the current audio content, Indicates the Cognitive load deviation and audio characteristics The generated sound effect parameter combination.
[0013] Furthermore, the dynamic parameter adjustment module adopts three adjustment strategies according to the cognitive load deviation: When the cognitive load deviation is higher than the preset high threshold, a load reduction strategy is adopted to reduce the interference of non-critical information; When the cognitive load deviation is lower than the preset low threshold, an enhancement strategy is adopted to improve the immersiveness of the audio experience; When the cognitive load deviation is between the high and low thresholds, a balancing strategy is adopted to keep the sound effect parameters within a moderate range.
[0014] Furthermore, the feedback learning and optimization module adopts a reinforcement learning method to construct an effect evaluation index system, including driving safety indicators and experience indicators, and optimizes the control strategy through a Q-learning algorithm.
[0015] Furthermore, the individualized baseline model is implemented using a three-layer fully connected network with the following structure: The input layer receives feature vectors from various physiological signals, and the dimension is dynamically determined according to the number of features; Hidden layer 1, contains 128 neurons and uses LeakyReLU as the activation function; Hidden layer 2, containing 64 neurons, uses LeakyReLU as the activation function; The output layer, a single neuron, outputs the cognitive load baseline value.
[0016] The present invention provides a computer storage medium for executing the above-mentioned in-vehicle sound system based on deep learning, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to execute the above-mentioned in-vehicle sound system based on deep learning.
[0017] The beneficial effects of the present invention are: through multi-source physiological signal acquisition and multi-modal CNN processing, real-time and accurate assessment of the driver's cognitive load status is achieved, driving distraction incidents are reduced, and the reaction time of key driving scenarios is shortened, while maintaining a high level of user audio experience satisfaction; scene recognition and multi-expert network structure are used to provide customized sound experience for different driving scenarios. Compared with the traditional single model solution, user satisfaction is improved and the multi-scenario adaptation problem is effectively solved; closed-loop control based on physiological feedback enables the system to accurately perceive changes in the driver's cognitive load status, achieve precise adjustment of sound parameters, improve control accuracy, and avoid attention distraction problems caused by information overload; it has the ability of continuous learning and knowledge accumulation, can continuously adapt to new scenarios and user preferences, and the long-term use effect is better than that of fixed strategy systems, and can realize knowledge transfer across vehicles, reducing the learning cost of users adapting to new cars. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a module diagram of a deep learning-based in-vehicle sound system in the present invention; Figure 2 is a flow chart of the physiological signal acquisition module of the present invention; Figure 3 is a flow chart of the cognitive load assessment module of the present invention; Figure 4 is a flow chart of the scene recognition and expert system activation module of the present invention; Figure 5 is a flow chart of the dynamic parameter adjustment module of the present invention; Figure 6 It is a flow chart of the feedback learning and optimization module of the present invention. DETAILED DESCRIPTION
[0019] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0020] At least one embodiment of the present invention discloses a vehicle-mounted sound system based on deep learning, such as Figures 1 to 6 Shown, including: Physiological signal acquisition module, used to collect multi-source physiological signal data; According to the embodiments of the present application, this module collects the driver's physiological data, driving behavior data, and environmental data through a variety of non-invasive sensing devices to achieve comprehensive perception of the driving status. Specifically, it includes: Step 1.1, collecting the driver's real-time ECG signal; The driver's real-time ECG signal is collected through the embedded ECG sensor installed on the steering wheel , including indicators such as heart rate and heart rate variability; Step 1.2, obtaining the driver's sitting posture pressure distribution map; The pressure distribution map of the driver's sitting posture is obtained through the pressure distribution sensor array installed on the seat , reflecting the driver’s body posture and tension; Step 1.3, capturing the driver's facial thermal distribution data; The infrared facial thermal imaging device installed in the instrument panel area captures the driver's facial thermal distribution data , for analyzing emotions and attention states; Step 1.4, collect environmental data; At the same time, the vehicle speed is collected through the vehicle CAN bus , acceleration , steering wheel angle Driving behavior data, as well as ambient noise levels , road type Other environmental data.
[0021] After completing the above data collection, the system synchronizes and preprocesses the various heterogeneous data to generate a multimodal driving state feature vector: ; in represents the multimodal driving state feature vector, 、 、 Respectively represent the first, second, and feature components, is the total dimension of the feature.
[0022] Optionally, in some embodiments, the system can also collect the driver's voice commands and conversation content through a voice interaction device, and analyze the driver's emotional state and attention level through natural language processing as a supplementary data source for physiological state assessment, thereby further improving the accuracy of cognitive load assessment.
[0023] Optionally, in some embodiments, the system can also record broader environmental characteristics such as ambient light conditions, weather conditions, and traffic density while collecting physiological signals, enhance the comprehensive understanding of the driving situation through multi-source information fusion, and provide richer contextual information for sound control.
[0024] The cognitive load assessment module establishes an individualized baseline model based on multi-source physiological signal data to assess the driver's current cognitive load deviation; In the embodiments of this application, a multimodal CNN network is used to process multi-source data collected by the physiological signal acquisition module in real time to build an individualized cognitive load assessment model. The physiological signal data, driving behavior data, and environmental data collected by the physiological signal acquisition module will serve as input to this module, and the driver's cognitive state will be analyzed through a deep learning model. Specifically, it includes: Step 2.1: Construct a feature extraction sub-network for different types of physiological signals and driving behavior data; For ECG signals , use one-dimensional CNN structure to extract time series features; for pressure distribution map and facial heat distribution A two-dimensional CNN structure is used to extract spatial features. For driving behavior data sequences, a long short-term memory (LSTM) structure is used to extract temporal patterns. The output features of each sub-network are combined through a feature fusion layer to generate a unified representation vector.
[0025] In this embodiment, the multimodal CNN network specifically includes the following structure: One-dimensional CNN sub-network: used to process ECG signals , contains 3 convolutional layers, with convolution kernel sizes of 5, 3, and 3 respectively, and the number of convolution kernels is 32, 64, and 128 respectively. Each convolution layer is followed by a batch normalization layer and a ReLU activation function, and finally a fixed-dimensional feature vector is obtained through a global average pooling layer; Two-dimensional CNN sub-network: used to process pressure distribution map and facial heat distribution , contains 4 convolutional layers, the convolution kernel size is 3×3, and the number of convolution kernels is 32, 64, 128, and 256 respectively. Each convolution layer is followed by a maximum pooling layer, a batch normalization layer, and a ReLU activation function. Finally, a fixed-dimensional feature vector is obtained through a fully connected layer; LSTM sub-network: used to process driving behavior data sequences, including two layers of bidirectional LSTM, with a hidden layer dimension of 128, and extracting key time series features through the attention mechanism; Feature fusion layer: The attention fusion mechanism is used to calculate the importance weights of each modal feature and obtain the final fused feature vector with a dimension of 256 by weighted averaging.
[0026] When a driver answers a phone call while driving on the highway, their ECG signal shows a 10% increase in heart rate, facial thermal distribution shows a slight rise in forehead temperature, and pressure distribution shows a more tense sitting posture. Simultaneously, steering wheel angle fluctuations increase. A multimodal CNN network processes these features to identify increased cognitive load on the driver, providing key insights for subsequent adjustments to sound parameters.
[0027] Step 2.2, construct an individualized baseline model; The baseline value of cognitive load for each driver is expressed as: ; in Indicates the The baseline value of cognitive load of each driver, 、 、 Respectively represent The first, second, and third physiological signal characteristics, represents the total number of physiological signal features, Represents the mapping function from physiological signals to cognitive load, implemented by a multi-layer perceptron.
[0028] The baseline model is trained using multiple sampling data of the driver in a no-load state.
[0029] In this embodiment, the individualized baseline model is implemented using a three-layer fully connected network with the following specific structure: Input layer: Receives feature vectors from various physiological signals, and the dimension is dynamically determined according to the number of features; Hidden layer 1: contains 128 neurons and uses LeakyReLU as the activation function; Hidden layer 2: contains 64 neurons and uses LeakyReLU as the activation function; Output layer: A single neuron that outputs a baseline value of cognitive load.
[0030] Optionally, in some embodiments, the personalized baseline model can further employ an attention mechanism to dynamically adjust the importance of different physiological signals in different driving scenarios. For example, in a high-noise environment, the system may increase the weight of facial thermal and pressure distribution features and reduce the weight of ECG signals, which may be affected by noise, to improve the robustness of the baseline assessment.
[0031] Optionally, in some embodiments, the system can construct a cognitive load prediction model through a self-supervised learning method, using abnormal driving behavior (such as sudden braking, lane departure) as an implicit supervision signal, reducing dependence on explicitly labeled data, and improving the adaptability of the model in actual driving environments.
[0032] Step 2.3, calculating the cognitive load value at the current moment based on the currently collected real-time data; The calculation formula is: ; in Indicates the current time The cognitive load value of Indicates the current time The multimodal driving state feature vector of Represents the mapping function from the current multimodal feature vector to the cognitive load evaluation value, which is implemented by the aforementioned CNN network.
[0033] Step 2.4, calculate cognitive load deviation; Cognitive load deviation is an important basis for subsequent sound effect control. The calculation formula is: ; in represents the cognitive load deviation, Indicates the current time The cognitive load value of Indicates the The baseline value of cognitive load of each driver; A positive value indicates that the current cognitive load is higher than the baseline, and the driver may be in a state of tension or distraction; A negative value indicates that the current cognitive load is lower than the baseline, and the driver may be in a relaxed or focused state.
[0034] This module outputs cognitive load deviation Serves as the key input for subsequent sound effect parameter adjustments.
[0035] The scene recognition and expert system activation module identifies the current driving scene and the type of driving scene in real time based on multi-source physiological signal data and cognitive load deviation, and activates the multi-expert network system to generate a combination of sound effect parameters; According to an embodiment of the present application, this module identifies the driving scenario type in real time based on the current driving environment and behavior data, and activates the corresponding expert network. The driving behavior data (vehicle speed, acceleration, steering wheel angle) and environmental data (noise level, road type) collected by the physiological signal acquisition module provide the basis for scene recognition. At the same time, the cognitive load deviation generated by the cognitive load assessment module serves as an important input parameter of the expert network. Specifically, it includes: Step 3.1, build a scene recognition CNN network; The network receives vehicle speed , acceleration , steering wheel angle , ambient noise level , road type The data is used as input and the probability distribution of scene categories is output: ; in represents the probability distribution vector of scene category, 、 、 Respectively represent the first, second, and The probability of scene categories, Indicates the total number of predefined scene categories.
[0036] The scenario categories include at least typical driving scenarios such as urban congestion, urban smooth driving, high-speed driving, and complex road conditions.
[0037] In this embodiment, the scene recognition CNN network adopts the ResNet structure, which specifically includes: Input layer: receives the driving environment feature sequence within the time window; Feature extraction layer: contains 5 residual blocks, each residual block contains two convolutional layers and skip connections, the convolutional layer kernel size is 3, and the number of channels is 64, 128, 256, 512, and 512 respectively; Feature aggregation layer: converts the feature map into a feature vector of fixed dimension through global average pooling; Classification layer: contains two fully connected layers, the number of hidden layer neurons is 256, and the number of output layer neurons is equal to the number of predefined scene categories , use the Softmax activation function to output the probability of each scene category.
[0038] When a vehicle enters a highway from an urban road, its speed rapidly increases from 40 km / h to 90 km / h. Acceleration increases initially but then stabilizes, steering wheel angle fluctuations decrease, and ambient noise levels increase but their distribution becomes more stable. The scene recognition CNN network accurately captures these characteristic changes, switching the scene category from "urban smooth driving" to "highway driving" and adjusting the sound effects strategy accordingly, such as enhancing low-frequency response to offset wind and engine noise.
[0039] Step 3.2, build a multi-expert network system; Expressed as: ; in 、 、 Represents the first, second, and Expert networks, each of which is responsible for optimizing the sound effects in a specific scenario. Represents the total number of predefined expert networks.
[0040] Each expert network receives the cognitive load assessment results and the current audio content features as input, and outputs the optimal sound effect parameter combination for the scenario.
[0041] In this embodiment, each expert network adopts the same structure but uses different parameter sets, specifically including: Input layer: receiving cognitive load deviation and audio content feature vector ; Hidden layer 1: contains 128 neurons and uses ReLU activation function; Hidden layer 2: contains 64 neurons and uses ReLU activation function; Output layer: Outputs a combination of sound effect parameters, including volume, equalizer settings, spatial sound field parameters, and priority weights.
[0042] Each expert network is trained and optimized for a specific scenario: The high-speed driving expert network, which focuses on highway and long-distance driving scenarios, pays more attention to the equalizer parameters that suppress environmental noise during training. Urban complex expert network, targeting urban congestion and complex traffic scenarios; Urban Smoothness Expert Network, targeting smooth driving scenarios on urban roads; Stationary waiting expert network, targeting traffic light waiting and parking scenarios When the vehicle is in a city congestion scenario, the system detects that the driver's cognitive load is high ( ), currently playing music and providing navigation instructions. The expert network for urban congestion scenarios generates a set of audio parameters that significantly reduce the music volume (by 30%), enhance the clarity of the navigation voice (+3dB gain in the mid- and high-frequency bands), and simplify the audio processing (disabling the surround sound effect) to reduce the driver's cognitive load.
[0043] Optionally, in some embodiments, the multi-expert network system may also include specialized expert networks for specific driving situations (e.g., emergency avoidance, extreme weather). For example, upon detecting an emergency, the emergency avoidance expert network will immediately take over audio control, muting all non-safety-related audio content and retaining only warning tones and necessary navigation instructions, minimizing the driver's cognitive burden.
[0044] Optionally, in some implementations, the system can employ a dynamic expert combination approach to automatically adjust the number and specialization of expert networks based on actual driving data. For example, for users who frequently drive in complex urban environments, the system might automatically generate a segmented urban driving expert network to provide more refined sound control for urban driving scenarios with varying traffic densities and road types.
[0045] Step 3.3, construct the gating network; To achieve dynamic fusion of the output results of each expert network, the calculation formula is: ; in Represents the probability distribution vector of scene category; Represents the mapping function from scene probability distribution to expert weight; is the weight vector of each expert network, expressed as: ; in represents the weight vector of each expert network, 、 、 Represents the first, second, and The weight of the expert network, Represents the total number of predefined expert networks.
[0046] In this embodiment, the gating network adopts a two-layer fully connected network structure: Input layer: receives the scene probability distribution vector ; Hidden layer: contains neurons, using the ReLU activation function; Output layer: contains neurons, using the Softmax activation function to ensure that the sum of all weights is 1.
[0047] The gating network not only considers the scene categories with the highest probability, but can also handle scene ambiguity or transition states, such as the transition stage where urban roads gradually become highways. By assigning reasonable weights to multiple related expert networks, smooth sound transitions are achieved.
[0048] When a vehicle exits a highway onto an on-ramp and prepares to enter a city road, the scene recognition CNN network might output a probability of 0.6 for "highway driving" and a probability of 0.4 for "city smoothness." The gating network then assigns weights accordingly, such as 0.55 for the highway expert network and 0.4 for the city smoothness expert network, with smaller weights assigned to other expert networks. This smooth transition avoids sudden changes in sound parameters and improves the user experience.
[0049] Optionally, in some implementations, the gating network can employ a recursive structure, incorporating historical scene transition information as additional input to smooth out scene transitions. For example, if the system detects a rapid transition from a highway to a city road, the recursive gating network can consider the transition speed and historical transition patterns to generate a smoother weight change curve, avoiding sudden changes in sound effect parameters.
[0050] Step 3.4, calculate the final sound effect parameter combination; The calculation is based on the output results of each expert network and the corresponding weights. The calculation formula is: ; in Represents the final sound effect parameter combination vector, Indicates that all The output of each expert network is weighted summed. Indicates the The weight of the expert network, The feature vector representing the current audio content, Indicates the Cognitive load deviation and audio characteristics The generated sound effect parameter combination.
[0051] Therefore, this module outputs a combination of sound effect parameters of a comprehensive multi-expert network , providing a basis for the next step of sound effect adjustment.
[0052] A dynamic parameter adjustment module adaptively adjusts sound parameters, including volume, frequency response, dynamic range, and sound field parameters, based on cognitive load deviation and sound parameter combinations, and outputs dynamically adjusted sound parameters. According to the embodiment of the present application, this module dynamically adjusts the sound parameters based on the current cognitive load assessment results and the output of multiple expert systems to achieve precise control of the sound effects. The cognitive load deviation assessed in the cognitive load assessment module and the sound effect parameter combination generated by the expert network system in the scene recognition and expert system activation module jointly guide the parameter adjustment strategy of this module. Specifically, the sound effect parameter combination output by the scene recognition and expert system activation module It needs to be further parsed into actual control signals that can be directly applied to the sound system, and according to the cognitive load deviation Determine appropriate tuning strategies to balance driving safety and audio experience. Specifically, include: Step 4.1, analyzing the sound effect parameter combination; Combining sound effect parameters Convert to actual control parameters: ; in represents the combination vector of sound effect parameters, 、 、 Represents the first, second, and sound effect control parameters, Indicates the total number of sound effect control parameters; Sound effect parameters include but are not limited to: Indicates the master volume control parameter, range [0, 1]; arrive Indicates 5-band equalizer parameters, corresponding to gain adjustment of different frequency bands, range [-12dB, +12dB]; arrive Represents spatial sound field parameters, including sound field width, depth, and positioning, ranging from [0, 1]; arrive Indicates the priority weight parameter of various audio sources (such as music, navigation, phone, warning sound, etc.), ranging from [0, 1], where Indicates the number of audio source types; arrive Indicates audio content characteristic parameters, such as dynamic range compression ratio, signal-to-noise ratio, etc. The range varies depending on the parameter. Indicates the number of audio content characteristic parameters In this embodiment, the control parameter analysis system adopts a decoupled mapping method. It first maps the unified representation vector output by the expert network to various parameter categories (such as volume, spectrum, and space), and then refines it to specific parameter values, including: Main classification layer: maps the 256-dimensional vector output by the expert network into four sub-vectors, corresponding to volume control, frequency control, spatial control, and priority control respectively; Sub-parameter mapping layer: Each sub-vector passes through a separate two-layer fully connected network to generate the specific parameter value of the corresponding category; Normalization layer: Normalizes the parameters according to their valid range to ensure that the output parameters are within the valid range.
[0053] Step 4.2, determine the sound effect adjustment strategy based on the cognitive load deviation; when When the cognitive load threshold is high, use a load-reducing strategy to reduce the interference of non-critical information: To reduce the volume of background music: ; in To adjust the coefficient, control the amplitude of volume reduction; is the volume parameter after adjustment; is the original volume parameter; To deflect cognitive load, prioritize key information (such as navigation and warnings): ; in The index of the corresponding key information, It is the adjustment coefficient to control the degree of priority improvement; is the adjusted priority weight; is the original priority weight; The function ensures that the adjusted weight does not exceed 1; simplifies sound processing: reduces the surround sound effect, reduces the dynamic range, and makes the sound more direct and clear. The driver receives a call while turning at a complex intersection. The system detects an increase in cognitive load ( , exceeding the high load threshold ), immediately activates the load-reduction strategy: it reduces the volume of the music being played from 75% to 45%, improves the clarity of the incoming call ringtone (increases the gain of the frequency band around 3kHz), and temporarily turns off the surround sound effect of the music, allowing the driver to more easily distinguish and respond to incoming calls while maintaining attention to the driving environment.
[0054] when (low cognitive load threshold), adopt an enhancement strategy to improve the audio experience quality: Optimize equalizer settings: ; in The value range is 2 to 6, corresponding to 5-band equalizer parameters; This is a scenario-related optimization increment, determined by the optimization requirements of different frequency bands; is the adjusted equalizer parameter; are the original equalizer parameters; is the enhancement coefficient, which increases as the cognitive load deviation decreases; Enhanced spatial sound field effect: Improves the width and depth of the sound field to create a more immersive listening experience; Enable advanced sound processing: such as dynamic range expansion, sound positioning enhancement, etc. The driver was driving steadily on an empty, straight highway, and the system detected that the cognitive load was at a low level ( , below the low load threshold At this point, the system activates an enhancement strategy: it gradually improves the dynamic range of the music, enhancing low-frequency response (+2.5dB at 80Hz) and high-frequency detail (+1.5dB at 10kHz), while also increasing the sound field width parameter from 0.6 to 0.85, creating a wider and more immersive sound field effect, improving the driver's music appreciation experience without affecting driving safety.
[0055] when When adjusting the sound quality, a balanced strategy is adopted to keep the sound parameters within a moderate range to ensure a balance between audio experience and driving safety.
[0056] Optionally, in some implementations, the system can also implement an adaptive threshold mechanism to dynamically adjust the high load threshold based on the user's historical cognitive load distribution. and low load threshold For example, for professional drivers who are accustomed to driving under high cognitive load, the system may appropriately increase value to avoid triggering the load reduction strategy too frequently; for older drivers or novice drivers, the system may reduce value, start the load-reduction strategy earlier and provide more driving support.
[0057] Optionally, in some implementations, the system can design specialized adjustment algorithms for specific audio content types (e.g., music, navigation, and phone calls). For example, for music content, the system not only adjusts the overall volume but also dynamically adjusts the rhythm and complexity of the music based on cognitive load. Under high cognitive load conditions, music with a stable rhythm and low complexity is prioritized, reducing the cognitive resources required to process the music.
[0058] Step 4.3: handling conflicts among multiple audio sources based on a hybrid priority mechanism; The calculation formula is: ; in Indicates the The final priority of the class audio source, represents the recommendation priority from the expert network, Indicates the priority preset by the user. Indicates the dynamic priority based on the current driving scenario, 、 and They are the weight coefficients of the recommendation priority from the expert network, the priority preset by the user, and the dynamic priority based on the current driving scenario.
[0059] In this embodiment, the priority mechanism is specifically implemented as a three-tier priority management system: Highest priority: safety-related information, such as collision warnings, emergency braking prompts, etc. Medium priority: driving-related information, such as navigation instructions, vehicle status prompts, etc. Basic priority: entertainment content such as music, radio, audiobooks, etc.
[0060] Within the same priority level, the calculated Specifically, when high-priority content needs to be played, the playback of low-priority content will be automatically slowed down or paused to ensure the effective delivery of key information.
[0061] A driver was listening to music while navigating to an unfamiliar destination when a vehicle approaching from behind triggered the rear collision warning. The system immediately prioritized the three audio sources: the rear collision warning received the highest priority (P=0.95), navigation instructions received a medium priority (P=0.75), and music received the lowest priority (P=0.35). The system then lowered the music volume by 70%, temporarily suspending the upcoming navigation instructions while clearly playing the collision warning, ensuring the driver was immediately aware of the potential danger. After the warning ended, the system resumed navigation instructions, gradually increasing the music volume as the navigation instructions completed.
[0062] Step 4.4, apply the adjusted sound effect parameters; Control each component of the car audio system to achieve precise adjustment of the sound effects. At the same time, record the adjustment results and cognitive load response data to provide data support for feedback learning in step 5.
[0063] This module outputs dynamically adjusted sound parameters and corresponding control signals to achieve real-time and precise control of the car audio system.
[0064] The feedback learning and optimization module combines dynamically adjusted sound parameters with multi-source physiological signal data, monitors changes in driving behavior and physiological indicators, evaluates the effectiveness of sound adjustments, and continuously optimizes control strategies through reinforcement learning algorithms. According to the embodiments of the present application, this module evaluates the effect of the sound effect adjustment by monitoring the changes in driving behavior and physiological indicators, and continuously optimizes the control strategy through the reinforcement learning algorithm. After the sound effect parameter adjustment is performed by the dynamic parameter adjustment module, this module collects the corresponding feedback data, evaluates the adjustment effect, and optimizes the parameters of the entire system. The physiological signal acquisition mechanism in step 1 provides an evaluation data source for this step, forming a complete closed-loop control system: from physiological signal acquisition to cognitive load assessment, to scene recognition, sound effect adjustment, and finally continuously optimizing the parameters of the entire process through feedback learning. This closed-loop design enables the system to continuously improve its performance as it is used and adapt to the characteristics of different drivers. Specifically including: Step 5.1, construct an effect evaluation indicator system; The effect evaluation indicator system includes: Safety indicators : Evaluate driving safety by monitoring driving behavior data (such as steering wheel angle stability, lane keeping ability, etc.); Experience indicators : Evaluate user satisfaction by monitoring physiological signals (such as heart rate variability, facial expressions, etc.); Comprehensive indicators: ; in is the trade-off coefficient between security and experience, , used to adjust the relative importance of security and user experience in the comprehensive evaluation; It is a comprehensive evaluation indicator; It is a safety indicator; Experience indicators In this embodiment, the specific implementation of the effect evaluation index system is as follows: Safety indicators The multi-feature fusion method is used for calculation, and the formula is: ; in For the Driving behavior data, For the The evaluation function of the driving behavior class maps the raw data to a safety score in the interval [0, 1]; For the The weight coefficient of the driving behavior type, Indicates the total number of driving behavior data types; Indicates that all The driving behavior features are weighted summed.
[0065] The main driving behavior characteristics include: Lane keeping stability: by calculating the variance of the vehicle's deviation from the lane centerline; Steering stability: through analysis of high-frequency components of steering wheel angle changes; Vehicle speed control stability: through analysis of acceleration and deceleration change rates; Reaction time: By measuring the time interval from the occurrence of a critical event to the driver's response Experience indicators It is calculated by combining physiological signals and subjective evaluation, and the formula is: ; in is an objective evaluation score based on physiological signals, is the collected subjective evaluation score, The weight coefficient for balancing the weight of objective physiological indicators and subjective evaluation in experience evaluation; The ultimate experience indicator After the system applied the load reduction strategy in the highway driving scenario, it monitored that the driver's steering wheel angle change rate decreased from the original 8.2° / s to 6.1° / s, the lane departure standard deviation decreased from 0.45m to 0.32m, and the heart rate variability returned to near the baseline level. The safety index was calculated. Improved from 0.72 to 0.85, experience index Keep it at 0.78. Use the trade-off coefficient , and get the comprehensive index: ; in It is a comprehensive evaluation indicator, indicating that the sound adjustment strategy effectively improves driving safety while maintaining a good user experience.
[0066] Step 5.2: Build a control strategy optimization system based on reinforcement learning; The system will be the current state (including cognitive load, driving scenario, etc.), and the actions performed (i.e. the adjusted sound parameters) and the rewards obtained (i.e., effect evaluation index) as input, and update the Q-value function through the Q-learning algorithm: ; in Indicates that the status Next action Q value, Indicates the next state Next action Q value; is the learning rate, which controls the speed of Q value update; is the discount factor, which determines the importance of future rewards; Indicates that the status Next action Expected cumulative reward of The immediate reward obtained for the current step; Indicates the next state The maximum Q value of all possible actions; Indicates an update operation, assigning the calculation result on the right to the variable on the left.
[0067] In this embodiment, the reinforcement learning system is implemented using a DoubleQ-Network structure, specifically including: State Space : Contains cognitive load deviation , current scene category probability distribution and audio content characteristics , forming a complete state representation by vector splicing; Action Space : Contains the adjustment amplitude and direction of the sound effect parameters, expressed in a discrete manner. For example, volume adjustment is divided into discrete levels such as {-20%, -10%, 0, +10%, +20%}; Reward Function : Directly use comprehensive evaluation indicators as a reward signal; Q network structure: Both the main Q network and the target Q network use a three-layer fully connected network with 256 and 128 hidden layer neurons respectively. They use the ReLU activation function, and the number of output layer neurons is equal to the size of the action space. Experience replay buffer: stores the last 10,000 conversion samples , mini-batch is randomly sampled for learning during training; Target network update: After every 100 training steps, the main Q network parameters are copied to the target Q network.
[0068] During driving on complex urban roads, the driver's cognitive load deviation The system detected the current scene categories as "city congestion" (with a probability of 0.82) and "complex road conditions" (with a probability of 0.15), while simultaneously playing music and navigation instructions. Based on this state, the reinforcement learning system selected an action: "Lower the music volume by 15%, improve navigation voice clarity, and simplify the spatial sound field." After executing this action, the safety metric improved by 12%, the experience metric decreased slightly by 5%, and the overall metric showed positive growth.
[0069] The system records this conversion sample and updates the Q network, gradually learning the optimal sound adjustment strategy under similar conditions.
[0070] Step 5.3, building a scene parameter mapping knowledge base; Record the optimal sound effect parameter combinations and their effect evaluation results in different scenarios. The knowledge base is in the form of key-value pairs, where the key is the scene feature vector and the value is the corresponding sound effect parameter combination and effect evaluation index.
[0071] In this embodiment, the knowledge base adopts a hierarchical storage structure, specifically including: Index layer: Uses the Locality-Sensitive Hashing (LSH) algorithm to build a multi-level index to enable fast query of scene feature vectors; Storage layer: stores data in the form of key-value pairs. Each record contains: Scene feature vector : Describes the characteristics of a specific driving scenario; Cognitive load interval : Applicable cognitive load deviation range; Sound effect parameter combination : The optimal parameter setting for this scenario; Effect evaluation indicators : Historical effect evaluation results; Frequency of use : The number of times this parameter combination is applied; Timestamp : The time of the last application; Retrieval algorithm: Use the weighted k-nearest neighbor algorithm to find the most similar historical records based on the current scene feature vector. The similarity calculation formula is: ; in Time decay weights so that recently used records receive higher weights; Increase the weight of frequency so that frequently used records receive higher weight; is the current scene feature vector; is the scene feature vector in the historical records; Represents the dot product of two vectors; and Respectively represent the modulus lengths of the two vectors; Represents the similarity between two scene feature vectors.
[0072] After repeatedly processing the "long-distance highway driving" scenario, the system has established a stable set of sound effect parameter configurations in its knowledge base. These include a moderate boost in low-frequency response (+2dB gain between 80-200Hz) to mask road noise, a slight compression of the dynamic range (compression ratio of 3:1) to ensure clarity in varying speeds, and a moderate surround sound field (width parameter of 0.65) to provide immersion without distraction. When the system encounters a similar scenario again, it retrieves this set of parameters directly from the knowledge base as the initial configuration, then fine-tunes it based on the specific circumstances, significantly reducing adaptation time and computing resource consumption.
[0073] Optionally, in some implementations, the knowledge base can adopt a hierarchical structure, comprising a global experience layer and a personal experience layer. The global experience layer stores universal optimization experience collected from multiple users, serving as an initial reference for new users or new scenarios. The personal experience layer records personalized optimization results for specific users, taking precedence over global experience. This structure leverages collective intelligence to rapidly adapt to new situations while also enabling the gradual establishment of highly personalized control strategies.
[0074] Optionally, in some implementations, the system can implement a cross-vehicle knowledge transfer mechanism, allowing users to carry their personalized sound control strategies with them when changing vehicles. The system automatically adjusts the learned parameter mappings based on the new vehicle's acoustic characteristics and sound system parameters, enabling seamless migration and reducing the learning curve for users adapting to a new vehicle.
[0075] Step 5.4, optimize the parameters of each expert network and gating network based on the historical data in the knowledge base; The optimization goal is to maximize the comprehensive index Expected value: ; in Represents the parameter set of the neural network model; Represents finding the parameters that maximize the objective function ; Indicates that the parameter Comprehensive indicators under conditions expected value.
[0076] In this embodiment, the model optimization system adopts a hybrid training method, which specifically includes: Offline batch training periodically (e.g. weekly or when the accumulated data volume reaches a threshold) uses historical data in the knowledge base to batch update each expert network and gating network, and uses the stochastic gradient descent algorithm to minimize the loss function: ; in For the The comprehensive index of the samples, is the regularization coefficient, which controls the regularization strength. It is the L2 regularization term to prevent the model from overfitting; is the total number of samples; Indicates taking the average of all N samples; is the loss function.
[0077] Online incremental update: During real-time interaction, small-batch incremental learning is performed on samples with outstanding performance (i.e., samples with comprehensive indicators above the average level). The update rule is: ; in is the incremental learning rate, which controls the parameter update step size; is the loss function with respect to the parameter gradient; is the parameter value before updating; is the updated parameter value.
[0078] Model pruning and quantization: To address the computing resource limitations of the vehicle environment, the model is regularly pruned (deleting unimportant connections) and quantized (reducing parameter accuracy) to reduce model size and computational complexity while maintaining performance.
[0079] After two months of use, the system optimized its expert network for highway driving scenarios based on approximately 3,500 collected sample data points. Before optimization, the average overall performance index for this scenario was 0.76; after optimization, the average overall performance index for the same scenario increased to 0.83, while simultaneously reducing model size by 18% and inference time by 25%. Specifically, the system learned to gradually enhance the dynamic range and spatial effects of music during stable highway driving with low cognitive load, while rapidly streamlining audio processing and highlighting safety-related information when encountering complex traffic conditions.
[0080] Step 5.5: Periodically update the user-individualized baseline model; Adapt to long-term changes in user driving habits and preferences to maintain the continued effectiveness of the system.
[0081] In this embodiment, the personalized baseline model update system adopts a method that combines incremental learning with a forgetting mechanism, specifically including: Data collection trigger conditions: When one of the following conditions is met, the system will collect the user's physiological data in a relaxed driving environment for baseline update: A predetermined time (e.g., 30 days) has passed since the last update; the system detects a change in the baseline value of the user's physiological indicators (e.g., the heart rate baseline value deviates by more than 15% during 5 consecutive driving sessions); the user actively requests recalibration of the system Incremental learning algorithm, using the exponentially weighted moving average (EWMA) method to update the baseline model parameters: ; in is the update rate, which controls the fusion ratio of the new and old model parameters; are the original baseline model parameters; are the model parameters trained based on new data; are the updated model parameters.
[0082] After three months of using the system, one user began a fitness program, and their resting heart rate gradually decreased from 76 beats per minute to 68 beats per minute. The system detected a change in heart rate baseline exceeding a threshold during five consecutive driving sessions, automatically triggering a baseline model update. During the update, the system collected new physiological data from the user in a relaxed driving environment and retrained the baseline model parameters, while retaining 70% of the original model's characteristics for stability. After the update, the system successfully adapted to the user's new physiological characteristics, accurately assessed cognitive load status, and avoided erroneous sound effect adjustments caused by baseline shift.
[0083] The feedback learning and sound optimization steps in the above examples enable the system's adaptive optimization and continuous learning capabilities, enabling the sound control strategy to continuously adapt to new scenarios and user preferences, while striking a good balance between audio experience and driving safety. The innovation of this step lies in the establishment of a closed-loop learning mechanism, enabling the system to continuously optimize the control strategy based on actual results, significantly improving long-term performance.
[0084] A computer storage medium for executing the above-mentioned deep learning-based in-vehicle sound system, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to execute the deep learning-based in-vehicle sound system according to any one of claims 1 to 9.
[0085] Here, the present invention provides an implementation example: a long-distance driving scenario on a highway; Driver Mr. Li drove a car equipped with the system of the present invention on a long-distance drive on the highway, covering a distance of about 300 kilometers with an estimated driving time of 3 hours. During the drive, Mr. Li played his personal music playlist and used the navigation system.
[0086] Initial stage (10 minutes after entering the highway): The system collected Mr. Li's physiological data through electrocardiogram (ECG), seat pressure distribution, and facial thermal imaging. It also collected driving behavior data such as vehicle speed (approximately 100 km / h) and steering angle (with minimal fluctuations). The scene recognition CNN network identified the current situation as a "high-speed driving scene" with a confidence level of 92%.
[0087] The cognitive load assessment module compared Mr. Li's individualized baseline model and assessed the current cognitive load deviation as -0.15 (below the average level), indicating that Mr. Li was in a relaxed driving state.
[0088] The gated network primarily activates the Highway Driving Expert Network (weighted 0.85) and the City Smoothness Expert Network (weighted 0.15). The system implements an "enhancement strategy," moderately increasing the music volume to 65% of the preferred setting, increasing the low-frequency response gain by 3dB for enhanced musical impact, and widening the soundstage to 120° for enhanced spatial perception, while maintaining the navigation prompt volume 10dB higher than the music.
[0089] The system detected a slight increase in Mr. Li's blinking rate, a slight change in his seat pressure distribution, and a 10% decrease in his heart rate variability. The cognitive load assessment module calculated a current cognitive load deviation of +0.12 (slightly above average).
[0090] The system implements a "balanced strategy," automatically reducing the music volume to 55%, reducing the low-frequency gain to 1dB, and narrowing the soundstage width to 90°. It also analyzes the current music genre and adjusts the EQ settings to reduce complex frequency components. Additionally, the navigation prompt volume is increased to 15dB above the music, and the sound processing is simplified.
[0091] Special situation handling (encountering a construction section after driving for 2 hours) The system detected a sudden decrease in vehicle speed (from 110 km / h to 60 km / h) and an increase in steering wheel operation. The scene recognition CNN network identified the current scenario as a combination of high-speed driving and complex road conditions. Simultaneously, physiological signals indicated a 10% increase in heart rate, and pressure distribution indicated increased muscle tension.
[0092] The cognitive load assessment module calculated the current cognitive load deviation to be +0.35 (above average). The gating network adjusted its weight distribution, reducing the weight of the highway expert network to 0.55 and increasing the weight of the urban complexity expert network to 0.45.
[0093] The system immediately implements the "burden reduction strategy": Automatically pause music playback and keep only necessary navigation prompts; Improve the clarity of navigation prompts and optimize frequency response to highlight the human voice frequency band; Simplify all sound processing effects; Turn off incoming call and message notifications.
[0094] After passing the construction section, the system detects that the cognitive load has returned to normal and gradually resumes music playback, but the volume setting is 10% lower than before.
[0095] Feedback learning records After the entire trip is completed, the system records the parameter adjustments and driving performance data during the driving process, including: Optimal volume range for highway scenarios (50%-65%); Effective load-reduction strategies in complex situations such as construction sections; A sound effect setting to relieve fatigue after driving for 2 hours.
[0096] This data is stored in Mr. Li’s personal experience layer knowledge base for future optimization of similar scenarios. At the same time, some anonymized data is uploaded to the global experience layer to improve the overall performance of the system.
[0097] Here, the present invention provides an implementation example: a complex urban traffic scenario; Ms. Zhang was driving a car equipped with the system of the present invention during the morning rush hour (8:30 a.m.) in the city center, where traffic was congested and she had to start and stop frequently. At the same time, she received multiple navigation turn instructions and important calls from work.
[0098] Initial status assessment; The system identified the current situation as a "complex urban scene" (confidence level 87%). Through physiological signal acquisition, the system detected that Ms. Zhang's cognitive load deviation was +0.25, indicating a high cognitive load.
[0099] The gated network mainly activates the urban complex expert network (weight 0.78) and the static waiting expert network (weight 0.22). The system defaults to the "burden reduction strategy": The music volume was automatically set to 40% below Ms. Zhang’s personal preference; Simplify sound processing and optimize mid-frequency vocal clarity; Navigation instructions use simple prompts and streamlined voice commands.
[0100] Multitasking: When Ms. Zhang received an incoming call, the system detected an increase in cognitive load deviation to +0.42. The system immediately implemented an enhanced "load reduction strategy": completely pausing music playback; reducing ambient noise in the car (via active noise cancellation); optimizing call quality to emphasize the clarity of the other party's voice; and streamlining navigation prompts, providing only the briefest instructions at key turning points. After the call ended, the system detected a decrease in cognitive load deviation to +0.28 and restored the previous audio settings, but with the volume reduced by 5%.
[0101] Adapting to Changing Traffic Conditions: When the vehicle exits a congested area and enters a smoother road, the scene recognition CNN network updates the scene to a "smooth urban scene" (confidence increases to 81%). Simultaneously, physiological signals indicate that the cognitive load deviation has decreased to +0.10.
[0102] The gated network redistributes weights, with the weight of the urban fluent expert network rising to 0.75 and the weight of the urban complex expert network falling to 0.25. The system gradually shifts to a "balanced strategy": Smoothly increase the music volume to 60%; Restore appropriate sound processing to enhance the three-dimensional sense of music; Navigation cues remain clear but require less intervention.
[0103] Personalized learning: The system records the physiological patterns Ms. Zhang exhibits while answering calls and associates them with a knowledge base. The next time a similar pattern is detected, it will more quickly trigger appropriate traffic-reduction strategies. The system also notes Ms. Zhang's preference for certain types of music in congested environments and prioritizes recommendations for similar scenarios in the future.
[0104] Here, the present invention provides an implementation example: an example of cross-vehicle knowledge transfer; Mr. Wang is a long-term user of this system and has been using it for 6 months on his car, brand A. He now buys a new car, brand B, which is also equipped with the system of the present invention, but the sound system and the acoustic environment inside the car are different.
[0105] knowledge transfer process; Basic information migration: The system downloads Mr. Wang’s personal experience layer data from the cloud, including: personalized baseline model parameters; scenario preference settings; historical optimization parameter records; and cognitive load threshold settings.
[0106] Acoustic Environment Adaptation: The system performs acoustic environment testing to measure the frequency response characteristics, sound field characteristics, and background noise level of the new vehicle. The test shows that Brand B vehicle has better low-frequency response (3.5dB higher low-frequency gain), lower background noise (5dB lower), and different speaker positioning configuration than Brand A vehicle.
[0107] Parameter mapping adjustment: The system automatically adjusts the parameter mapping relationship based on acoustic differences: the volume setting ratio is adjusted to 0.85 times (due to lower background noise); the low-frequency gain is adjusted to 0.7 times (due to better low-frequency response); the optimal sound field parameters are recalculated to adapt to the new speaker configuration; and the noise reduction parameters for voice recognition and calls are optimized accordingly.
[0108] Adaptability Verification: During the first drive, the system collected feedback data more frequently to monitor the effectiveness of parameter mapping adjustments: physiological data showed that Mr. Wang's cognitive load pattern was consistent with historical data; the system fine-tuned some parameter mappings, such as the volume curve and frequency response curve; after three drives, the system completed the adaptive adjustment of major parameters and switched to normal operating mode.
[0109] Here, the present invention provides an implementation example: differentiated processing of different audio contents; Ms. Liu is driving a suburban road in a vehicle equipped with this system, playing music, receiving navigation instructions, and occasionally answering calls and voice messages.
[0110] Differentiated processing of audio content; The system recognizes that Ms. Liu's current cognitive load deviation is -0.05 (close to the baseline) and implements the "balance strategy."
[0111] Regarding music content: Analyze the type of music currently being played (classical music); set an EQ curve suitable for this type to enhance mid-frequency details and high-frequency clarity; apply a moderate reverberation effect to enhance the sense of space; set reasonable dynamic range compression to ensure that quiet passages are not masked by ambient noise; set the volume to 65% of your personal preference.
[0112] Navigation command processing; When navigation instructions need to be read: audio spectrum analysis is used to identify "quiet gaps" in the current music; navigation information is intelligently broadcast during these gaps to reduce interference; "audio ducking" technology is applied to temporarily reduce the music volume by 20% instead of completely interrupting it; adaptive equalization is used for navigation voice to highlight the 2-4kHz frequency band for vocal clarity; navigation prompts use a personalized voice, which has been tested to have the least impact on Ms. Liu's cognitive load.
[0113] phone and messaging handling; When a call comes in: the system evaluates the current driving complexity and cognitive load status; when the cognitive load deviation is less than +0.3, it allows incoming call reminders and enables noise reduction processing; automatically reduces the music volume to 20% and simplifies sound processing; activates the directional microphone array to improve the quality of call voice collection; and applies real-time noise reduction and echo cancellation algorithms to optimize call quality.
[0114] For voice messages: The system decides whether to play them immediately or delay them based on the current cognitive load; during playback, sound effects similar to those of a phone call are applied, but the volume is 10% lower than that of a phone call; and a simple voice reply function is provided to minimize distraction.
[0115] Alarm information processing: When the system needs to play a safety-related alarm: regardless of the currently playing content, the safety alarm is given the highest priority; the volume and notification method are adjusted according to the alarm level (emergency alarms use full-band prompt sounds); visual and tactile feedback are combined to reduce the burden on a single sense; after the alarm ends, the previous audio state is intelligently restored to avoid sudden changes that cause distraction.
[0116] These application examples demonstrate that the deep learning-based in-car audio system of this invention can intelligently adjust various sound parameters in real time based on the driver's cognitive load, driving scenario, and audio content type, maximizing the audio experience while ensuring driving safety. Through continuous learning and feedback adjustments, the system continuously optimizes its performance, adapting to individual differences and user habits, achieving truly intelligent and personalized in-car audio control.
[0117] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A deep learning-based in-vehicle sound system, characterized by: include: Physiological signal acquisition module, used to collect multi-source physiological signal data; The cognitive load assessment module establishes an individualized baseline model based on multi-source physiological signal data to assess the driver's current cognitive load deviation; The scene recognition and expert system activation module identifies the current driving scene and the type of driving scene in real time based on multi-source physiological signal data and cognitive load deviation, and activates the multi-expert network system to generate a combination of sound effect parameters; A dynamic parameter adjustment module adaptively adjusts sound parameters, including volume, frequency response, dynamic range, and sound field parameters, based on cognitive load deviation and sound parameter combinations, and outputs dynamically adjusted sound parameters. The feedback learning and optimization module combines dynamically adjusted sound parameters and multi-source physiological signal data, monitors changes in driving behavior and physiological indicators, evaluates the effectiveness of sound adjustment, and continuously optimizes the control strategy through reinforcement learning algorithms.
2. The in-vehicle sound system based on deep learning according to claim 1, characterized in that: The multi-source physiological signal data includes the driver's real-time electrocardiogram signal, the driver's sitting posture pressure distribution map, the driver's facial thermal distribution data and environmental data.
3. The in-vehicle sound system based on deep learning according to claim 1, characterized in that: The cognitive load assessment module uses a multimodal CNN structure to receive time-synchronized physiological signal data, outputs a normalized cognitive load value, and calculates the cognitive load deviation by comparing it with an individualized baseline model.
4. The in-vehicle sound system based on deep learning according to claim 1, characterized in that: The scene recognition adopts the ResNet structure, which specifically includes: Input layer: receives the driving environment feature sequence within the time window; Feature extraction layer: contains 5 residual blocks, each residual block contains two convolutional layers and skip connections, the convolutional layer kernel size is 3, and the number of channels is 64, 128, 256, 512, and 512 respectively; Feature aggregation layer: converts the feature map into a feature vector of fixed dimension through global average pooling; Classification layer: contains two fully connected layers, the number of hidden layer neurons is 256, and the number of output layer neurons is equal to the number of predefined scene categories , use the Softmax activation function to output the probability of each scene category.
5. The in-vehicle sound system based on deep learning according to claim 1, characterized in that: The multi-expert network system includes four types of expert networks: High-speed driving expert network, targeting highway and long-distance driving scenarios; Urban complex expert network, targeting urban congestion and complex traffic scenarios; Urban Smoothness Expert Network, targeting smooth driving scenarios on urban roads; Stationary waiting expert network, for traffic light waiting and parking scenarios.
6. The in-vehicle sound system based on deep learning according to claim 1, characterized in that: The generated sound effect parameter combination is calculated based on the output results of each expert network and the corresponding weights, and the calculation formula is: ; in Represents the final sound effect parameter combination vector, Indicates that all The outputs of the expert networks are weighted summed. Indicates the The weight of the expert network, The feature vector representing the current audio content, Indicates the Cognitive load deviation and audio characteristics The generated sound effect parameter combination.
7. The deep learning-based in-vehicle sound system according to claim 1, characterized in that: The dynamic parameter adjustment module adopts three adjustment strategies according to the cognitive load deviation: When the cognitive load deviation is higher than the preset high threshold, a load reduction strategy is adopted to reduce the interference of non-critical information; When the cognitive load deviation is lower than the preset low threshold, an enhancement strategy is adopted to improve the immersiveness of the audio experience; When the cognitive load deviation is between the high and low thresholds, a balancing strategy is adopted to keep the sound effect parameters within a moderate range.
8. The in-vehicle sound system based on deep learning according to claim 1, characterized in that: The feedback learning and optimization module adopts the reinforcement learning method to build an effect evaluation index system, including driving safety indicators and experience indicators, and optimizes the control strategy through the Q-learning algorithm.
9. The in-vehicle sound system based on deep learning according to claim 3, characterized in that: The individualized baseline model is implemented using a three-layer fully connected network with the following structure: The input layer receives feature vectors from various physiological signals, and the dimension is dynamically determined according to the number of features; Hidden layer 1, contains 128 neurons and uses LeakyReLU as the activation function; Hidden layer 2, containing 64 neurons, uses LeakyReLU as the activation function; The output layer, a single neuron, outputs the cognitive load baseline value.
10. A computer storage medium, characterized in that It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to execute the deep learning-based in-vehicle sound system described in any one of claims 1-9.
Citation Information
Cited By
Multi-channel audio linkage playing device based on Bluetooth broadcast and implementation method
CN120916105A