Vehicle-mounted multi-sensory collaborative interaction method and system based on multi-dimensional audio features

By combining semantic and acoustic dual-stream analysis of the in-vehicle audio system with environmental perception, a multi-dimensional sensory control vector is generated to drive the multi-sensory execution unit to work together. This solves the problems of monotonous experience and noise interference in the in-vehicle audio system, and realizes a multi-sensory immersive experience and low-latency feedback.

CN121963738APending Publication Date: 2026-05-01ZHEJIANG HEQIAN ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG HEQIAN ELECTRONIC TECH CO LTD
Filing Date
2025-12-12
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing in-vehicle audio systems suffer from a lack of diverse experience dimensions, a lack of coordinated interaction among multi-sensory units, an inability to achieve a multi-sensory immersive experience, and a decline in user perception in high-noise environments.

Method used

By acquiring in-vehicle audio signals in real time, using semantic analysis and acoustic analysis in parallel, combining environmental perception for feature fusion and dynamic scheduling, multi-dimensional sensory control vectors are generated to drive multi-sensory execution units to work collaboratively, and the output is optimized through adaptive filtering and PID closed-loop control.

Benefits of technology

It achieves multi-sensory collaborative interaction, enhances the multi-dimensional matching of the audio experience, reduces environmental noise interference, ensures driving safety and low-latency sensory feedback, and has low power consumption and strong adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963738A_ABST
    Figure CN121963738A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted multi-sensory collaborative interaction method and system based on multi-dimensional audio features, and the method comprises the steps: obtaining a vehicle-mounted original audio source signal, employing a double-flow parallel analysis architecture, extracting the semantic features of lyrics through employing an NLP model, and extracting the acoustic physical features through employing a CNN model; weighted arbitration is carried out on the double-flow features based on the vehicle driving state, and a basic sensory control vector is generated; processing an in-vehicle environment audio signal by using an adaptive noise cancellation algorithm, calculating an environment masking coefficient, and performing dynamic gain correction on a control vector; and finally, in combination with reverse time sequence scheduling and PID closed-loop control, executing units such as fragrance, an atmosphere lamp and a seat vibration motor are driven to work cooperatively. According to the method, the problem of emotion deficiency of single physical feature control is solved through double-flow fusion, the road noise masking effect is eliminated through environmental perception compensation, and accurate immersive interactive experience is achieved through hardware time sequence synchronization.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for in-vehicle multi-sensory collaborative interaction based on multi-dimensional audio features Technical Field

[0001] This invention relates to the field of smart cockpit technology, and in particular to an in-vehicle multi-sensory collaborative interaction method and system based on multi-dimensional audio features. Background Technology

[0002] With the rapid development of the automotive industry and multimedia technology, in-vehicle audio systems have become one of the core configurations of car interiors, and their main function is to provide passengers with audio playback services such as music and radio. Currently, some in-vehicle scenarios have seen attempts at single-sensory assisted experience technologies, such as simple "voice-controlled lights," which mechanically drive ambient lights to flash by extracting the amplitude or frequency characteristics of audio signals.

[0003] Currently, some sensory assistance technologies used in in-vehicle scenarios have significant shortcomings: on the one hand, the experience is limited in scope and lacks deep integration with audio content. Sensory experiences are often fixed or require manual adjustment, failing to dynamically match the audio content, emotion, and context. On the other hand, a multi-sensory collaboration mechanism has not been established; each sensory unit works independently without collaborative interaction, making it difficult to achieve a synergistic effect where 1+1>2. These technologies fail to fully engage passengers' other sensory systems and cannot meet users' demands for an immersive, multi-dimensional experience.

[0004] Therefore, there is an urgent need for an in-vehicle interactive system that can integrate audio semantics and acoustic features, and combine in-vehicle environmental perception to achieve multi-sensory collaborative output. Summary of the Invention

[0005] This invention addresses the shortcomings of existing in-vehicle audio experiences, such as limited variety, lack of coordinated multi-sensory unit interaction, and poor adaptability to in-vehicle scenarios. It proposes a complete technical solution combining "multi-module collaboration + AI-precise analysis + real-time dynamic scheduling," as detailed below:

[0006] In a first aspect, the present invention provides an in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features, comprising:

[0007] Step S1: Data acquisition, real-time acquisition of the original audio source signal output by the in-vehicle audio playback source;

[0008] Step S2: Parallel parsing of dual-stream features, directly using the original audio source signal as input data, and processing it separately into the semantic analysis branch and the acoustic analysis branch:

[0009] The semantic analysis branch is used to identify text content in audio and extract semantic tags and sentiment features.

[0010] The acoustic analysis branch is used to extract the physical acoustic features of the audio, which include at least rhythm, intensity, or spectral energy information.

[0011] Step S3: Feature fusion decision: Based on the preset bimodal arbitration strategy, the semantic tags, emotional tendency features and physical acoustic features are weighted and fused to generate a basic sensory control vector.

[0012] Step S4: Cooperative conversion output, converting the basic sensory control vector into hardware control instructions for multiple vehicle-mounted sensory execution units, and driving the sensory execution units to perform actions.

[0013] Furthermore, step S1 also includes collecting in-vehicle ambient audio signals through an in-vehicle audio pickup device;

[0014] The method also includes a vector gain correction step:

[0015] Using the original audio source signal as a reference, the in-vehicle ambient audio signal is subjected to signal differential processing to calculate the ambient masking coefficient;

[0016] The intensity parameter dimension of the basic sensory control vector generated in step S3 is dynamically superimposed using the environmental masking coefficient to generate a corrected multidimensional sensory control vector.

[0017] In step S4, the modified multidimensional sensory control vector is converted into hardware control instructions.

[0018] Furthermore, the vector gain correction step specifically includes:

[0019] Using the original audio source signal as the reference signal of the adaptive filter, convolution and difference operations are performed on the in-vehicle ambient audio signal to filter out the music echo component and separate the ambient noise component.

[0020] The environmental masking coefficient is calculated based on the energy amplitude of the environmental noise component.

[0021] When the environmental masking coefficient indicates an increase in environmental noise, the intensity parameter values ​​of the tactile and auditory dimensions in the basic sensory control vector are increased by a preset ratio.

[0022] Furthermore, the dual-modal arbitration strategy in step S3 specifically includes:

[0023] The semantic analysis branch uses a natural language processing model to extract scene keywords and emotional polarity from the lyrics;

[0024] The acoustic analysis branch uses a convolutional neural network to extract the energy spectrum, tempo, and pitch variation features of the audio signal;

[0025] When there is a logical conflict between the sentiment tendency output by the semantic analysis branch and the physical acoustic features output by the acoustic analysis branch, the weight ratio of the two is dynamically adjusted according to the current vehicle driving scenario or user preference mode to generate the basic sensory control vector.

[0026] Furthermore, the method also includes a security gating step based on dual-channel signal verification:

[0027] Real-time comparison of the energy correspondence between the original audio source signal and the in-vehicle environmental audio signal;

[0028] When the original audio source signal is detected to have signal strength, and the effective energy of the in-vehicle ambient audio signal is lower than a preset safety threshold, the system is determined to be in silent mode, circuit fault, or abnormal sound field state.

[0029] When an abnormal state is determined, the intensity parameters of the tactile and olfactory dimensions in the corrected multidimensional sensory control vector are forcibly set to zero.

[0030] Furthermore, the collaborative conversion output in step S4 also includes a timing synchronization step:

[0031] Obtain the inherent physical response delay time of each in-vehicle sensor execution unit;

[0032] Based on the inherent physical response delay time, the sending time of the hardware control command is reversed and scheduled, and the command is sent to each sensory execution unit at different advance times, so that the sensory effects produced by each sensory execution unit are synchronized at the same target time.

[0033] Furthermore, the timing synchronization step also includes closed-loop feedback control:

[0034] Real-time feedback on the actual working status of the sensory execution unit;

[0035] The deviation between the actual working state and the target hardware control command is calculated using the PID control algorithm, and the parameters of the hardware control command are dynamically corrected in real time.

[0036] Furthermore, the sensory execution unit includes at least:

[0037] Olfactory actuators, including in-vehicle fragrance modules or scent generators, are used to adjust the type of scent;

[0038] The visual execution unit includes an in-vehicle ambient lighting module and an in-vehicle display screen. The in-vehicle ambient lighting module is used to adjust the color of the light effect and the flashing frequency, and the in-vehicle display screen is used to display dynamic wallpapers or dynamic images that match the scene.

[0039] The tactile actuator includes a vehicle seat vibration module for providing tactile feedback based on the basic sensory control vector or hardware control commands.

[0040] An environmental simulation unit is configured to adjust the airflow state and physical properties inside the vehicle according to the basic sensory control vector; the airflow state includes at least airflow speed or airflow mode, and the physical properties include at least air humidity or temperature; it is used to provide meteorological tactile feedback that matches the semantic tag.

[0041] Secondly, the present invention provides an in-vehicle multi-sensory collaborative interaction system based on multi-dimensional audio features, used to implement the aforementioned in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features, characterized in that it includes:

[0042] The signal acquisition module is configured to acquire the original audio source signal of the vehicle audio system;

[0043] A dual-stream parsing processor configured to perform semantic and acoustic analysis;

[0044] The feature fusion unit is configured to execute a bimodal arbitration strategy and generate basic sensory control vectors;

[0045] The execution drive network is configured to convert vectors into instructions and drive the sensory execution units via the vehicle bus.

[0046] Furthermore, it also includes:

[0047] An environmental perception module is configured to connect to an in-vehicle audio pickup device and execute a differential algorithm to calculate an environmental masking coefficient, wherein the environmental masking coefficient characterizes the degree to which in-vehicle environmental noise masks the original audio source signal.

[0048] The feature fusion unit is also configured to receive the environment masking coefficient and use it to perform gain correction on the basic sensory control vector to generate a corrected multidimensional sensory control vector for driving the execution driving network.

[0049] Beneficial effects:

[0050] This invention overcomes the shortcomings of existing technologies from the root by employing a closed-loop architecture of "data acquisition - precise analysis - collaborative control - multi-sensory output".

[0051] 1. Dual-channel acquisition + adaptive noise reduction ensures the purity of audio data, providing reliable input for AI. It solves the problems of large resolution errors caused by environmental interference and poor interference of single acquisition in existing technologies, improving the AI ​​resolution accuracy to over 92% and providing a reliable basis for scene matching.

[0052] 2. The semantic-acoustic dual-branch fusion analysis breaks through the limitations of single feature analysis. Compared with vibration control that relies solely on rhythm, it can achieve multi-dimensional matching of "scene-emotion-rhythm". For example, ocean-themed songs can simultaneously trigger sea salt fragrance, sea breeze simulation and blue light to form a multi-sensory immersive experience.

[0053] 3. The synchronous scheduling mechanism of "prediction + PID correction" establishes reverse timing scheduling and PID closed-loop control for the physical delay characteristics of different sensory hardware. Combined with the low latency characteristics of CAN / LIN hybrid communication, the audio and sensory output delay is controlled within 50ms, which is far lower than the threshold of human hearing and other senses (100ms), thus avoiding a disconnect in the experience.

[0054] 4. A safety gating system based on dual-channel signal verification was established, which effectively prevented accidental triggering in silent or faulty states. In particular, it forcibly blocked highly invasive tactile and olfactory feedback, ensuring driving safety.

[0055] 5. By using dual-channel differential processing, the original audio signal is used as the reference source for the adaptive filter to perform echo cancellation (AEC) on the in-vehicle sound field signal collected by the microphone, and the in-vehicle environmental noise component is separated in real time. Based on this noise component, the system can dynamically compensate the output intensity of the sensory execution unit (e.g., automatically enhance tactile vibration feedback when the ambient noise is high), thereby solving the technical problem that the existing technology relies solely on the sound source signal for control, resulting in a decline in user perception experience in high-speed and high-noise environments.

[0056] 6. Modular design and energy consumption scheduling strategy enable the system to have standby power consumption ≤5W and operating power consumption ≤30W, and can be adapted to different vehicle interior layouts, solving the problems of high energy consumption and poor adaptability of existing technologies. Attached Figure Description

[0057] The following figures are for illustrative purposes only and do not limit the scope of the invention.

[0058] Figure 1 is a system hardware architecture diagram of an embodiment of the present invention.

[0059] Figure 2 is a flowchart of the interaction method according to an embodiment of the present invention.

[0060] Figure 3 is a flowchart of the dual-stream parsing and fusion sub-flow of an interactive method according to an embodiment of the present invention.

[0061] Figure 4 is a flowchart of the environment perception and vector correction sub-process of an interaction method according to an embodiment of the present invention.

[0062] Figure 5 is a flowchart of the collaborative transformation output sub-process of the interactive method according to an embodiment of the present invention. Detailed Implementation

[0063] To provide a clearer understanding of the technical features, objectives, and effects of this invention, specific embodiments are now described with reference to the accompanying drawings, in which the same reference numerals denote the same parts. For the sake of simplicity, the parts related to this invention are shown schematically in each drawing and do not represent their actual structure as a product. Furthermore, for the sake of clarity and ease of understanding, in some drawings, components with the same structure or function are shown only schematically, or only one is labeled.

[0064] In this invention, "connection" can include direct connection, indirect connection, communication connection, electrical connection, unless otherwise specified.

[0065] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly specifies otherwise. It will also be understood that, when used in the specification, the terms “comprising” and / or “including” mean the presence of the stated features, values, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, values, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the listed related items.

[0066] The "original audio source signal" described in this invention refers to the audio electrical signal generated by the in-vehicle entertainment system that has not yet been propagated through the physical sound field inside the vehicle. It can be in digital form (e.g., I2S or PCM data streams directly obtained from the SOC chip, or audio data packets at the APP level) or analog form (e.g., voltage signals derived from the preamplifier or power amplifier circuitry of the in-vehicle amplifier). Regardless of the signal format, its core characteristic is that the signal represents the intended playback content of the in-vehicle system and does not include environmental noise such as wind noise, tire noise, and occupant voices collected by the microphone.

[0067] As shown in Figure 1, Embodiment 1 of the present invention provides an in-vehicle multi-sensory collaborative interaction system based on multi-dimensional audio features. The system adopts a computing architecture with an in-vehicle microcontroller unit (MCU) as the core and realizes distributed control of sensory execution units through a hybrid bus network.

[0068] Specifically, the system first includes a signal acquisition module configured to construct a dual-channel data input link. The first channel connects directly to the audio output port of the in-vehicle infotainment system via an I2S (Inter-IC Sound) digital audio bus to acquire the pure, raw audio source signal without digital-to-analog conversion. The second channel connects to the in-vehicle audio pickup device via an A2B (Automotive Audio Bus) digital bus; in this embodiment, an in-vehicle microphone array is used to acquire the in-vehicle ambient audio signal, including music echoes and environmental noise. This dual-channel design ensures that the system possesses both pure source data for precise analysis and on-site feedback data for environmental perception.

[0069] At the computational level, the system is equipped with a main control and computing unit, which preferably adopts a heterogeneous chip solution of automotive-grade MCU and DSP. The DSP unit is dedicated to performing high-performance signal processing tasks, including running adaptive filtering algorithms for echo cancellation and differential noise reduction, and performing FFT (Fast Fourier Transform) to extract frequency domain features of audio (such as Mel spectrum); it also deploys lightweight CNN inference models (such as models optimized based on the MobileNet architecture after pruning) for millisecond-level acoustic feature extraction. The MCU unit is responsible for logic control tasks and has a built-in collaborative control strategy library. This strategy library is a pre-built key-value mapping table that stores thousands of mapping rules, such as [tag: ocean] -> [fragrance: sea salt, lighting: RGB(0,0,255), wind speed: level 2]. The MCU also runs PID control algorithms and a task scheduler, responsible for distributing final instructions.

[0070] The system also includes a dual-stream parsing processor for performing semantic and acoustic analysis. In this embodiment, this module employs a heterogeneous parallel computing architecture. Although logically it is handled by a single processor, in its physical implementation, its computational tasks are distributed across different hardware cores to achieve optimal performance.

[0071] Acoustic Analysis Branch (running on DSP): Utilizing the powerful floating-point arithmetic and matrix parallel processing capabilities of the DSP unit, it performs FFT transformation and CNN model inference in real time, and is responsible for extracting high-frequency updated physical acoustic features, which include at least rhythm, intensity or spectral energy information.

[0072] Semantic Analysis Branch (running on MCU): Utilizing the logic processing capabilities of the MCU unit, it runs a lightweight NLP model, responsible for processing infrequently updated text / lyrics data, and extracting semantic tags and sentiment features.

[0073] The DSP and MCU interact with each other through inter-chip shared memory or high-speed buses (such as SPI / IPC) to ensure that the two analysis streams remain parallel and synchronized on the time axis.

[0074] Accordingly, the system also includes a feature fusion unit, which runs in the collaborative control strategy library of the MCU unit. It is configured to execute a bimodal arbitration strategy, that is, to receive semantic tags and physical acoustic features from the dual-stream parsing processor, perform weighted calculations according to preset logical rules, and finally generate a basic sensory control vector.

[0075] To address perception bias and false triggering issues in in-vehicle environments, the system includes an environmental perception module. This module primarily runs on the DSP unit and is configured to execute an adaptive filtering algorithm. Using the original audio source signal as a reference, it performs differential processing on the in-vehicle ambient audio signal to calculate the environmental masking coefficient. This module is also responsible for feeding back the calculated coefficients to the feature fusion unit. The feature fusion unit uses these coefficients to perform dynamic gain superposition on the intensity parameters in the basic sensory control vector, generating the final corrected multi-dimensional sensory control vector used to drive the driver network.

[0076] At the execution level, the system connects various sensory actuators through an execution drive network. To balance control response speed and hardware cost, this drive network adopts a CAN / LIN hybrid bus architecture. For latency-sensitive devices, such as tactile and visual actuators, they are connected to the high-speed CAN bus; for comfort devices with high inertia, such as olfactory actuators and environmental simulation units, they are connected to the low-speed LIN bus.

[0077] In this embodiment, the haptic actuator includes a vehicle seat vibration module. This module embeds at least four 3020-type linear vibration motors with a response time of less than 20ms within the seat, supporting frequency adjustment from 20-200Hz and amplitude adjustment from 0-100% to match music rhythm or ambient atmosphere. The CAN bus can transmit vibration waveform data at a rate of 500kbps, achieving millisecond-level haptic start and stop, closely following the music's beat.

[0078] The visual execution unit includes an in-vehicle ambient lighting module and an in-vehicle display screen. The in-vehicle ambient lighting module includes RGB ambient light groups (2700K-6500K color temperature) distributed throughout the vehicle, used to adjust the light effect color and flashing frequency. The in-vehicle display screen is used to display dynamic wallpapers or dynamic images that match the scene.

[0079] To optimize the experience, the visual execution unit is configured to support an immersive light and shadow linkage mode: when the in-vehicle display screen is detected to be in an idle display state that is not for navigation or safety prompts, a matching dynamic wallpaper is displayed on the in-vehicle display screen according to the scene atmosphere, and the RGB ambient light group is controlled to present a lighting effect consistent with the main color of the dynamic wallpaper; when the in-vehicle display screen is in a function occupied state, only the RGB ambient light group is controlled to adjust independently.

[0080] The system features a display priority arbitration module. High priority includes reversing camera, navigation map, and vehicle alarm information; low priority includes music player interface and live wallpaper. When a high-priority task occupies the screen, the visual feedback strategy automatically downgrades to "lights-only response" mode. In this mode, the central control screen displays the navigation image, while the ambient lighting continues to pulsate according to the meaning and rhythm of the music, ensuring that the cabin atmosphere is maintained without interfering with the acquisition of driving information. When the high-priority task ends (such as exiting the reversing mode), the system automatically reverts to "screen-light linkage" mode, and the screen background is synchronized with the lighting again, achieving a seamless integration of visual experience.

[0081] The olfactory actuator, including an in-vehicle fragrance module or a scent generator, is used to adjust the scent type. In this embodiment, the in-vehicle fragrance module is preferably installed in a 3-compartment fragrance box at the air vent of the center console, and the fragrance type switching and concentration (0-100% adjustable) are controlled by a solenoid valve.

[0082] The environmental simulation unit is configured to adjust the airflow and physical properties inside the vehicle. In this embodiment, it is configured to integrate a blower (0.5-5 m / s wind speed) and an ultrasonic atomizer (0-50% RH water vapor volume) at the front air vent to simulate natural environmental scenarios (such as sea breeze and mountain forest breeze).

[0083] To ensure stable operation of the system in the complex electromagnetic environment of an in-vehicle environment, this system integrates an electromagnetic shielding module into its hardware architecture. Specifically, this includes signal isolation and filtering circuits on the audio signal transmission path between the microphone array and the DSP, and on the bus control path between the MCU and the execution unit. For high-frequency computing cores such as the MCU and DSP, conductive shielding covers or metal shielding layers are used for physical enclosure and grounding. An electromagnetic interference filter is integrated at the system power input to filter out high-frequency noise in the in-vehicle power network, providing clean power to the dual-stream analytical processor and ensuring the stability of AI operations.

[0084] As shown in Figure 2, Embodiment 2 of the present invention provides an in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features. This method operates according to the following logical flow based on the aforementioned hardware architecture:

[0085] First, in step S1 (data acquisition), the system acquires the original audio source signal and the in-vehicle ambient audio signal in real time. Since there is a slight physical time difference between the two signals on the transmission path, the DSP unit will pre-perform cross-correlation calculation to calculate the time offset and apply a corresponding delay buffer to the original audio source signal to achieve microsecond-level alignment of the two signals on the time axis.

[0086] Next, the process proceeds to step S2 (parallel parsing of dual-stream features). In this step, the system directly uses the pure original audio source signal as input to ensure that the analysis results are not affected by environmental noise. As shown in Figure 3, this step is divided into two parallel branches: the semantic analysis branch uses an NLP (Natural Language Processing) model to extract keywords (such as "ocean," "forest," "rose," "sunshine"), emotional tendencies (such as "joyful," "sad," "soothing," "exhilarating"), and scene information from the lyrics, and outputs semantic labels L based on the scene depicted in the audio. scene (e.g., ocean, forest), output emotional tendency feature E based on the emotional polarity of the audio text content. sem ([-1,1], -1 represents sadness, 1 represents joy). The acoustic analysis branch first preprocesses the audio signal to extract Mel-frequency cepstral coefficients (MFCCs), then inputs this spectrum into a convolutional neural network (CNN), outputting physical acoustic features (Ii) containing rhythm, energy, and emotional intensity. ac and E ac ). Among them, I ac Represents energy intensity (0-100), reflecting the loudness and rhythmic density of sound; E ac Represents acoustic emotion (-1 to +1), reflecting whether the melody is in a major key (cheerful) or a minor key (sad).

[0087] Next, in step S3 (feature fusion decision), the system performs weighted fusion of the two-stream features based on a preset dual-modal arbitration strategy to generate a basic sensory control vector V. control Specifically, when semantic emotions (such as sadness) logically conflict with acoustic characteristics (such as excitement), the system incorporates the vehicle's driving state as an adjustment factor. The "vehicle state" mentioned in this step is a driving aggression index (λ) calculated in real-time by the MCU based on chassis CAN bus data. The MCU collects accelerator pedal opening, engine speed / motor torque, and vehicle speed change rate (acceleration) in real-time. When rapid acceleration or high engine speed is detected, it is determined to be an aggressive driving condition (λ approaches 1); when constant speed cruising or smooth driving is detected, it is determined to be a comfortable driving condition (λ approaches 0).

[0088] The system generates control vectors dimension by dimension according to the following logic:

[0089] First Dimension: Scene Dimension (D) sceneThis dimension determines the "main atmosphere tone" (such as fragrance type, wallpaper theme). The generation strategy is to directly use semantic tags (L) from S2. scene If there are no lyrics, the acoustic classification scenario will be used instead.

[0090] Second Dimension: Emotional Dimension (D) emotion This dimension determines the "light and shadow tone" (such as warm and cool colors). This dimension is most prone to conflict (e.g., lyrics expressing sadness). sem = -0.8, but the melody is exciting E ac =+0.6). The fusion strategy uses the vehicle state λ for weighted arbitration. Formula: D emotion = (1-λ)·E sem +λ·E ac When in a comfortable operating condition (λ small), the driver is in a calm state of mind, and the system assigns higher weight to semantic features, prioritizing the display of the delicate emotions in the lyrics (output sadness); when in an aggressive operating condition (λ large), the driver seeks excitement, and the system assigns higher weight to acoustic features, prioritizing the display of the melody's excitement (output cheerfulness).

[0091] Third dimension: Intensity dimension (D) intensity This dimension determines the "physical feedback strength" (such as amplitude and wind speed). The fusion strategy is based on acoustic energy I. ac Based on this, the excitation gain of the vehicle state λ is superimposed. Formula: D intensity =I ac ·(1+k·λ), where k is the aggressive gain coefficient (e.g., k=0.5). The more aggressive the driving (the larger λ is), the stronger the physical feedback intensity (vibration, wind sensation) output by the system relative to the original music energy, in order to match the driver's highly aroused state.

[0092] Ultimately, the system outputs the basic sensory control vector V. control =[D scene D emotion D intensity This is used to map subsequent steps into specific hardware instructions.

[0093] To address perception bias and false triggering issues in in-vehicle environments, the system performs environmental masking calculations. As shown in Figure 4, the DSP unit runs an adaptive filter algorithm, such as an LMS (Least Mean Square Error) adaptive filter, to estimate the transfer function from the speaker to the microphone. Specifically, using the original audio source signal as a reference, a convolution operation is performed on the signal using the weight coefficient vector of the adaptive filter. This convolution operation simulates the reflection, delay, and attenuation characteristics of the original audio propagating in the vehicle's physical space, thereby generating an echo estimation signal. Subsequently, the in-vehicle ambient audio signal captured by the microphone is compared with this echo estimation signal using a difference operation. Since the echo estimation signal highly fits the music echo component in the microphone signal, the difference between the two can cancel out music interference, thus separating the clean ambient noise component.

[0094] Based on this noise component, the system performs two key operations: first, safety gating, which forcibly blocks highly intrusive seat vibrations and fragrance release when the original audio source signal energy is high but the ambient signal energy is extremely low (indicating silence or line fault), preventing "ghost triggering"; second, masking effect calculation, which calculates the degree of audible masking of the original audio by the ambient noise, and calculates the energy ratio of this noise component relative to the original audio based on the Weber-Fechner law regarding the logarithmic relationship between the sensory threshold and the intensity of the background stimulus. Specifically, when the signal-to-noise ratio in the low-frequency band is lower than a preset threshold (e.g., 15dB), the system generates a positive environmental masking coefficient based on the logarithmic decay curve.

[0095] If the security check passes and the environmental masking coefficient indicates that road noise is masking the low frequencies of the music, the system performs a vector gain correction step. The system automatically adds gain to the intensity dimension of the control vector using the environmental masking coefficient, generating a corrected multi-dimensional sensory control vector (e.g., for every 3dB increase in road noise, the vibration intensity increases by 10%). Utilizing the "multi-sensory complementation effect" of touch on hearing, even if the actual signal-to-noise ratio at the user's ear decreases, the enhanced tactile feedback maintains the user's subjective psychological experience at a "shocking" level, achieving immersive and consistent perception.

[0096] Finally, in step S4 (cooperative conversion output), the system converts the multidimensional sensory control vector into the final hardware driver instructions. The MCU decouples and distributes this multidimensional sensory control vector, for example, the fragrance module subscribes to D... scene When you receive "Ocean," switch to Fragrance Chamber #2. Ambient Light Module Subscription D emotion and D intensity D emotion Determine the color (blue / red), D intensity Determines the peak brightness. Vibration / wind sensor module subscription D intensity D intensity It is directly mapped to the motor PWM duty cycle and fan speed.

[0097] Step S4 also includes two key technical aspects, as shown in Figure 5:

[0098] The first step is timing synchronization. To address the varying physical response delays of different hardware components (e.g., a 500ms delay for the fragrance, a 50ms delay for the seat vibration motor, and a 20ms delay for the ambient lighting), the system establishes a data buffer for pre-aiming control. The MCU reads the audio characteristics at the end of the buffer in advance. Commands are sent at different pre-emptive moments to ensure that all sensory feedback reaches its peak synchronously at the climax of the music. For example, when a climax is anticipated 500ms later, a LIN command is immediately sent to the fragrance module; a CAN command is sent to the seat motor after a 450ms delay; and a command is sent to the ambient lighting after a 480ms delay. Despite the different command sending times, the fragrance, vibration, and lighting all reach their peak simultaneously at 500ms, eliminating any sense of disjointed experience.

[0099] Secondly, there is PID closed-loop control. Addressing the actual output error of the actuator, the system dynamically adjusts the duty cycle of the control signal using a PID algorithm, combined with sensor feedback, to ensure precise execution of the control command. The sensor actuator receives hardware drive commands, and the system collects the operating status signals of the sensor actuator in real time. For motor-type loads, this status signal can be a sampled value of the drive current or a speed pulse signal; for fluid-type loads, it can be a flow monitoring signal or a pressure signal. The MCU performs the following calculations in millisecond cycles (e.g., 10ms): First, it calculates the deviation between the setpoint and the feedback value, and outputs the basic drive energy proportionally based on the current deviation. When the deviation is large (e.g., when the command has just been issued and the device has not yet started), it outputs a high-level signal to overcome static friction and achieve a rapid response. If the output is consistently slightly lower than the setpoint (steady-state error) due to aging or environmental resistance, the integral term gradually increases the duty cycle of the drive signal until the error returns to zero. This effectively solves the problem of "commands arriving but insufficient force." When the device output approaches the target value, the differential term will generate a reverse suppression effect to prevent output overshoot, ensure a smooth transition of sensory effects, and avoid abrupt vibrations or airflow impacts.

[0100] The final correction value obtained from the PID calculation will be mapped to specific electrical signal parameters for output:

[0101] For devices driven by PWM (Pulse Width Modulation) (such as ambient light brightness and motor speed): the PID output value directly corrects the high-level duration (duty cycle) of the PWM waveform. For example, when the feedback display shows insufficient brightness, the PWM waveform is automatically widened.

[0102] For voltage-driven devices: the PID output value corrects the output voltage amplitude of the DAC (digital-to-analog converter).

[0103] The above descriptions are merely preferred embodiments of the present invention, and the present invention is not limited to the above embodiments. Those skilled in the art will understand that the forms in these embodiments are not limited thereto, nor are the adjustments possible. It is understood that other improvements and variations directly derived or conceived by those skilled in the art without departing from the basic concept of the present invention should be considered to be included within the scope of protection of the present invention.

Claims

1. A vehicle-mounted multi-sensory collaborative interaction method based on multi-dimensional audio features, characterized in that: include: Step S1: Data acquisition, real-time acquisition of the original audio source signal output by the in-vehicle audio playback source; Step S2: Parallel parsing of dual-stream features, directly using the original audio source signal as input data, and processing it into the semantic analysis branch and the acoustic analysis branch respectively: the semantic analysis branch is used to identify the text content in the audio and extract semantic tags and sentiment features; The acoustic analysis branch is used to extract the physical acoustic features of the audio, which include at least rhythm, intensity, or spectral energy information; Step S3: Feature fusion decision, based on a preset bimodal arbitration strategy, the semantic tags, emotional tendency features, and physical acoustic features are weighted and fused to generate a basic sensory control vector; Step S4: Collaborative conversion output, the basic sensory control vector is converted into hardware control instructions for multiple in-vehicle sensory execution units, and the sensory execution units are driven to perform actions.

2. The in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features as described in claim 1, characterized in that, Step S1 further includes acquiring in-vehicle ambient audio signals through an in-vehicle audio pickup device; the method further includes a vector gain correction step: using the original audio source signal as a reference, performing signal differential processing on the in-vehicle ambient audio signal to calculate the ambient masking coefficient; using the ambient masking coefficient to dynamically superimpose the intensity parameter dimension in the basic sensory control vector generated in step S3 to generate a corrected multidimensional sensory control vector; in step S4, the corrected multidimensional sensory control vector is converted into hardware control instructions.

3. The in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features as described in claim 2, characterized in that, The vector gain correction step specifically includes: using the original audio source signal as the reference signal of the adaptive filter, performing convolution and difference operations on the in-vehicle environmental audio signal to filter out the music echo component and separate the environmental noise component; calculating the environmental masking coefficient based on the energy amplitude of the environmental noise component; when the environmental masking coefficient indicates that the environmental noise has increased, increasing the intensity parameter values ​​of the tactile and auditory dimensions in the basic sensory control vector by a preset ratio.

4. The in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features as described in claim 1, characterized in that, The bimodal arbitration strategy in step S3 specifically includes: the semantic analysis branch uses a natural language processing model to extract scene keywords and emotional polarity from the lyrics; the acoustic analysis branch uses a convolutional neural network to extract the energy spectrum, beat, and pitch variation features of the audio signal; when there is a logical conflict between the emotional tendency output by the semantic analysis branch and the physical acoustic features output by the acoustic analysis branch, the weight ratio of the two is dynamically adjusted according to the current vehicle driving scenario or user preference mode to generate the basic sensory control vector.

5. The in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features as described in claim 2, characterized in that, The method further includes a safety gating step based on dual-channel signal verification: real-time comparison of the energy correspondence between the original audio source signal and the in-vehicle ambient audio signal; when the original audio source signal is detected to have signal strength, while the effective energy of the in-vehicle ambient audio signal is lower than a preset safety threshold, the system is determined to be in silent mode, circuit fault, or abnormal sound field state. When an abnormal state is determined, the intensity parameters of the tactile and olfactory dimensions in the corrected multidimensional sensory control vector are forcibly set to zero.

6. The in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features as described in claim 1, characterized in that, The collaborative conversion output in step S4 also includes a timing synchronization step: obtaining the inherent physical response delay time of each vehicle-mounted sensory execution unit; based on the inherent physical response delay time, performing reverse timing scheduling on the sending time of the hardware control command, sending commands to each sensory execution unit at different advance times, so that the sensory effects produced by each sensory execution unit are synchronized at the same target time.

7. The in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features as described in claim 6, characterized in that, The timing synchronization step also includes closed-loop feedback control: real-time acquisition of the actual working status feedback of the sensor execution unit; calculation of the deviation between the actual working status and the target hardware control command using a PID control algorithm, and real-time dynamic correction of the parameters of the hardware control command.

8. The in-vehicle multi-sensory collaborative interaction method based on multi-dimensional audio features as described in claim 1, characterized in that, The sensory execution unit includes at least: an olfactory execution unit, including an in-vehicle fragrance module or odor generator, for adjusting the odor type; a visual execution unit, including an in-vehicle ambient lighting module and an in-vehicle display screen, wherein the in-vehicle ambient lighting module is used to adjust the light effect color and flashing frequency, and the in-vehicle display screen is used to display dynamic wallpapers or dynamic images matching the scene; a tactile execution unit, including an in-vehicle seat vibration module, for providing tactile feedback according to the basic sensory control vector or hardware control commands; and an environmental simulation unit configured to adjust the airflow state and physical properties inside the vehicle according to the basic sensory control vector; wherein the airflow state includes at least airflow speed or airflow mode, and the physical properties include at least air humidity or temperature; and is used to provide meteorological tactile feedback matching the semantic tags.

9. A vehicle-mounted multi-sensory collaborative interaction system based on multi-dimensional audio features, used to implement the method of any one of claims 1 to 8, characterized in that, include: The signal acquisition module is configured to acquire the original audio source signal of the vehicle audio system; A dual-stream parsing processor configured to perform semantic and acoustic analysis; The feature fusion unit is configured to execute a bimodal arbitration strategy and generate basic sensory control vectors; the execution drive network is configured to convert the vectors into instructions and drive each sensory execution unit through the vehicle bus.

10. The system as described in claim 9, characterized in that, Also includes: An environmental perception module is configured to connect to an in-vehicle audio pickup device and execute a differential algorithm to calculate an environmental masking coefficient, wherein the environmental masking coefficient characterizes the degree to which in-vehicle environmental noise masks the original audio source signal; the feature fusion unit is further configured to receive the environmental masking coefficient and use it to perform gain correction on the basic sensory control vector to generate a corrected multidimensional sensory control vector for driving the execution driving network.