Vehicle-mounted audio processing method and vehicle
By using AI models to analyze, render, and tune audio data, the dependence of Dolby Atmos technology on dedicated audio sources has been resolved. This enables real-time panoramic sound conversion of any stereo system, providing an intelligent and personalized in-car immersive audio experience and lowering the barrier to entry for high-end audio functions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-14
AI Technical Summary
Current Dolby Atmos technology relies heavily on dedicated audio sources, making it difficult to enhance the experience of massive amounts of existing traditional audio data. The immersive experience depends entirely on the production quality and lacks intelligent adaptive capabilities.
The AI model analyzes the audio data, separates it into multiple mono audio streams, extracts multi-dimensional metadata, renders and adjusts the data based on this metadata, and sends the rendered audio stream to the speaker for playback, thus achieving real-time panoramic sound conversion for any stereo sound.
It breaks away from the reliance on pre-built panoramic sound audio data, providing a highly intelligent and personalized in-car immersive audio experience without the need for high licensing fees, significantly lowering the threshold and cost of high-end audio functions.
Smart Images

Figure CN121865196A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of automotive electronics technology and audio signal processing, and more particularly to a method for processing in-vehicle audio, a vehicle, an electronic device, and a computer-readable storage medium. Background Technology
[0002] Currently, with the increasing popularity of electric vehicles, the interiors are quieter, providing a better environment for high-quality audio. At the same time, this also places higher demands on the capabilities of audio systems and the sources of content.
[0003] Dolby Atmos and similar technologies represent the current state of immersive audio. However, these technologies have significant limitations, such as a high dependence on dedicated audio sources, difficulty in enhancing the experience of massive amounts of existing traditional audio data, and a lack of intelligent adaptive capabilities on the playback end, meaning the immersive experience depends entirely on the production quality. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and computer-readable storage medium for processing in-vehicle audio, aiming to improve the problems of Dolby Atmos and similar technologies, which are highly dependent on dedicated audio sources, making it difficult to enhance the experience of massive amounts of existing traditional audio data, and the lack of intelligent adaptive capabilities at the playback end, with the immersive experience entirely dependent on the production level.
[0005] This application discloses a method for processing in-vehicle audio, the method comprising: Obtain the audio data to be played in the vehicle; The audio data is parsed to obtain multiple mono audio streams of the audio data and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. Based on the multi-dimensional metadata, the multiple mono audio streams are rendered and tuned, and the rendered and tuned multiple mono audio streams are sent to the corresponding speakers in the vehicle for playback.
[0006] Optionally, the multi-dimensional metadata also includes music style information, and the rendering and tuning of the multiple mono audio streams based on the multi-dimensional metadata includes: Based on the music style information, determine the target equalization curve of the audio data; Based on the target equalization curve, equalization processing is performed on the multiple mono audio streams.
[0007] Optionally, after performing equalization processing on the plurality of mono audio streams based on the target equalization curve, the method further includes: Obtain the speaker layout information of the vehicle; For each mono audio stream, the target speaker and speaker playback parameters are determined based on the corresponding sound image position coordinates and the speaker layout information.
[0008] Optionally, after determining the target speaker and speaker playback parameters for each mono audio stream based on the corresponding sound image position coordinates and the speaker layout information, the method further includes: Determine the audio playback energy for each mono audio stream; Based on the audio playback energy of each mono audio stream, the dominant mono audio stream is determined from the plurality of mono audio streams; The audio playback energy of other mono audio streams is adjusted based on the ratio of the audio playback energy of the dominant mono audio stream to the playback energy of the channel.
[0009] Optionally, based on the audio playback energy of each mono audio stream, determining the dominant mono audio stream from the plurality of mono audio streams includes: Based on the audio playback energy of each mono audio stream, determine the average playback energy of each mono audio stream; The dominant mono audio stream is determined from the plurality of mono audio streams based on the average playback energy of each mono audio stream.
[0010] Optionally, sending the rendered and tuned multiple mono audio streams to the corresponding speakers in the vehicle for playback includes: Based on the speaker playback parameters, the multiple mono audio streams are sent to their respective target speakers for playback.
[0011] Optionally, the speaker playback parameters include any one or more of the following: Amplitude adjustment parameters, phase adjustment parameters, and delay adjustment parameters.
[0012] This application discloses a vehicle, the vehicle comprising: The audio acquisition module is used to acquire the audio data to be played in the vehicle; An audio processing module is used to parse the audio data to obtain multiple mono audio streams of the audio data and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. The audio rendering module is used to render and adjust the multiple mono audio streams based on the multi-dimensional metadata, and send the rendered and adjusted multiple mono audio streams to the corresponding speakers in the vehicle for playback.
[0013] This application also discloses an electronic device, including a processor and a memory, wherein... Memory, used to store computer programs; A processor is used to execute a program stored in memory to implement the method described in the embodiments of this application.
[0014] This application also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method described in this application.
[0015] The embodiments of this application have the following advantages: In this embodiment, audio data to be played in a vehicle is acquired, and the audio data is parsed to obtain multiple mono audio streams and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. Based on the multi-dimensional metadata, the multiple mono audio streams are rendered and adjusted, and the rendered and adjusted multiple mono audio streams are sent to the corresponding speakers in the vehicle for playback. This realizes real-time panoramic sound conversion of traditional stereo audio data, breaks the dependence on pre-made panoramic sound audio data, and provides a highly intelligent and personalized in-vehicle immersive audio experience. It eliminates the need to pay high licensing fees for scarce panoramic sound content or engage in complex cooperation with content providers. It can provide a top-notch audio experience using existing music libraries, significantly reducing the threshold and cost of providing high-end audio functions. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the steps of a vehicle audio processing method according to an embodiment of this application; Figure 2This is a schematic flowchart of a prior art audio playback technology provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the steps of an AI model audio processing procedure provided in an embodiment of this application. Figure 4 This is a flowchart illustrating a method for processing in-vehicle audio according to an embodiment of this application; Figure 5 This is a block diagram of a module structure in a vehicle provided in an embodiment of this application; Figure 6 This is a structural diagram of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the technical problems, technical solutions, and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] To facilitate understanding of the technical solutions and effects of the embodiments of this application, the relevant technologies of this application will be briefly described below.
[0019] Currently, Dolby Atmos and similar technologies represent the advanced level of immersive audio. Its core is the concept of "object audio," which means that during the content production stage, the content is divided into beds and objects, and given three-dimensional spatial metadata. At the playback end, the codec renderer restores the three-dimensional sound field of the recording studio based on this metadata.
[0020] However, such technologies have significant limitations: Highly dependent on content: Users must obtain specially produced panoramic sound sources, and the vast amount of existing stereo or ordinary multi-channel sound sources cannot provide an immersive experience.
[0021] Static and fixed: The audio object and its spatial position are pre-made during production and cannot be dynamically adjusted according to user preferences or the in-car environment during playback.
[0022] Creativity is limited: the immersive experience of the audio depends entirely on the level of post-production, and the playback device is just a "replay" tool, lacking intelligence and adaptability.
[0023] like Figure 1The diagram illustrates a technical solution for Audio Vivid (immersive audio) in existing technologies. It shows the entire process from inputting the audio source produced by Audio Vivid into the car's infotainment system, to unpacking, decoding, rendering, and finally tuning in the DSP (Digital Signal Processor). Dolby Atmos uses the same implementation path as Audio Vivid, but the SDKs (Software Development Kits) for packaging, encoding, decoding, and rendering differ. For both Dolby Atmos and Audio Vivid, users must obtain specially produced immersive sound sources. The vast amount of existing stereo or ordinary multi-channel audio sources cannot provide an immersive experience. Furthermore, current upmixing algorithms require inferring and separating the spatial attributes of sound elements from the mixed audio without knowing the original multi-channel recording information, leading to poor transient response and significant artificial artifacts. Wet upmixing, on the other hand, is too simplistic, lacks variation, and produces mediocre results.
[0024] With the increasing prevalence of electric vehicles, the quieter interiors provide a better environment for high-quality audio, but also place higher demands on the capabilities of audio systems and the sources of content. Therefore, there is an urgent need in this field for an innovative solution that can overcome the limitations of pre-built content and intelligently, in real-time, and personalizedly render immersive sound from any audio source.
[0025] This application provides a method for processing in-vehicle audio. It acquires audio data to be played in a vehicle, processes the audio data using a pre-set AI model to obtain multiple mono audio streams and multi-dimensional metadata, generates a control data stream based on the multi-dimensional metadata, and synchronously transmits the control data stream and the multiple mono audio streams to a pre-set digital signal processor. The digital signal processor, based on the control data stream, renders and adjusts the multiple mono audio streams, and sends the rendered and adjusted mono audio streams to the corresponding speakers in the vehicle for playback. This achieves real-time panoramic sound conversion of any stereo traditional audio data, breaking the dependence on pre-made panoramic sound audio data and providing a highly intelligent and personalized in-vehicle immersive audio experience. It eliminates the need to pay high licensing fees for scarce panoramic sound content or engage in complex collaborations with content providers; a top-tier audio experience can be provided using existing music libraries, significantly reducing the barrier and cost of providing high-end audio functionality.
[0026] Reference Figure 2 The diagram illustrates a flowchart of a method for processing in-vehicle audio according to an embodiment of this application, specifically including the following steps: Step 101: Obtain the audio data to be played in the vehicle.
[0027] In step 101, conventional audio data with arbitrary stereo to be played can be acquired first.
[0028] Step 102: Parse the audio data to obtain multiple mono audio streams of the audio data and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams.
[0029] In step 102, the acquired audio data to be played can be input into a preset AI model. The AI model parses the audio data to be played to obtain multiple mono audio streams and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio represents the audio playback energy ratio relationship between the multiple mono audio streams.
[0030] Specifically, such as Figure 3 As shown, the AI model can perform two processes on the audio data to be played: audio element separation and metadata extraction, to obtain multiple mono audio streams and multi-dimensional metadata of the audio data.
[0031] For audio element separation, the AI model can perform semantic-level track segmentation of the audio data to be played. The AI model does not perform simple frequency band separation, but rather, based on a semantic understanding of the music's composition, it separates the mixed stereo signal (i.e., the audio data to be played) in real time, such as separating it into at least six independent mono element tracks, specifically including: vocals, drums, bass, guitar, piano, and other instruments. Simultaneously, to preserve some spatial information of the audio data to be played, the AI model can perform the above separation process separately for the left and right channels, thereby generating a total of 12 high-fidelity PCM (Pulse-Code Modulation) audio streams (i.e., mono audio streams). It should be noted that the number of mono element tracks separated in real time can also be other numbers.
[0032] For metadata extraction, the AI model can parse the audio data to be processed in the same forward processing process to obtain multi-dimensional metadata associated with multiple mono audio streams. This multi-dimensional metadata can include at least the channel playback energy ratio of multiple mono audio streams and the sound image position information of each mono audio stream. In addition, the multi-dimensional metadata can also include the music style information of the audio data to be played.
[0033] Among them, the sound image position coordinates can be the ideal three-dimensional sound image position (X, Y, Z coordinates) of each mono element corresponding to the separated mono audio stream in a standard listening room environment, inferred by the AI model. For example, the human voice is usually located in the center of the front, while the drum kit may be located in the rear. The music genre information can be the music genre and style tags of the audio data to be played, such as pop, rock, classical, jazz, etc., identified and output by the AI model. The channel playback energy ratio can be the ratio of the relative volume energy (i.e., audio playback energy) between each separated element in the audio data to be played within the current preset time segment, analyzed and output by the AI model.
[0034] As an example, when analyzing audio data, the AI model can determine the music style information based on the entire audio data, and it can also determine the channel playback energy ratio and sound image position coordinates of multiple mono audio streams within the current preset time segment, and then perform subsequent processing. The preset time segment can be determined by the AI model based on the time step of the audio data to be played, and is not specifically limited here.
[0035] In addition, AI models can be dedicated large models pre-trained in the cloud, such as audio understanding models based on Transformer and convolutional neural networks, which can be deployed on in-vehicle system-on-a-chip in practical applications.
[0036] In some embodiments of this application, a control data stream is generated based on the multi-dimensional metadata, and the control data stream and the plurality of mono audio streams are synchronously transmitted to a preset digital signal processor.
[0037] After acquiring multiple mono audio streams for the audio data to be processed, as well as multi-dimensional metadata associated with each mono audio stream, a control data stream can be generated based on the multi-dimensional metadata, and the control data stream and multiple mono audio streams can be synchronously transmitted to a preset digital signal processor.
[0038] Specifically, multiple mono audio streams and multi-dimensional metadata can be transmitted to the DSP digital signal processor through two independent and synchronous paths. For the generated 12-channel PCM audio stream (i.e., multiple mono audio streams), it can be directly transmitted to the input interface of the in-vehicle dedicated audio DSP via a high-speed serial audio bus, such as TDM (Time-Division Multiplexing). For the multi-dimensional metadata, the parsed metadata can be defined as corresponding control data streams (instruction sets) at the SOC (System on a Chip) level. These instruction sets can then be sent to the DSP in real-time and synchronously via Serial Peripheral Interface (SPI) commands. This ensures precise alignment between the mono audio streams and their corresponding spatial and style control parameters.
[0039] Step 103: Render and adjust the multiple mono audio streams based on the multi-dimensional metadata, and send the rendered and adjusted multiple mono audio streams to the corresponding speakers in the vehicle for playback.
[0040] In step 103, after the digital signal processor receives multiple mono audio streams and a control data stream, it can render and adjust the multiple mono audio streams based on the control data stream, and then send the rendered and adjusted mono audio streams to the corresponding speakers in the vehicle for playback. The control data stream and the multiple mono audio streams can be processed separately based on the preset time segments mentioned above; that is, the channel playback energy ratio and sound image position coordinates of the multiple mono audio streams in different preset time segments may be different.
[0041] In some embodiments of this application, the multi-dimensional metadata further includes music style information, and the rendering and tuning of the plurality of mono audio streams based on the multi-dimensional metadata includes: Sub-step 11: Based on the music style information, determine the target equalization curve of the audio data.
[0042] In sub-step 11, specifically, the 12 PCM audio streams, i.e. multiple mono audio streams, can first enter a global equalizer. This equalizer can also receive music style information from SPI instructions (i.e., control data streams) and can determine the target equalization curve of the audio data to be played from the internally stored or dynamically generated equalization curves corresponding to different styles based on the music style information.
[0043] Sub-step 12: Based on the target equalization curve, perform equalization processing on the multiple mono audio streams.
[0044] In sub-step 12, after determining the target equalization curve for the audio data to be played, the global equalizer can perform equalization processing on multiple mono audio streams based on the target equalization curve. Specifically, the global equalizer can be implemented through a multi-band parametric equalizer, whose key frequency points cover the audible range of 20Hz to 20kHz. It can then intelligently adjust the gain values of each frequency band according to the musical style information, for example, enhancing the low-frequency impact of rock music or strengthening the high-frequency extension of classical music. Furthermore, this function can be designed as an "AI intelligent sound effect" mode, operating transparently to the user, or it can provide limited personalized adjustments through the user interface.
[0045] In some embodiments of this application, after equalizing the plurality of mono audio streams based on the target equalization curve, the method further includes: Sub-step 21: Obtain the speaker layout information of the vehicle.
[0046] In sub-step 21, as described above, the multi-dimensional metadata includes the sound image position coordinates of multiple mono audio streams, and the generated control data stream also includes the sound image position coordinates of multiple mono audio streams. Therefore, the vehicle's speaker layout information can be obtained.
[0047] In practical applications, a digital signal processor may include a sound effect algorithm module, which contains a precise mathematical model of the speaker system layout in a vehicle, thereby obtaining the speaker layout information of the vehicle.
[0048] Sub-step 22: For each mono audio stream, determine the target speaker and speaker playback parameters for each mono audio stream based on the corresponding sound image position coordinates and the speaker layout information.
[0049] In sub-step 22, after obtaining the vehicle's speaker layout information, for each of the multiple mono audio streams, the target speaker and speaker playback parameters of each mono audio stream can be determined based on the sound image position coordinates corresponding to each mono audio stream and the speaker layout information.
[0050] Specifically, after the sound effect algorithm module receives the sound image position coordinates (X, Y, Z) of a mono audio stream in the current preset time segment, the localization algorithm is activated. This localization algorithm can calculate a gain matrix in real time. This gain matrix can accurately determine how to allocate the mono audio stream of the current preset time segment to the corresponding target speaker in the vehicle with the corresponding speaker playback parameters. The localization algorithm can be a vector basis amplitude translation algorithm.
[0051] As an example, the mono audio stream can be distributed to the physical speakers closest to the sound image's location coordinates. By coordinating the sound output of multiple speakers and utilizing the wavefront synthesis principle, a stable sound image from a specified three-dimensional coordinate can be accurately "drawn" in the listener's auditory perception.
[0052] It should be noted that, since the sound image position coordinates of the multiple mono audio streams of the audio data to be played in traditional stereo are in the ideal three-dimensional sound image position in a standard listening room environment, it is necessary to make corresponding sound image position adjustments for the space of the vehicle and the ideal position in the vehicle.
[0053] In some embodiments of this application, the speaker playback parameters include any one or more of the following: Amplitude adjustment parameters, phase adjustment parameters, and delay adjustment parameters.
[0054] Specifically, the speaker playback parameters can include amplitude adjustment parameters, phase adjustment parameters, and delay adjustment parameters. That is, the gain matrix can precisely determine how to allocate the mono audio stream to the corresponding target speaker in the vehicle with the corresponding amplitude adjustment parameters, phase adjustment parameters, and delay adjustment parameters.
[0055] In addition, since the ideal position for the listener in a vehicle is the driver's seat, the speaker playback parameters were determined by taking into account the physical distance between the speakers and the home location, in order to ensure that the sound waves from all speakers reach the listener's position, i.e., the driver's seat, simultaneously.
[0056] In some embodiments of this application, after determining the target speaker and speaker playback parameters for each mono audio stream based on the corresponding sound image position coordinates and the speaker layout information, the method further includes: Sub-step 31: Determine the audio playback energy of each mono audio stream.
[0057] In sub-step 31, after determining the target speaker and speaker playback parameters for each mono audio stream, the audio playback energy of each mono audio stream within the current preset time segment can also be determined.
[0058] Sub-step 32: Based on the audio playback energy of each mono audio stream, determine the dominant mono audio stream from the plurality of mono audio streams.
[0059] In sub-step 32, after determining the audio playback energy of each mono audio stream within the current preset time segment, the dominant mono audio stream of the current preset time segment can be determined from multiple mono audio streams based on the audio playback energy of each mono audio stream within the current preset time segment.
[0060] Specifically, the audio effect algorithm module in the digital signal processor performs multi-segment analysis on the audio playback energy of multiple mono audio streams within the current preset time segment in order to identify the dominant mono audio stream in the current preset time segment.
[0061] Sub-step 33: Adjust the audio playback energy of other mono audio streams according to the ratio of the audio playback energy of the dominant mono audio stream to the channel playback energy.
[0062] In sub-step 33, the audio effect algorithm module in the digital signal processor receives the channel energy ratio of multiple mono audio streams within the current preset time segment. After identifying the dominant mono audio stream in the current preset time segment, it can adjust the audio playback energy of other mono audio streams based on the audio playback energy of the dominant mono audio stream and the channel playback energy ratio.
[0063] Specifically, based on the channel energy ratio between each mono audio stream and the dominant mono audio stream, the gain of other mono audio streams can be dynamically and gently attenuated or slightly increased. This ensures that the overall output level remains stable when multiple mono audio streams are mixed, avoiding sudden over-loudness or under-loudness in some channels due to channel separation caused by the AI model. This provides a good basis for users to make unified adjustments using the vehicle's volume knob.
[0064] In some embodiments of this application, determining the dominant mono audio stream from the plurality of mono audio streams based on the audio playback energy of each mono audio stream includes: Sub-step 41: Determine the average playback energy of each mono audio stream based on the audio playback energy of each mono audio stream.
[0065] In sub-step 41, the audio effect algorithm module in the digital signal processor performs multi-segment analysis on the audio playback energy of multiple mono audio streams within the current preset time segment, that is, it determines the average playback energy of multiple mono audio streams within the current preset time segment based on the audio playback energy of multiple mono audio streams within the current preset time segment.
[0066] Sub-step 42: Based on the average playback energy of each mono audio stream, determine the dominant mono audio stream from the plurality of mono audio streams.
[0067] In sub-step 42, after determining the average playback energy of multiple mono audio streams within the current preset time segment, the mono audio stream with the highest average playback energy can be identified as the dominant mono audio stream from among the multiple mono audio streams.
[0068] In some embodiments of this application, sending the rendered and tuned multiple mono audio streams to corresponding speakers in the vehicle for playback includes: Based on the speaker playback parameters, the multiple mono audio streams are sent to their respective target speakers for playback.
[0069] Specifically, based on the speaker playback parameters corresponding to multiple mono audio streams within a preset time period, multiple mono audio streams can be sent to their respective target speakers for playback. For example, the mono audio stream of the "left front channel" can be assigned to the tweeter, midrange, and woofer units in the left front position.
[0070] In addition, when multiple mono audio streams are played through corresponding target speakers, high-pass and low-pass filters, eq (equalization / equalizer) can be used to correct the frequency response defects of the speaker unit itself, and limiter processing can be applied to the mono audio stream to prevent overload distortion and protect the speaker hardware.
[0071] Ultimately, all speakers in the vehicle are processed in parallel according to the above scheme, driving the entire speaker array to regenerate any stereo traditional audio data into an immersive panoramic sound field that adapts to the vehicle's space, matches the musical style, has distinct layers, and precise sound image positioning. This achieves precise sound image positioning tailored to the vehicle, comparable to or even surpassing that of a standard listening room. Moreover, since all mono channels have been separated in real time, the system or user can theoretically independently adjust the volume and equalization of any instrument, and even dynamically change its spatial position (for example, by swiping the guitar sound image from the front left to the rear right using a gesture UI).
[0072] In this embodiment, audio data to be played in a vehicle is acquired, and the audio data is parsed to obtain multiple mono audio streams and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. Based on the multi-dimensional metadata, the multiple mono audio streams are rendered and adjusted, and the rendered and adjusted multiple mono audio streams are sent to the corresponding speakers in the vehicle for playback. This realizes real-time panoramic sound conversion of traditional stereo audio data, breaks the dependence on pre-made panoramic sound audio data, and provides a highly intelligent and personalized in-vehicle immersive audio experience. It eliminates the need to pay high licensing fees for scarce panoramic sound content or engage in complex cooperation with content providers. It can provide a top-notch audio experience using existing music libraries, significantly reducing the threshold and cost of providing high-end audio functions.
[0073] Reference Figure 4 The diagram illustrates a flowchart of a vehicle audio processing method according to an embodiment of this application, which may include the following steps: 1. Acquire any stereo traditional audio data for the left and right channels.
[0074] 2. The AI model in the GPU / NPU processes the audio data to obtain multiple mono audio streams and multi-dimensional metadata, i.e., the elements and metadata in the figure.
[0075] 3. For multiple mono audio streams, they can be directly transmitted to the input interface of the vehicle-mounted audio DSP in TDM form via a high-speed serial audio bus, namely A2B (Automotive Audio Bus) in the diagram. For multi-dimensional metadata, the parsed multi-dimensional metadata can be defined as the corresponding control data stream (instruction set) through the MCU (Microcontroller Unit) of the SOC and transmitted through the CAN (Controller Area Network) bus. Thus, the instruction set can be sent to the DSP in real time and synchronously with multiple mono audio streams in the form of serial peripheral interface instruction SPI.
[0076] 4. GEQ (Graphic Equalizer) first performs global equalization on multiple mono audio streams, then processes them through an audio effect algorithm module to determine the target speakers and corresponding speaker playback parameters for the multiple mono audio streams. Based on the speaker playback parameters, the PA (Power Amplifier) sends the multiple mono audio streams to their respective target speakers for playback.
[0077] In this embodiment, audio data to be played in a vehicle is acquired, and the audio data is parsed to obtain multiple mono audio streams and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. Based on the multi-dimensional metadata, the multiple mono audio streams are rendered and adjusted, and the rendered and adjusted multiple mono audio streams are sent to the corresponding speakers in the vehicle for playback. This realizes real-time panoramic sound conversion of traditional stereo audio data, breaks the dependence on pre-made panoramic sound audio data, and provides a highly intelligent and personalized in-vehicle immersive audio experience. It eliminates the need to pay high licensing fees for scarce panoramic sound content or engage in complex cooperation with content providers. It can provide a top-notch audio experience using existing music libraries, significantly reducing the threshold and cost of providing high-end audio functions.
[0078] Reference Figure 5 This diagram illustrates a modular structure block diagram of a vehicle according to an embodiment of this application. The vehicle may specifically include the following modules: The audio acquisition module 501 is used to acquire audio data to be played in the vehicle; The audio processing module 502 is used to parse the audio data to obtain multiple mono audio streams of the audio data and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. The audio rendering module 503 is used to render and adjust the multiple mono audio streams based on the multi-dimensional metadata, and send the rendered and adjusted multiple mono audio streams to the corresponding speakers in the vehicle for playback.
[0079] In one embodiment of this application, the multi-dimensional metadata further includes music genre information, and the audio rendering module 504 includes: The curve determination module is used to determine the target equalization curve of the audio data based on the music style information. An equalization processing module is used to perform equalization processing on the multiple mono audio streams based on the target equalization curve.
[0080] In one embodiment of this application, after equalizing the plurality of mono audio streams based on the target equalization curve, the method further includes: A layout information acquisition module is used to acquire the speaker layout information of the vehicle; The playback parameter determination module is used to determine the target speaker and speaker playback parameters for each mono audio stream based on the corresponding sound image position coordinates and the speaker layout information.
[0081] In one embodiment of this application, after determining the target speaker and speaker playback parameters for each mono audio stream based on the corresponding sound image position coordinates and the speaker layout information, the method further includes: The energy determination module is used to determine the audio playback energy of each mono audio stream separately. A dominant channel determination module is used to determine the dominant mono audio stream from the plurality of mono audio streams based on the audio playback energy of each mono audio stream. The capability adjustment module is used to adjust the audio playback energy of other mono audio streams based on the ratio of the audio playback energy of the dominant mono audio stream to the channel playback energy.
[0082] In one embodiment of this application, the dominant channel determination module includes: The average energy calculation submodule is used to determine the average playback energy of each mono audio stream based on the audio playback energy of each mono audio stream. The dominant channel determination submodule is used to determine the dominant mono audio stream from the plurality of mono audio streams based on the average playback energy of each mono audio stream.
[0083] In one embodiment of this application, sending the rendered and tuned multiple mono audio streams to corresponding speakers in the vehicle for playback includes: The playback submodule is used to send the multiple mono audio streams to the corresponding target speakers for playback based on the speaker playback parameters.
[0084] In one embodiment of this application, the speaker playback parameters include any one or more of the following: Amplitude adjustment parameters, phase adjustment parameters, and delay adjustment parameters.
[0085] This application also provides an electronic device 60, please refer to... Figure 6 It includes a processor 601 and a memory 602, wherein the memory 601 is used to store computer programs; the processor 602 is used to execute the programs stored in the memory 601 to implement a method for processing in-vehicle audio as described in any embodiment of this application.
[0086] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a method for processing in-vehicle audio as described in any embodiment of this application.
[0087] In this embodiment, audio data to be played in a vehicle is acquired, and the audio data is parsed to obtain multiple mono audio streams and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. Based on the multi-dimensional metadata, the multiple mono audio streams are rendered and adjusted, and the rendered and adjusted multiple mono audio streams are sent to the corresponding speakers in the vehicle for playback. This realizes real-time panoramic sound conversion of traditional stereo audio data, breaks the dependence on pre-made panoramic sound audio data, and provides a highly intelligent and personalized in-vehicle immersive audio experience. It eliminates the need to pay high licensing fees for scarce panoramic sound content or engage in complex cooperation with content providers. It can provide a top-notch audio experience using existing music libraries, significantly reducing the threshold and cost of providing high-end audio functions.
[0088] In this application, "multiple" refers to two or more.
[0089] In this application, unless otherwise expressly defined, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0090] The terms “first,” “second,” “third,” “fourth,” etc., in this application (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0091] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0092] Unless otherwise specified, all steps in this application may be performed sequentially or randomly. For example, if the method includes steps A and B, it means that the method may include steps A and B performed sequentially, or it may include steps B and A performed sequentially. For example, if the method may also include step C, it means that step C may be added to the method in any order. For example, the method may include steps A, B, and C, or it may include steps A, C, and B, or it may include steps C, A, and B, etc.
[0093] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for processing in-vehicle audio, characterized in that, The method includes: Obtain the audio data to be played in the vehicle; The audio data is parsed to obtain multiple mono audio streams of the audio data and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. Based on the multi-dimensional metadata, the multiple mono audio streams are rendered and tuned, and the rendered and tuned multiple mono audio streams are sent to the corresponding speakers in the vehicle for playback.
2. The method according to claim 1, characterized in that, The multi-dimensional metadata also includes music style information, and the rendering and tuning of the multiple mono audio streams based on the multi-dimensional metadata includes: Based on the music style information, determine the target equalization curve of the audio data; Based on the target equalization curve, equalization processing is performed on the multiple mono audio streams.
3. The method according to claim 2, characterized in that, After equalizing the plurality of mono audio streams based on the target equalization curve, the process further includes: Obtain the speaker layout information of the vehicle; For each mono audio stream, the target speaker and speaker playback parameters are determined based on the corresponding sound image position coordinates and the speaker layout information.
4. The method according to claim 3, characterized in that, After determining the target speaker and speaker playback parameters for each mono audio stream based on the corresponding sound image position coordinates and the speaker layout information, the process further includes: Determine the audio playback energy for each mono audio stream; Based on the audio playback energy of each mono audio stream, the dominant mono audio stream is determined from the plurality of mono audio streams; The audio playback energy of other mono audio streams is adjusted based on the ratio of the audio playback energy of the dominant mono audio stream to the playback energy of the channel.
5. The method according to claim 4, characterized in that, Based on the audio playback energy of each mono audio stream, the dominant mono audio stream is determined from the plurality of mono audio streams, including: Based on the audio playback energy of each mono audio stream, determine the average playback energy of each mono audio stream; The dominant mono audio stream is determined from the plurality of mono audio streams based on the average playback energy of each mono audio stream.
6. The method according to claim 3, characterized in that, The step of sending the rendered and tuned multiple mono audio streams to the corresponding speakers in the vehicle for playback includes: Based on the speaker playback parameters, the multiple mono audio streams are sent to their respective target speakers for playback.
7. The method according to any one of claims 3-6, characterized in that, The speaker playback parameters include any one or more of the following: Amplitude adjustment parameters, phase adjustment parameters, and delay adjustment parameters.
8. A vehicle, characterized in that, The vehicles include: The audio acquisition module is used to acquire the audio data to be played in the vehicle; An audio processing module is used to parse the audio data to obtain multiple mono audio streams of the audio data and multi-dimensional metadata associated with the multiple mono audio streams. The multi-dimensional metadata includes at least the channel playback energy ratio of the multiple mono audio streams and the sound image position information of each mono audio stream. The channel playback energy ratio characterizes the audio playback energy ratio relationship between the multiple mono audio streams. The audio rendering module is used to render and adjust the multiple mono audio streams based on the multi-dimensional metadata, and send the rendered and adjusted multiple mono audio streams to the corresponding speakers in the vehicle for playback.
9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, which, when executed by a processor, implements the method as described in any one of claims 1 to 7.