Audio processing method and apparatus, and device, system, vehicle, medium and product

WO2026179598A1PCT designated stage Publication Date: 2026-09-03BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/076127
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2026-01-30
Publication Date
2026-09-03

Smart Images

  • Figure CN2026076127_03092026_PF_FP_ABST
    Figure CN2026076127_03092026_PF_FP_ABST
Patent Text Reader

Abstract

An audio processing method and apparatus, and a device, a system, a vehicle, a medium and a product. The method comprises: on the basis of an audio signal, determining a plurality of stem audio signals; and outputting each stem audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, equipment, systems, vehicles, media and products for audio processing

[0001] Cross-reference of related applications

[0002] This application claims priority to Chinese Patent Application No. 202510246299.7, filed on February 28, 2025, entitled “Method, Apparatus, Device, System, Vehicle, Medium and Product for Audio Processing”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of audio technology, and in particular to methods, apparatus, devices, systems, vehicles, media and products for audio processing. Background Technology

[0004] Currently, users have increasingly higher requirements for audio playback, but existing audio playback methods can only output the acquired audio signal directly, resulting in poor audio playback quality and making it difficult to meet user needs. Summary of the Invention

[0005] In view of the above problems, this application proposes methods, apparatus, devices, systems, vehicles, media and products for audio processing to improve audio playback effects and meet user needs.

[0006] Firstly, an audio processing method is provided, the method comprising:

[0007] Based on the audio signal, multiple track-specific audio signals are determined;

[0008] Each audio track is output separately.

[0009] In some embodiments of this application, each track audio signal is output separately, including:

[0010] Determine the spatial acoustic image position of each audio track;

[0011] Based on the spatial sound image location, each track audio signal is output separately.

[0012] In some embodiments of this application, determining the spatial acoustic image position of each track audio signal includes:

[0013] Determine the audio playback mode of the audio signal;

[0014] Based on the audio playback mode, determine the spatial sound image position of each audio track.

[0015] In some embodiments of this application, determining the spatial acoustic image position of each audio track signal according to the audio playback mode includes:

[0016] Based on the audio playback mode, determine the default position of each audio track and set the default position as the spatial sound image position of each audio track.

[0017] Alternatively, based on the audio playback mode, determine the historical record position of each audio track and set the historical record position as the spatial sound image position of each audio track.

[0018] In some embodiments of this application, it also includes:

[0019] In response to user input, the spatial image position of the split-track audio signals is adjusted.

[0020] In some embodiments of this application, adjusting the spatial image position of the track-by-track audio signals in response to user operation includes:

[0021] The corresponding object control for each audio track signal is displayed on the screen area corresponding to the spatial sound image position via the in-vehicle display screen.

[0022] In response to the user's movement of the object control, the spatial image position of the audio signal of the corresponding track is adjusted.

[0023] In some embodiments of this application, adjusting the spatial image position of the track-by-track audio signals in response to user operation includes:

[0024] In response to the user's voice control command, determine the track audio signal and target spatial sound image position indicated by the voice control command;

[0025] The spatial image positions of the segmented audio signals are adjusted according to the target spatial image position.

[0026] In some embodiments of this application, the audio playback mode includes any of the following: full vehicle mode, driver's seat mode, passenger seat mode, front row mode, rear row mode, left column mode, right column mode, and designated seat mode.

[0027] In some embodiments of this application, each track audio signal is output separately, including:

[0028] Based on the spatial acoustic image location, determine the target loudspeaker for each audio track;

[0029] Based on the target loudspeaker, each audio track is output separately.

[0030] In some embodiments of this application, determining the target loudspeaker for each track audio signal based on its spatial acoustic image location includes:

[0031] Based on each audio track, determine the distance between the positions of multiple speakers and the spatial sound image positions;

[0032] Based on distance, the target speaker for each audio track is determined from multiple speakers.

[0033] In some embodiments of this application, determining the target loudspeaker for each track audio signal from a plurality of loudspeakers based on distance includes:

[0034] According to the preset number of speakers, determine multiple speaker combinations from multiple speakers;

[0035] The target loudspeaker combination for each track audio signal is determined from multiple loudspeaker combinations based on the distance between the loudspeaker position and the spatial sound image position in each loudspeaker combination.

[0036] In some embodiments of this application, each audio track is output separately based on the target loudspeaker, including:

[0037] Each audio track is rendered to obtain a rendered audio track, and then each rendered audio track is output through a target loudspeaker.

[0038] In some embodiments of this application, each audio track is subjected to audio-visual rendering to obtain each audio track after audio-visual rendering, including:

[0039] Based on the distance between each audio track and the target speaker, the amplitude of each audio track is allocated to the audio signal emitted by the target speaker to obtain the audio track signal after sound image rendering.

[0040] In some embodiments of this application, the audio-visual rendering of each track audio signal is achieved using any of the following techniques: amplitude frequency shifting technique based on distance vector synthesis, or amplitude frequency shifting technique based on amplitude vector synthesis.

[0041] In some embodiments of this application, each audio track is output separately based on the target loudspeaker, including:

[0042] Based on each audio track signal, generate the control signal corresponding to the target speaker;

[0043] Based on the control signal, the target speaker is controlled to output audio.

[0044] In some embodiments of this application, before controlling the target speaker to output audio based on the control signal, the method further includes:

[0045] When multiple control signals exist for the target loudspeaker, the multiple control signals are synthesized.

[0046] In some embodiments of this application, a control signal corresponding to the target loudspeaker is generated based on each track audio signal, including:

[0047] Acquire in-vehicle audio signals;

[0048] Based on the in-vehicle audio signal and the audio signal of each individual track, a control signal corresponding to the target speaker is generated.

[0049] In some embodiments of this application, before generating the control signal corresponding to the target speaker based on the in-vehicle audio signal and each track audio signal, the method further includes:

[0050] Perform echo cancellation and / or noise reduction and / or howling suppression on the in-vehicle audio signal.

[0051] In some embodiments of this application, the split-track audio signal includes a first split-track audio signal and a second split-track audio signal, wherein the first split-track audio signal is a human voice audio signal and the second split-track audio signal is an accompaniment audio signal.

[0052] In some embodiments of this application, before outputting each track audio signal separately, the method further includes:

[0053] In response to user input, adjust the volume of the first audio track.

[0054] In some embodiments of this application, before determining multiple track-specific audio signals based on audio signals, the method further includes:

[0055] Get the preset number of channels and the number of channels of the audio signal;

[0056] If the number of channels in the audio signal is not the preset number, adjust the audio signal to use the preset number of channels.

[0057] In some embodiments of this application, the preset number of channels is two, with the two channels corresponding to the first and second track audio signals, respectively.

[0058] In some embodiments of this application, before determining multiple track-specific audio signals based on audio signals, the method further includes:

[0059] Obtain the preset sampling rate and the sampling rate of the audio signal;

[0060] If the audio signal's sampling rate is not the preset sampling rate, adjust the audio signal to use the preset sampling rate.

[0061] In some embodiments of this application, multiple track-specific audio signals are determined based on audio signals, including:

[0062] Based on the audio signal and sound object separation model, the audio signal is multitracked to obtain multiple separate audio signals.

[0063] In some embodiments of this application, before performing multitrack separation of the audio signal based on the audio signal and sound object separation model to obtain multiple track-separated audio signals, the method further includes:

[0064] A sound object separation model is established based on the synthetic database; the synthetic database is generated based on the acquired sample audio signals and the multitrack separation results of the sample audio signals.

[0065] In some embodiments of this application, an acoustic object separation model is established based on a synthetic database, including:

[0066] Based on the sample audio signals in the synthetic database and the multitrack separation results of the sample audio signals, the model parameters of the acoustic object separation model are adjusted to establish the acoustic object separation model.

[0067] In some embodiments of this application, the audio signal includes any of the following: audio signal in an in-vehicle application, audio signal in a storage medium, audio signal in a terminal device, audio signal in the cloud, and audio signal preset in an in-vehicle system.

[0068] Secondly, an audio processing apparatus is provided for implementing the above method.

[0069] Thirdly, an electronic device is provided, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the method described above.

[0070] Fourthly, an in-vehicle system is provided for implementing the above-described audio processing method, or including the above-described audio processing device, or including the above-described electronic device.

[0071] Fifthly, a vehicle is provided, including the audio processing device as described above, or the electronic device as described above, or the in-vehicle system as described above.

[0072] Sixthly, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method described above is implemented.

[0073] In a seventh aspect, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described above.

[0074] In this embodiment, multiple audio tracks are determined based on the audio signal, and each audio track is output separately. This achieves multi-track separation of the audio signal and outputs multiple audio tracks, thereby improving the audio playback effect and meeting user needs. Attached Figure Description

[0075] To more clearly illustrate the technical solution of this application, the drawings used in the description of this application will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0076] Figure 1 is a flowchart of the steps of an audio processing method provided in some embodiments of this application;

[0077] Figure 2 is a schematic diagram of an audio signal acquisition and preprocessing process provided in some embodiments of this application;

[0078] Figure 3 is a schematic diagram of a multi-track separation process provided in some embodiments of this application;

[0079] Figure 4a is a schematic diagram of an in-vehicle interactive interface provided in some embodiments of this application;

[0080] Figure 4b is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application;

[0081] Figure 4c is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application;

[0082] Figure 4d is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application;

[0083] Figure 4e is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application;

[0084] Figure 5 is a schematic diagram of a user adjustment process provided in some embodiments of this application;

[0085] Figure 6 is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application;

[0086] Figure 7 is a schematic diagram of a loudspeaker determination process provided in some embodiments of this application;

[0087] Figure 8 is a schematic diagram of an effect processing procedure provided by some embodiments of this application;

[0088] Figure 9 is a schematic diagram of a system signal processing procedure provided in some embodiments of this application;

[0089] Figure 10 is a flowchart of another audio processing method provided in some embodiments of this application. Detailed Implementation

[0090] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0091] In practical applications, from the perspective of a singer's live performance, singing takes place in the center of the stage, while the instrumental accompaniment is usually distributed around the singer. When singing, the singer can personally feel the feeling of being surrounded by stage instruments. Related technologies such as in-vehicle karaoke cannot provide the same karaoke experience as on a real stage.

[0092] Specifically, in-car karaoke and other similar functions rely solely on accompaniment and vocals, lacking the layered feel of different instruments arranged on a stage, thus failing to provide users with the experience of singing as a singer in the center of the stage.

[0093] With the increasing number of car speakers, it is possible to render the spatial sound image position of the instrument track, making it possible to achieve better karaoke effects.

[0094] In this embodiment, audio data from any audio source can be acquired, such as music apps, video apps, storage media, and music transmitted via Bluetooth. The in-vehicle computer is equipped with a music source separation algorithm, which can separate various audio tracks, such as bass, drums, piano, guitar, and vocals. During karaoke, the user can adjust the spatial sound image position of each audio track. Utilizing the multi-speaker layout in the vehicle, the sound images of different audio tracks are rendered in real-time to the user's set position, providing a completely new karaoke experience as if singing on a stage.

[0095] In a car environment, the relevant technology relies on third-party karaoke software for karaoke singing. The music library depends on the music accompaniment provider of the third-party software, and it cannot allow users to sing karaoke instantly from any music source, such as music software, video software, music videos, USB flash drives, etc.

[0096] To address this issue, this application provides a method for multitrack separation of audio data from any music within an in-vehicle processor. All sound objects can be used as accompaniment, allowing users to instantly perform karaoke with any music within the vehicle, breaking the limitations of traditional karaoke software. Furthermore, users can select different karaoke modes and customize the spatial position of the virtual sound images of each sound object, thereby obtaining an in-vehicle karaoke experience similar to performing on a stage.

[0097] In this application embodiment, the following specific contents are included:

[0098] 1. Collect and create a synthesis database, which consists of original music sounds, vocals, bass, drums, piano, violin, guitar and other objects, covering most commonly used instruments and types of music.

[0099] 2. Establish a deep learning-based sound object separation model, train it based on the above database, and after the model converges, it can be used to infer and extract the track-specific audio signals (i.e., accompaniment sound sources) from any music source.

[0100] 3. Based on the aforementioned segmented audio signals, users can select audio playback modes via voice control or button selection, such as full vehicle mode, driver's seat mode, passenger seat mode, front row mode, rear row mode, left row mode, right row mode, and designated seat mode. In practice, full vehicle mode is the default mode. The designated seat mode can be freely combined according to the user's selection.

[0101] Taking the karaoke full-vehicle mode as an example, the spatial sound image positions of the separated multi-object accompaniment sound sources are customized by the user, and then sound field rendering is performed. Specifically, by combining the virtual sound image positions of the sound objects adjusted by the user, the number, type, and layout of the speakers in the car, the sound image rendering algorithm is used to render the multi-object accompaniment sound sources in the spatial positions set by the karaoke user, so that the spatial sound images of multiple track audio signals surround the user, bringing the user a karaoke stage experience.

[0102] 4. For the human voice separated from the deep learning music source, the position of its sound image can also be adjusted by the user, and the volume of the human voice in the music source can be freely adjusted by the user, making it easier for the user to sing on stage with the original singer.

[0103] To address the aforementioned problems, this disclosure provides, in several embodiments, a method, apparatus, device, system, vehicle, medium, and product for audio processing. The following description, in conjunction with the accompanying drawings, provides an exemplary account of this application:

[0104] Referring to Figure 10, which illustrates a flowchart of another audio processing method provided in some embodiments of this application, the method may specifically include the following steps:

[0105] Step 1001: Based on the audio signal, determine multiple track-specific audio signals;

[0106] In some examples, the audio signal may include any of the following: audio signal in an in-vehicle application, audio signal in a storage medium, audio signal in a terminal device, audio signal in the cloud, or audio signal preset in an in-vehicle system.

[0107] Figure 2 is a schematic diagram of an audio signal acquisition and preprocessing process provided in some embodiments of this application. The in-vehicle application may include an in-vehicle music app and an in-vehicle video app. The in-vehicle system (201) can install the in-vehicle music app (202), allowing the system to acquire audio signals from music and music videos. The in-vehicle video app (203) can also be installed, allowing the video software to play music-related videos and acquire audio signals from them. Users can also acquire audio data from storage media such as USB flash drives, hard drives, and internal storage (204), and then read the audio signals through the in-vehicle system. Users can also connect to the in-vehicle system using mobile phones, tablets, or other terminal devices (205) via Bluetooth or wireless connections, sending the audio signals played by the terminal devices to the in-vehicle system, enabling the system to acquire the audio signals provided in the above ways. In some examples, the audio signals to be played by the user can also be obtained from a cloud server, or some audio signals can be pre-set on the in-vehicle system and input according to the user's selection.

[0108] After obtaining the audio signal, it can be processed by an algorithm, which mainly involves splitting the input audio signal into multiple tracks to obtain multiple split audio signals. These multiple split audio signals can be used for in-vehicle karaoke.

[0109] In some examples, the split audio signal includes a first split audio signal and a second split audio signal, where the first split audio signal is a vocal audio signal and the second split audio signal is an accompaniment audio signal.

[0110] In some embodiments of this application, determining multiple track-specific audio signals based on an audio signal includes: performing multitrack separation on the audio signal based on an audio signal and a sound object separation model to obtain multiple track-specific audio signals.

[0111] In practical applications, a sound object separation model can be established in advance, and then the input audio signal can be multitracked using the sound object separation model to obtain multiple track-separated audio signals.

[0112] In some embodiments of this application, before performing multitrack separation of the audio signal based on the audio signal and sound object separation model to obtain multiple track-separated audio signals, the method further includes:

[0113] A sound object separation model is established based on the synthetic database; the synthetic database is generated based on the acquired sample audio signals and the multitrack separation results of the sample audio signals.

[0114] As an example, the multitrack separation result signal of the sample audio signal can be multiple sample track audio signals obtained by multitrack separation of the sample audio signal.

[0115] In some embodiments of this application, establishing an acoustic object separation model based on a synthetic database includes: adjusting the model parameters of the acoustic object separation model based on sample audio signals in the synthetic database and the multitrack separation result signals of the sample audio signals, thereby establishing the acoustic object separation model.

[0116] Figure 3 is a schematic diagram of a multitrack separation process provided in some embodiments of this application. Step 301 collects human voice signals and accompaniment signals (i.e., multitrack separation result signals), establishes a signal library, and then synthesizes sample audio signals through this signal library. A synthesis database 302 (i.e., the synthesis database includes sample audio signals and multitrack separation results of sample audio signals) can be established to separate human voice and multitrack music sources. After establishing the synthesis database, a sound object separation model 303 can be established based on the synthesis database.

[0117] In establishing an acoustic object separation model, sample data from a synthesis database can be input into a pre-set acoustic object separation model. The acoustic object separation model is trained based on the sample data, and the parameters of the model are continuously optimized and adjusted to obtain a well-trained acoustic object separation model.

[0118] In some examples, the input and output of the model are both dual-channel signals. For example, the sound object separation model can be a network type such as CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), or Transformer (a neural network architecture based on self-attention mechanism). The algorithm model 303 is trained using the 302 synthesis database to meet the requirements of multi-track separation performance of human voice and music sources.

[0119] Any audio signal 304 can be processed by this model algorithm to obtain its first-track audio signal 305 (i.e., vocal signal) and second-track audio signal 306 (i.e., multi-track accompaniment signal). In some examples, the second-track audio signal 306 may contain bass, drums, piano, guitar, etc., and is not limited to these. Besides the separated objects and vocals, other musical components are assigned to other tracks. It should be noted that this model algorithm can process any audio signal to obtain vocals and multi-track accompaniment.

[0120] In some embodiments of this application, before determining multiple track-specific audio signals based on the audio signal, the method further includes: obtaining a preset number of channels and the number of channels of the audio signal; and adjusting the audio signal to use an audio signal with the preset number of channels when the number of channels of the audio signal is not the preset number of channels.

[0121] In some examples, the preset number of channels can be two, with the two channels corresponding to the first and second audio tracks, respectively.

[0122] Figure 2 is a schematic diagram of an audio signal acquisition and preprocessing process provided in some embodiments of this application. The preprocessing 206 mainly includes channel number unification and sampling rate conversion. If the audio signal acquired by the vehicle system is dual-channel, that is, two channels on the left and right, no channel preprocessing is performed. If the audio signal is single-channel, the single-channel signal is copied to make it dual-channel data. If the number of channels is greater than 2, such as 5.1 or 7.1 surround sound audio, or 7.1.4 Dolby Atmos audio, etc., the number of channels is greater than 2, and it is converted into 2-channel audio through downmixing.

[0123] In some embodiments of this application, before determining multiple track-specific audio signals based on the audio signal, the method further includes: obtaining a preset sampling rate and the sampling rate of the audio signal; and adjusting the audio signal to use the preset sampling rate if the sampling rate of the audio signal is not the preset sampling rate.

[0124] Figure 2 is a schematic diagram of an audio signal acquisition and preprocessing process provided by some embodiments of this application. If the sampling rate of the music is not the set sampling rate, it is upsampled or downsampled to make its sampling rate the preset sampling rate. Finally, in this preprocessing, a dual-channel audio signal 207 with the preset sampling rate is obtained. This preprocessing can accept audio signals of any format and any source.

[0125] Step 1002: Output each audio track separately.

[0126] In this embodiment, by determining multiple audio tracks based on the audio signal and outputting each audio track separately, the audio signal is separated into multiple tracks and outputs multiple audio tracks, thereby improving the audio playback effect and meeting user needs.

[0127] In some embodiments of this application, each track audio signal is output separately, including:

[0128] Sub-step 11: Determine the spatial image position of each audio track.

[0129] In some examples, the spatial sound image position can correspond to the spatial position inside the vehicle, used to indicate the position where the multi-track audio signal is played, such as the spatial sound image position indicating the driver's seat, passenger seat, front row position, rear row position, left column position, right column position, etc.

[0130] In some embodiments of this application, determining the spatial acoustic image position of each audio track includes: determining the audio playback mode of the audio signal; and determining the spatial acoustic image position of each audio track based on the audio playback mode.

[0131] As an example, the audio playback modes include any of the following: full vehicle mode, driver's seat mode, passenger seat mode, front row mode, rear row mode, left column mode, right column mode, and designated seat mode.

[0132] In practical applications, users can control and configure related settings, including setting switches, selecting different audio playback modes (such as karaoke mode), adjusting the spatial image position of the placed sound source object, and adjusting the volume.

[0133] After selecting the appropriate audio playback mode, the spatial sound image position of each audio track can be determined, and then input can be made according to the corresponding spatial sound image position.

[0134] Figure 4a is a schematic diagram of an in-vehicle interactive interface provided in some embodiments of this application. The UI interface of the in-vehicle tablet has a switch button (i.e., an audio playback indicator) 401 for in-vehicle stage karaoke. After turning on the switch, there are multiple karaoke mode (i.e., audio playback mode) selections 402, including full vehicle mode, driver's seat mode, passenger seat mode, front row mode, and rear row mode, which the user can select. The UI interface can set the audio playback indicator 401 and the audio playback mode 402. It can also display a layout diagram 403 of the corresponding spatial sound image position for the audio playback mode, and can also provide controls 408 for volume adjustment and controls 409 for adjusting the spatial sound image position.

[0135] Figure 4b is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application. The UI interface can display the layout diagram 404 of the spatial sound image position corresponding to the audio playback mode.

[0136] Figure 4c is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application. The UI interface can display the layout diagram 405 of the spatial sound image position corresponding to the audio playback mode.

[0137] Figure 4d is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application. The UI interface can display the layout diagram 406 of the spatial sound image position corresponding to the audio playback mode.

[0138] Figure 4e is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application. The UI interface can display the layout diagram 407 of the spatial sound image position corresponding to the audio playback mode.

[0139] The main effect of different audio playback modes is to make the spatial sound images surrounding different audio tracks inconsistent. For example, when the driver's seat mode is turned on and the whole car mode is turned on, the former will make the spatial sound images of the multi-track objects in the music source surround the driver's seat, while the passenger seat or other seats will not have this acoustic experience. In the whole car mode, the multi-track objects in the music source will surround the entire car, so that everyone in the car can feel the acoustic experience of being surrounded by sound objects.

[0140] In some embodiments of this application, determining the spatial acoustic image position of each audio track signal according to the audio playback mode includes:

[0141] Based on the audio playback mode, determine the default position of each audio track and set the default position as the spatial image position of each audio track; or, based on the audio playback mode, determine the historical record position of each audio track and set the historical record position as the spatial image position of each audio track.

[0142] Figure 5 is a schematic diagram of a user adjustment process provided by some embodiments of this application. User 501 selects an audio playback mode 502 (e.g., front row mode). First, it needs to determine if this is the user's first adjustment 503. When the user uses this function for the first time, if the multiple sound sources are bass, drums, piano, guitar, vocals, and others, the spatial sound image of the multiple sound sources is in a default position 504. This default position allows all the spatial sound images of the objects to be evenly distributed around the selected mode area. This could mean they are evenly distributed along a line directly in front of the selected mode area, or all sound images are concentrated in one spatial position. Then, the user can adjust the spatial position of the sound objects at the default position 505 to obtain the final adjusted spatial sound image position 506. The adjusted object can be recorded 507 as a historical record position for the user's next selection. After obtaining the spatial sound image position, a sound image rendering operation 508 can be performed on the first and second track audio signals, and a volume adjustment operation 509 can be performed on the second track audio signal. Then, multiple audio signals can be merged and output 510. In some embodiments of this application, it also includes:

[0143] In response to user input, the spatial image position of the split-track audio signals is adjusted.

[0144] In practical applications, to improve audio playback quality, users can adjust the spatial image position of the multi-track audio signals. As shown in Figure 4a, which is a schematic diagram of an in-vehicle interactive interface provided by some embodiments of this application, operation 409 is the user's sound object position adjustment function, which allows the user to arbitrarily change the default sound image position layout of the sound object space in the selected mode.

[0145] In some embodiments of this application, adjusting the spatial image position of the track-by-track audio signals in response to user operation includes:

[0146] The in-vehicle display screen shows the corresponding object control for each audio track on the screen area corresponding to the spatial sound image position; in response to the user's movement operation of the object control, the spatial sound image position of the audio track corresponding to the object control is adjusted.

[0147] Figure 6 is a schematic diagram of another in-vehicle interactive interface provided by some embodiments of this application. Taking the driver's seat mode as an example, if the multi-track separated objects (i.e., the separate audio signals) are guitar, bass, drums, piano, vocals, and others, the user can manually drag the spatial position of the instruments and source vocal objects in the screen UI. 601 shows an example of the default sound image position in the driver's seat mode that is adjusted by the user. In this default mode, the source vocal object is directly in front of the driver, the bass is on the right, the piano is on the right rear, the drums are on the left rear, the guitar is on the left, and the sound image positions of the other tracks are located at the driver's head. After the user adjusts and drags the spatial sound image positions of each track in step 602, the source vocal object is adjusted to the right front, the bass on the right is adjusted to a closer distance, the piano on the left rear is adjusted to a farther position, and the drums on the left and the guitar on the left rear are swapped, finally forming the spatial sound image position distribution of the multi-track objects of the in-vehicle stage karaoke sound source after user self-adjustment 603. Through the above operations, users can be provided with a user-customized acoustic experience of the instrument arrangement in a car stage karaoke. At the same time, the position adjustment and movement of the source vocals can also bring users a new karaoke experience of singing on the same stage as the original singer.

[0148] In some embodiments of this application, adjusting the spatial image position of the track-by-track audio signals in response to user operation includes:

[0149] In response to a user's voice control command, determine the track-specific audio signal and the target spatial sound image position indicated by the voice control command; adjust the spatial sound image position of the track-specific audio signal according to the target spatial sound image position.

[0150] In practical applications, to facilitate user operation, it can also receive voice control commands input by the user, identify the track audio signals and target spatial sound image positions indicated in the voice control commands, and then adjust the indicated track audio signals to the target spatial sound image positions.

[0151] Sub-step 12: Output each track audio signal separately according to the spatial sound image position.

[0152] After determining the spatial image position of each audio track, each audio track can be output separately at its spatial image position.

[0153] In some embodiments of this application, each track audio signal is output separately according to its spatial acoustic image position, including:

[0154] Sub-step 121: Determine the target loudspeaker for each track audio signal based on the spatial acoustic image position.

[0155] In practical applications, multiple speakers can be set up, and different speakers can be set up in different positions. Then, the target speaker for each audio track can be determined according to the spatial sound image position.

[0156] In some embodiments of this application, determining the target loudspeaker for each track audio signal based on its spatial acoustic image location includes:

[0157] Based on each audio track, determine the distance between the positions of multiple loudspeakers and the spatial sound image positions; based on the distance, determine the target loudspeaker for each audio track from among the multiple loudspeakers.

[0158] In some embodiments of this application, determining the target loudspeaker for each track audio signal from a plurality of loudspeakers based on distance includes:

[0159] According to the preset number of speakers, multiple speaker combinations are determined from multiple speakers; based on the distance between the speaker position in each speaker combination and the spatial sound image position, the target speaker combination for each track audio signal is determined from multiple speaker combinations.

[0160] After determining the spatial sound image location, a specific speaker combination needs to be selected and then rendering algorithms such as DBAP are used to generate multi-speaker signals for the specific sound image location. The principle for selecting the speaker combination is the proximity principle, that is, the speakers are selected based on the distance.

[0161] Taking the driver's seat mode two-dimensional plane adjustment as an example, as shown in Figure 7, Figure 7 is a schematic diagram of a speaker determination process provided by some embodiments of this application. On the two-dimensional plane, a Cartesian coordinate system o-xy is established with the position of the head as the origin, the right side as the horizontal axis, and the front as the vertical axis. If a mode with multiple seats is selected, such as the full vehicle mode, the coordinate system is established with the spatial center of the multiple seats as the origin. The plane calibration position of each speaker in the vehicle is (xi,yi), i = 1, 2, ..., 8, for a total of 8 speaker positions. If the spatial position of the speaker is not on the plane of this coordinate system, the projection point of the speaker's spatial position on this plane is taken as the calibration position.

[0162] If the user adjusts the piano position to (xt, yt), the line it lies on in this coordinate system is kx - y = 0, where k = yt / xt is the slope. During the rendering of the sound object's spatial position, the selected number of speakers is designed to be three. The specific speaker numbers are determined based on the principle of proximity to obtain the speaker combination. The calculation formula is as follows: by obtaining the three speakers that minimize the distance dis, this combination is used to render the sound object. Then, the DBAP algorithm is used to render the sound object, generating control signals for the three speakers.

[0163] Sub-step 122: Based on the target loudspeaker, output each track audio signal separately.

[0164] In some embodiments of this application, each audio track is output separately based on the target loudspeaker, including:

[0165] Each audio track is rendered to obtain a rendered audio track, and then each rendered audio track is output through a target loudspeaker.

[0166] In some embodiments of this application, each audio track is subjected to audio-visual rendering to obtain each audio track after audio-visual rendering, including:

[0167] Based on the distance between each audio track and the target speaker, the amplitude of each audio track is allocated to the audio signal emitted by the target speaker to obtain the audio track signal after sound image rendering.

[0168] In some examples, the audio-visual rendering of each track is achieved using either of the following techniques: amplitude shifting based on distance vector synthesis or amplitude shifting based on amplitude vector synthesis.

[0169] In practical applications, each audio track can be rendered with sound images, as shown in Figure 5. Figure 5 is a schematic diagram of a user adjustment process provided by some embodiments of this application. After determining the position of each accompaniment object, the rendering operation 508 of the human voice and sound object is performed to render each sound object to the spatial position adjusted by the user. The rendering method may include amplitude panning (VBAP) based on amplitude vector synthesis and amplitude panning (DBAP) based on distance vector synthesis. The latter spatial sound image rendering method is not limited by the number of speakers and has higher flexibility. In some examples, this method can be used for spatial sound image rendering in some embodiments of this application.

[0170] DBAP technology can be based on the principle of distance vector synthesis, which calculates the distance between the sound object and each speaker to determine the amplitude of the sound signal that each speaker should emit.

[0171] In some embodiments of this application, each audio track is output separately based on the target loudspeaker, including:

[0172] Based on each track audio signal, a control signal corresponding to the target speaker is generated; based on the control signal, the target speaker is controlled to output audio.

[0173] In some embodiments of this application, before controlling the target speaker to output audio based on the control signal, the method further includes: when there are multiple control signals for the target speaker, performing signal synthesis on the multiple control signals.

[0174] After rendering the sound images of all sound objects, as shown in Figure 5, each speaker will synthesize and add all the signals assigned to it to obtain a composite signal 510 of multiple audio tracks, such as the accompaniment track of a car stage karaoke.

[0175] For example, with 4 speakers, there are 2 sound objects (i.e., control signals for the track-by-track audio signals). The spatial sound image rendering of sound object 1 is composed of speakers 1, 2, and 3, resulting in three sets of signals x11, x12, and x13. The spatial sound image of sound object 2 is composed of speakers 2, 3, and 4, with signals x22, x23, and x24 respectively. Finally, speaker 1 only outputs the x11 signal, speaker 2 outputs x12+x22, speaker 3 outputs x13+x23, and speaker 4 outputs the x24 signal. The signals from these 4 speakers constitute the accompaniment track.

[0176] In some embodiments of this application, a control signal corresponding to the target loudspeaker is generated based on each track audio signal, including:

[0177] Acquire the in-vehicle audio signal; based on the in-vehicle audio signal and the audio signal of each track, generate the control signal corresponding to the target speaker.

[0178] In some embodiments of this application, before generating the control signal corresponding to the target speaker based on the in-vehicle audio signal and each track audio signal, the method further includes: performing echo cancellation processing and / or noise reduction processing and / or howling suppression processing on the in-vehicle audio signal.

[0179] In practical applications, the sound output effects of karaoke are processed, mainly including the spatial image rendering of different audio tracks, echo cancellation, noise reduction, and reverberation processing. The processed karaoke audio signals are then output through a power amplifier and a car speaker system.

[0180] Figure 8 is a schematic diagram of an effect processing process provided by some embodiments of this application. The vehicle microphone 801 collects data such as the user's voice, echo, and noise inside the vehicle. The echo cancellation 802 is used to eliminate the music echo signal generated by the music played by the vehicle speaker. At the same time, the noise reduction 803 is used to reduce the noise inside and outside the vehicle that interferes with the user's voice. Finally, the user's voice with the echo and noise removed is obtained. The user's voice is mixed with the composite signal 804 of multiple audio signals (such as the backing track of a car stage karaoke) to obtain the car karaoke audio 805. After being processed by the reverb unit 806, it is finally output to the speaker 807 for playback.

[0181] In some embodiments of this application, the split audio signal includes a first split audio signal and a second split audio signal. Before outputting each split audio signal separately, the method further includes: adjusting the volume of the first split audio signal in response to a user operation.

[0182] In practical applications, volume adjustment controls can be set up so that users can adjust the volume of the first audio track by operating the controls. Of course, the volume of the first audio track can also be adjusted by voice control, as shown in Figure 4a. Operation 408 can adjust the output volume of the human voice (i.e., the first audio track) of the music source, so that the original vocals can be accompanied by the user's karaoke singing.

[0183] The present application will be described by way of example below with reference to the accompanying drawings:

[0184] Figure 1 is a flowchart of an audio processing method provided by some embodiments of this application. Step 101: Acquire audio signal; Step 102: Perform multitrack separation processing through a multitrack separation algorithm; Step 103: The user makes relevant settings (such as adjusting the spatial sound image position and adjusting the volume) through selection and control; Step 104: Sound image rendering and effect processing; Step 105: Output through power amplifier and speaker.

[0185] Figure 9 is a schematic diagram of a system signal processing procedure provided in some embodiments of this application. User 901 can obtain source music signals 903 from different audio signal sources 902. The user can obtain different in-vehicle stage karaoke modes and adjust different spatial sound image positions of sound objects through UI / voice control 905 in the vehicle host 904. When the user sings, their voice is collected by the in-vehicle microphone 906, and after digital-to-analog conversion 907, it enters the algorithm processing module. The algorithm processing 908 includes a sound object separation model, rendering algorithm 909, echo cancellation, howling suppression, noise reduction algorithm 910, and reverberation algorithm 911. Finally, the karaoke audio is played through a power amplifier and the in-vehicle multi-channel speaker system 912.

[0186] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.

[0187] Some embodiments of this application also provide an in-vehicle audio processing apparatus, which can be specifically used for:

[0188] Based on the audio signal, multiple track-specific audio signals are determined;

[0189] Each audio track is output separately.

[0190] In some embodiments of this application, each track audio signal is output separately, including:

[0191] Determine the spatial acoustic image position of each audio track;

[0192] Based on the spatial sound image location, each track audio signal is output separately.

[0193] In some embodiments of this application, determining the spatial acoustic image position of each track audio signal includes:

[0194] Determine the audio playback mode of the audio signal;

[0195] Based on the audio playback mode, determine the spatial sound image position of each audio track.

[0196] In some embodiments of this application, determining the spatial acoustic image position of each audio track signal according to the audio playback mode includes:

[0197] Based on the audio playback mode, determine the default position of each audio track and set the default position as the spatial sound image position of each audio track.

[0198] Alternatively, based on the audio playback mode, determine the historical record position of each audio track and set the historical record position as the spatial sound image position of each audio track.

[0199] In some embodiments of this application, it also includes:

[0200] In response to user input, the spatial image position of the split-track audio signals is adjusted.

[0201] In some embodiments of this application, adjusting the spatial image position of the track-by-track audio signals in response to user operation includes:

[0202] The corresponding object control for each audio track signal is displayed on the screen area corresponding to the spatial sound image position via the in-vehicle display screen.

[0203] In response to the user's movement of the object control, the spatial image position of the audio signal of the corresponding track is adjusted.

[0204] In some embodiments of this application, adjusting the spatial image position of the track-by-track audio signals in response to user operation includes:

[0205] In response to the user's voice control command, determine the track audio signal and target spatial sound image position indicated by the voice control command;

[0206] The spatial image positions of the segmented audio signals are adjusted according to the target spatial image position.

[0207] In some embodiments of this application, the audio playback mode includes any of the following: full vehicle mode, driver's seat mode, passenger seat mode, front row mode, rear row mode, left column mode, right column mode, and designated seat mode.

[0208] In some embodiments of this application, the split-track audio signal includes a first split-track audio signal and a second split-track audio signal, wherein the first split-track audio signal is a human voice audio signal and the second split-track audio signal is an accompaniment audio signal.

[0209] In some embodiments of this application, before outputting each track audio signal separately, the method further includes:

[0210] In response to user input, adjust the volume of the first audio track.

[0211] In some embodiments of this application, before determining multiple track-specific audio signals based on audio signals, the method further includes:

[0212] Get the preset number of channels and the number of channels of the audio signal;

[0213] If the number of channels in the audio signal is not the preset number, adjust the audio signal to use the preset number of channels.

[0214] In some embodiments of this application, the preset number of channels is two, with the two channels corresponding to a first split-track audio signal and at least one second split-track audio signal, respectively.

[0215] In some embodiments of this application, before determining multiple track-specific audio signals based on audio signals, the method further includes:

[0216] Obtain the preset sampling rate and the sampling rate of the audio signal;

[0217] If the audio signal's sampling rate is not the preset sampling rate, adjust the audio signal to use the preset sampling rate.

[0218] In some embodiments of this application, each track audio signal is output separately, including:

[0219] Based on the spatial acoustic image location, determine the target loudspeaker for each audio track;

[0220] Based on the target loudspeaker, each audio track is output separately.

[0221] In some embodiments of this application, determining the target loudspeaker for each track audio signal based on its spatial acoustic image location includes:

[0222] Based on each audio track, determine the distance between the positions of multiple speakers and the spatial sound image positions;

[0223] Based on distance, the target speaker for each audio track is determined from multiple speakers.

[0224] In some embodiments of this application, determining the target loudspeaker for each track audio signal from a plurality of loudspeakers based on distance includes:

[0225] According to the preset number of speakers, determine multiple speaker combinations from multiple speakers;

[0226] The target loudspeaker combination for each track audio signal is determined from multiple loudspeaker combinations based on the distance between the loudspeaker position and the spatial sound image position in each loudspeaker combination.

[0227] In some embodiments of this application, each audio track is output separately based on the target loudspeaker, including:

[0228] Each audio track is rendered to obtain a rendered audio track, and then each rendered audio track is output through a target loudspeaker.

[0229] In some embodiments of this application, each audio track is subjected to audio-visual rendering to obtain each audio track after audio-visual rendering, including:

[0230] Based on the distance between each audio track and the target speaker, the amplitude of each audio track is allocated to the audio signal emitted by the target speaker to obtain the audio track signal after sound image rendering.

[0231] In some embodiments of this application, the audio-visual rendering of each track audio signal is achieved using any of the following techniques: amplitude frequency shifting technique based on distance vector synthesis, or amplitude frequency shifting technique based on amplitude vector synthesis.

[0232] In some embodiments of this application, each audio track is output separately based on the target loudspeaker, including:

[0233] Based on each audio track signal, generate the control signal corresponding to the target speaker;

[0234] Based on the control signal, the target speaker is controlled to output audio.

[0235] In some embodiments of this application, before controlling the target speaker to output audio based on the control signal, the method further includes:

[0236] When multiple control signals exist for the target loudspeaker, the multiple control signals are synthesized.

[0237] In some embodiments of this application, a control signal corresponding to the target loudspeaker is generated based on each track audio signal, including:

[0238] Acquire in-vehicle audio signals;

[0239] Based on the in-vehicle audio signal and the audio signal of each individual track, a control signal corresponding to the target speaker is generated.

[0240] In some embodiments of this application, before generating the control signal corresponding to the target speaker based on the in-vehicle audio signal and each track audio signal, the method further includes:

[0241] Perform echo cancellation and / or noise reduction and / or howling suppression on the in-vehicle audio signal.

[0242] In some embodiments of this application, multiple track-specific audio signals are determined based on audio signals, including:

[0243] Based on the audio signal and sound object separation model, the audio signal is multitracked to obtain multiple separate audio signals.

[0244] In some embodiments of this application, before performing multitrack separation of the audio signal based on the audio signal and sound object separation model to obtain multiple track-separated audio signals, the method further includes:

[0245] A sound object separation model is established based on the synthetic database; the synthetic database is generated based on the acquired sample audio signals and the multitrack separation results of the sample audio signals.

[0246] In some embodiments of this application, an acoustic object separation model is established based on a synthetic database, including:

[0247] Based on the sample audio signals in the synthetic database and the multitrack separation results of the sample audio signals, the model parameters of the acoustic object separation model are adjusted to establish the acoustic object separation model.

[0248] In some embodiments of this application, the audio signal includes any of the following: audio signal in an in-vehicle application, audio signal in a storage medium, audio signal in a terminal device, audio signal in the cloud, and audio signal preset in an in-vehicle system.

[0249] Some embodiments of this application also provide an audio processing apparatus for implementing the audio processing method described above.

[0250] Some embodiments of this application also provide an electronic device, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the method described above.

[0251] Some embodiments of this application also provide an in-vehicle system for implementing the above-described audio processing method, or including the above-described audio processing device, or including the above-described electronic device.

[0252] Some embodiments of this application also provide a vehicle including the audio processing device as described above, or the electronic device as described above, or the in-vehicle system as described above.

[0253] Some embodiments of this application also provide a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described above.

[0254] Some embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the method described above.

[0255] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0256] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0257] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0258] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0259] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0260] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0261] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0262] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0263] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the aforementioned element.

[0264] The above provides a detailed description of the audio processing methods, apparatus, devices, systems, vehicles, media, and products. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. An audio processing method, wherein, The method includes: Based on the audio signal, multiple track-specific audio signals are determined; Each of the aforementioned audio tracks is output separately.

2. The method according to claim 1, wherein, The step of outputting each of the separate audio tracks includes: Determine the spatial acoustic image position of each of the said segmented audio signals; Based on the spatial acoustic image position, each of the separate audio tracks is output.

3. The method according to claim 2, wherein, Determining the spatial acoustic image position of each of the said track audio signals includes: Determine the audio playback mode of the audio signal; Based on the audio playback mode, the spatial sound image position of each of the audio tracks is determined.

4. The method according to claim 3, wherein, Determining the spatial sound image position of each of the audio tracks according to the audio playback mode includes: Based on the audio playback mode, determine the default position of each of the audio tracks, and set the default position as the spatial sound image position of each of the audio tracks. Alternatively, based on the audio playback mode, determine the historical record position of each of the audio tracks, and set the historical record position as the spatial sound image position of each of the audio tracks.

5. The method according to any one of claims 2-4, wherein, Also includes: In response to user operation, the spatial image position of the split-track audio signal is adjusted.

6. The method according to claim 5, wherein, The adjustment of the spatial image position of the split-track audio signal in response to user operation includes: The corresponding object control for each of the separate audio tracks is displayed on the screen area corresponding to the spatial sound image position via the vehicle-mounted display screen. In response to the user's movement operation on the object control, the spatial sound image position of the audio signal of the track corresponding to the object control is adjusted.

7. The method according to claim 5, wherein, The adjustment of the spatial image position of the split-track audio signal in response to user operation includes: In response to a user-inputted voice control command, determine the track-specific audio signal and the target spatial sound image position indicated by the voice control command; The spatial sound image position of the split-track audio signal is adjusted according to the target spatial sound image position.

8. The method according to claim 3 or 4, wherein, The audio playback modes include any of the following: full vehicle mode, driver's seat mode, passenger seat mode, front row mode, rear row mode, left column mode, right column mode, and designated seat mode.

9. The method according to any one of claims 2-8, wherein, The step of outputting each of the separate audio tracks includes: Based on the spatial acoustic image position, determine the target loudspeaker for each of the segmented audio signals; Based on the target loudspeaker, each of the separate audio tracks is output.

10. The method according to claim 9, wherein, The step of determining the target loudspeaker for each of the segmented audio signals based on the spatial acoustic image position includes: Based on each of the said segmented audio signals, the distance between the positions of multiple speakers and the spatial sound image positions is determined; Based on the distance, the target loudspeaker for each of the plurality of loudspeakers is determined from the plurality of loudspeakers.

11. The method according to claim 10, wherein, The step of determining the target loudspeaker for each of the plurality of loudspeakers for each of the segmented audio signals based on the distance includes: According to a preset number of speakers, determine multiple speaker combinations from the plurality of speakers; The target loudspeaker combination for each of the multiple loudspeaker combinations is determined from the plurality of loudspeaker combinations based on the distance between the loudspeaker position in each loudspeaker combination and the spatial sound image position.

12. The method according to any one of claims 9-11, wherein, The step of outputting each of the separate audio tracks based on the target loudspeaker includes: Each of the audio tracks is rendered to obtain a rendered audio track, and each rendered audio track is output through the target loudspeaker.

13. The method according to claim 12, wherein, The step of performing audio-visual rendering on each of the said segmented audio signals to obtain each of the said segmented audio signals after audio-visual rendering includes: Based on the distance between each of the audio tracks and the target speaker, the amplitude of each audio track is assigned to the audio signal emitted by the target speaker to obtain the audio track signal after sound image rendering.

14. The method according to claim 12, wherein, The audio-visual rendering of each of the said track audio signals is achieved using any of the following techniques: amplitude frequency shifting technique based on distance vector synthesis, or amplitude frequency shifting technique based on amplitude vector synthesis.

15. The method according to any one of claims 9-11, wherein, The step of outputting each of the separate audio tracks based on the target loudspeaker includes: Based on each of the said separate audio tracks, a control signal corresponding to the target loudspeaker is generated; Based on the control signal, the target speaker is controlled to output audio.

16. The method according to claim 15, wherein, Before controlling the target speaker to output audio based on the control signal, the method further includes: When multiple control signals exist for the target loudspeaker, the multiple control signals are synthesized.

17. The method according to claim 15, wherein, The step of generating a control signal corresponding to the target loudspeaker based on each of the segmented audio signals includes: Acquire in-vehicle audio signals; Based on the in-vehicle audio signal and each of the separate audio tracks, the control signal corresponding to the target speaker is generated.

18. The method according to claim 17, wherein, Before generating the control signal corresponding to the target speaker based on the in-vehicle audio signal and each of the separate audio tracks, the method further includes: The in-vehicle audio signal is subjected to echo cancellation and / or noise reduction and / or howling suppression processing.

19. The method according to any one of claims 1-8, wherein, The split-track audio signal includes a first split-track audio signal and a second split-track audio signal. The first split-track audio signal is a human voice audio signal, and the second split-track audio signal is an accompaniment audio signal.

20. The method according to claim 19, wherein, Before outputting each of the separate audio tracks individually, the method further includes: In response to user input, the volume of the first audio track is adjusted.

21. The method according to claim 19, wherein, Before determining multiple track audio signals based on the audio signals, the method further includes: Obtain the preset number of channels and the number of channels of the audio signal; If the number of channels in the audio signal is not a preset number, the audio signal is adjusted to use the preset number of channels.

22. The method according to claim 21, wherein, The preset number of channels is two, with the two channels corresponding to the first track audio signal and the second track audio signal, respectively.

23. The method according to any one of claims 1-8, wherein, Before determining multiple track audio signals based on the audio signals, the method further includes: Obtain the preset sampling rate and the sampling rate of the audio signal; If the sampling rate of the audio signal is not a preset sampling rate, the audio signal is adjusted to use the preset sampling rate.

24. The method according to any one of claims 1-8, wherein, The determination of multiple track-specific audio signals based on audio signals includes: Based on the audio signal and sound object separation model, the audio signal is multitracked to obtain multiple separate audio signals.

25. The method according to claim 24, wherein, Before performing multitrack separation of the audio signal based on the audio signal and sound object separation model to obtain multiple track-separated audio signals, the method further includes: A sound object separation model is established based on a synthetic database; the synthetic database is generated based on the acquired sample audio signals and the multitrack separation results of the sample audio signals.

26. The method according to claim 25, wherein, The establishment of an acoustic object separation model based on a synthetic database includes: Based on the sample audio signals in the synthetic database and the multitrack separation result signals of the sample audio signals, the model parameters of the acoustic object separation model are adjusted to establish the acoustic object separation model.

27. The method according to any one of claims 1-8, wherein, The audio signal includes any of the following: the audio signal in an in-vehicle application, the audio signal in a storage medium, the audio signal in a terminal device, the audio signal in the cloud, or the audio signal preset in the in-vehicle system.

28. An audio processing apparatus, wherein, Used to implement the method as described in any one of claims 1-27.

29. An electronic device, wherein, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1-27.

30. A vehicle-mounted system, wherein, Used to implement the audio processing method as described in any one of claims 1-27, or includes the audio processing apparatus as described in claim 28, or includes the electronic device as described in claim 29.

31. A vehicle, wherein, This includes the audio processing device as described in claim 28, or the electronic device as described in claim 29, or the vehicle system as described in claim 30.

32. A computer-readable storage medium, wherein, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the method as described in any one of claims 1-27.

33. A computer program product, wherein, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-27.