Audio processing method and related apparatus, system and storage medium

By employing audio separation and upmixing strategies, combined with precise speaker deployment, the problem of insufficient spatial sense and presence in single-channel mixed signals of vehicle audio data is solved, resulting in a richer audio experience.

CN122120692APending Publication Date: 2026-05-29IFLYTEK CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2026-02-05
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In vehicle driving scenarios, audio data itself is a single-channel mixed signal, which makes it difficult to improve the sense of space and presence.

Method used

Several first audio channels of different types are obtained through audio separation. An upmixing strategy matching the first channel type is selected to perform upmixing on the first audio. Then, a speaker whose device properties match the second channel type is selected in the three-dimensional sound space for playback.

Benefits of technology

It increases the number of audio channels and the panoramic sound channel mapping in three-dimensional sound space, enhancing the sense of space and presence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120692A_ABST
    Figure CN122120692A_ABST
Patent Text Reader

Abstract

The application discloses an audio processing method and related device, system and storage medium, wherein the audio processing method comprises: performing audio separation based on to-be-played audio to obtain first audio of a plurality of first channel types; selecting an upmix strategy matched with the first channel type as a target strategy of the first channel type based on the first channel type to which the first audio belongs; performing upmix operation on the first audio based on the target strategy of the first channel type to obtain second audio of a second channel type subordinate to the first channel type; and selecting a loudspeaker with a device attribute matched with the second channel type to which the second audio belongs as a target loudspeaker of the second audio in a three-dimensional sound space. The above scheme can realize optimized processing of to-be-played audio, so as to break through the limitation of audio data itself and as far as possible improve the sense of space and the sense of presence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method and related apparatus, system and storage medium. Background Technology

[0002] Besides its widespread popularity in professional fields such as vocal music, the audio listening experience has also gained increasing attention in many ordinary scenarios in recent years. For example, in the context of driving and riding in a vehicle, the needs of drivers and passengers for in-car audio experience have shifted from simply listening to music to an immersive experience.

[0003] Currently, in scenarios such as driving and riding in vehicles, although multiple speakers are typically installed inside the space to create a good listening experience, the audio data itself is usually a single-channel mixed signal. Therefore, it is difficult to improve the sense of space and presence due to the limitations of the audio data itself. In view of this, how to optimize the playback audio to overcome the limitations of the audio data itself and maximize the sense of space and presence has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide an audio processing method and related apparatus, system, and storage medium that can optimize the processing of audio to be played, thereby overcoming the limitations of the audio data itself to maximize the sense of space and presence.

[0005] To address the aforementioned technical problems, the first aspect of this application provides an audio processing method, comprising: performing audio separation based on the audio to be played to obtain a plurality of first audio of a first channel type; selecting an upmixing strategy matching the first channel type as a target strategy for the first channel type based on the first channel type to which the first audio belongs; performing an upmixing operation on the first audio based on the target strategy for the first channel type to obtain a second audio of a second channel type under the first channel type; selecting a loudspeaker in a three-dimensional sound space whose device attributes match the second channel type to which the second audio belongs as a target loudspeaker for the second audio; wherein, the device attributes include at least the deployment location; and playing the second audio based on each target loudspeaker of the second audio.

[0006] To address the aforementioned technical problems, a second aspect of this application provides an audio processing apparatus, comprising: an audio separation module, a strategy selection module, an audio upmixing module, a device selection module, and an audio playback module. The audio separation module is used to perform audio separation based on the audio to be played, obtaining a plurality of first audio channels of a first audio type. The strategy selection module is used to select an upmixing strategy matching the first channel type as the target strategy for the first channel type. The audio upmixing module is used to perform upmixing operations on the first audio based on the target strategy for the first channel type, obtaining a second audio channel of a second channel type belonging to the first channel type. The device selection module is used to select, in a three-dimensional sound space, a loudspeaker whose device attributes match the second channel type of the second audio as the target loudspeaker for the second audio; wherein the device attributes include at least the deployment location. The audio playback module is used to play the second audio based on each target loudspeaker of the second audio.

[0007] To address the aforementioned technical problems, a third aspect of this application provides an electronic device comprising at least a memory and a processor coupled to each other, wherein the memory stores at least program instructions, and the processor executes the program instructions to implement the audio processing method of the first aspect described above.

[0008] To address the aforementioned technical problems, a fourth aspect of this application provides a vehicle audio system, comprising: a control device and speakers deployed at different locations in the vehicle, wherein the control device is electrically connected to each speaker, and the control device includes the electronic device described in the third aspect above.

[0009] To address the aforementioned technical problems, the fifth aspect of this application provides a computer-readable storage medium storing program instructions executable by a processor, the program instructions being used to implement the audio processing method of the first aspect described above.

[0010] The above scheme performs audio separation based on the audio to be played, obtaining several first audio channels of different types. Based on the first channel type of each first audio channel, it selects an upmixing strategy that matches the first channel type as the target strategy. The first audio channels are then upmixed using this target strategy to obtain second audio channels of different types. This allows for the selection of speakers in a three-dimensional sound space whose device attributes match the second channel type of the second audio channel, serving as the target speakers for the second audio channel. The device attributes at least include the deployment location. Finally, the second audio channel is played based on each target speaker. This approach, by separating the audio to be played, allows for the acquisition of first audio channels of different types. The audio signal is processed using an upmixing strategy that matches the first channel type of the first audio signal. This upmixing operation yields a second audio signal of a second channel type under the first channel type. Therefore, a single-channel mixed audio signal can be separated and expanded into more refined second audio signals of various second channel types under different first channel types, which helps increase the number of channels for the audio to be played. Furthermore, after obtaining the second audio signals of each second channel type, a speaker whose device properties match the second channel type of the second audio signal is selected in the three-dimensional sound space as the target speaker for playing the second audio signal. This allows for panoramic channel mapping in the three-dimensional sound space as much as possible, improving the distinction between the depth and layering of each sound element. Therefore, optimized processing of the audio to be played can be achieved, overcoming the limitations of the audio data itself to maximize the sense of space and presence. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating an embodiment of the audio processing method of this application; Figure 2 This is a schematic diagram of an embodiment of the audio processing method of this application; Figure 3 This is a schematic diagram of the framework of an embodiment of the audio processing apparatus of this application; Figure 4 This is a schematic diagram of the framework of an embodiment of the electronic device of this application; Figure 5 This is a schematic diagram of the framework of an embodiment of the vehicle audio system of this application; Figure 6 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0012] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0013] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0014] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the slash " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper indicates two or more objects.

[0015] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the audio processing method of this application. It should be noted that the process operations in this embodiment can be executed by an electronic device with computing capabilities or related equipment containing an electronic device. The specific structure and type of the electronic device and related equipment containing the electronic device are not limited herein. Specifically, this embodiment may include the following steps: Step S11: Perform audio separation based on the audio to be played to obtain several first audio channels.

[0016] In this embodiment, the specific content of the audio to be played can vary depending on the actual application scenario. For example, in an in-vehicle scenario, the audio to be played may include, but is not limited to, local offline music or streaming music, live radio broadcast audio, etc. Alternatively, in a home scenario, the audio to be played may include, but is not limited to, local offline music or streaming music, live TV playback audio, etc. Please refer to the relevant documentation. Figure 2 , Figure 2 This is a schematic diagram illustrating an embodiment of the audio processing method of this application. Figure 2 As shown, several first channel types may include, but are not limited to: vocals, drums, bass, and others (such as keyboards and other instrument sounds). The latter three can also be collectively referred to as background music. It should be noted that in practical applications, the audio to be played may contain all of the above first channel types (such as...). Figure 2 The audio to be played may be a mixture of vocals, bass, drums, and keyboard sounds. However, it's also possible that not all of the above first channel types are present, but only one or a few. In this case, the first audio of the corresponding first channel type can be empty. For example, in a car setting, if the audio to be played is a cappella music, audio separation will only yield the first audio of the first channel type "vocals." Of course, the above examples are merely a few possible cases of the audio to be played and the first channel type in practical applications. Other possible scenarios regarding the audio to be played and the first channel type are not limited here, nor will they be listed one by one.

[0017] In one implementation scenario, as a possible approach, to achieve audio separation based on the audio to be played, a blind source separation algorithm can be used to perform audio separation on the audio to be played, obtaining several first audio channels. It should be noted that blind source separation (BSS) is a type of signal processing technique that recovers independent original source signals from a mixed signal when only the mixed observed signal is known, and the source signals and mixed system parameters are unknown. For example, blind source separation algorithms may include, but are not limited to, Independent Component Analysis (ICA), Principal Component Analysis (PCA), Non-negative Matrix Factorization (NFF), Sparse Component Analysis (SCA), etc. The specific types of blind source separation algorithms are not limited here, nor will they be listed individually.

[0018] In another implementation scenario, as a different possible approach, distinct from the aforementioned implementation, to perform audio separation based on the audio to be played and obtain several first audio tracks of the first channel type, data conversion can be performed first on the audio to be played to obtain its time spectrum. Then, the time spectrum is segmented to obtain several spectral blocks belonging to different sub-frequency bands. Based on these spectral blocks in each sub-frequency band, prediction can be made to determine the first channel type to which the spectral blocks belong. Finally, waveform recovery can be performed based on spectral blocks belonging to the same first channel type to obtain the corresponding first audio track of the first channel type. This method, by segmenting the time spectrum to obtain several spectral blocks belonging to different sub-frequency bands, and predicting the spectral blocks in each sub-frequency band to determine the first channel type to which the spectral blocks belong, and finally combining spectral blocks belonging to the same first channel type for waveform recovery to obtain the corresponding first audio track of the first channel type, can further refine the processing granularity to the sub-frequency band spectral blocks during audio separation, thus improving the precision of audio separation.

[0019] In a specific implementation scenario, after obtaining the audio to be played, it can be processed using methods such as STFT (Short-Time Fourier Transform) to achieve data transformation and obtain the time-frequency spectrum of the audio. Alternatively, as a possible example, to reduce the number of parameters, the time-frequency spectrum can be mapped to a lower dimension, such as by mapping it based on complex convolution. Of course, the above examples are merely a few possible examples of data transformation; other possible implementation methods are not limited here, nor will they be listed in detail.

[0020] In a specific implementation scenario, after obtaining the time spectrum of the audio to be played, it can be segmented to obtain several spectral blocks belonging to different sub-frequency bands. Specifically, the time spectrum can be segmented based on each preset frequency band interval to obtain sub-frequency bands of each preset frequency band interval. Then, based on the cutting interval matching the preset frequency band interval to which the sub-frequency band belongs, the sub-frequency bands are divided into blocks to obtain several spectral blocks of the sub-frequency bands. It should be noted that the cutting interval can be positively correlated with the representative frequency of the preset frequency band interval. That is, the higher the representative frequency of the preset frequency band interval (e.g., the frequency value at the left end of the interval, the frequency value at the right end of the interval, the frequency value at the center of the interval, etc.), the larger the cutting interval can be; conversely, the lower the representative frequency of the preset frequency band interval, the smaller the cutting interval can be. For ease of understanding, please refer to Table 1, which is a schematic table of an embodiment of the preset frequency band interval and its cutting interval.

[0021] Table 1. Schematic diagram of a preset frequency band interval and its cutting interval in one embodiment.

[0022] As shown in Table 1, the preset frequency bands can include: 0-1kHz, 1kHz-4kHz, 4kHz-8kHz, 8kHz-16kHz, and 16kHz-24kHz. Among them, the cutting interval of the preset frequency band 0-1kHz is 100Hz (that is, a spectrum block is divided every 100Hz in the sub-band "0-1kHz"), and the cutting interval of the preset frequency band 1kHz-4kHz is 250Hz (that is, a spectrum block is divided every 250Hz in the sub-band "1kHz-4kHz"). The preset frequency band interval of 4kHz-8kHz is divided into 500Hz segments (i.e., a spectral block is divided every 500Hz in the sub-band "4kHz-8kHz"); the preset frequency band interval of 8kHz-16kHz is divided into 1kHz segments (i.e., a spectral block is divided every 1kHz in the sub-band "8kHz-16kHz"); and the preset frequency band interval of 16kHz-24kHz is divided into 2kHz segments (i.e., a spectral block is divided every 2kHz in the sub-band "16kHz-24kHz"). Of course, the above examples are merely one possible illustration of the preset frequency band intervals and their division intervals. Other possible settings are not limited here, nor will they be listed in detail.

[0023] In a specific implementation scenario, after obtaining several spectral blocks belonging to different sub-frequency bands, prediction can be made based on the different spectral blocks in each sub-frequency band to obtain the first channel type to which the spectral block belongs. As one possible example, feature extraction networks such as Transformers can be used to extract features from different spectral blocks in each sub-frequency band, obtaining the spectral features of the spectral blocks. Then, a multi-classification network containing layers such as fully connected layers can be used to classify and predict the spectral features of the spectral blocks, obtaining the probability values ​​of each spectral block belonging to each first channel type. Based on the probability values ​​of each spectral block belonging to each first channel type, the first channel type to which the spectral block belongs can be determined (e.g., the first channel type corresponding to the highest probability value can be selected as the first channel type to which the spectral block belongs). Alternatively, as another possible example, several sets of feature modeling (e.g., one set of feature modeling, two sets of feature modeling, or three or more sets of feature modeling) can be performed based on several feature blocks in each sub-frequency band to obtain the spectral features of different spectral blocks in each sub-frequency band. It should be noted that each set of feature modeling can include: first performing intra-band modeling in the time domain dimension, and then performing inter-band modeling in the frequency domain dimension. In the time domain, intra-band modeling can be performed by arranging spectral blocks within the same sub-band in chronological order, treating them as time-domain sequence data. This time-domain sequence data is then processed sequentially through a regularization layer, a bidirectional long short-term memory (LSTM) unit, and a fully connected layer to obtain the spectral features of the spectral blocks in the time domain. Similarly, in the frequency domain, inter-band modeling can be performed by arranging spectral blocks according to their frequency order, treating them as frequency-domain sequence data. This frequency-domain sequence data is then processed sequentially through a regularization layer, a bidirectional LSM unit, and a fully connected layer to obtain the spectral features of the spectral blocks in the frequency domain. When performing a new set of feature modeling, the spectral features of the spectral blocks output from the inter-band modeling in the previous set can be used as input data for the intra-band modeling in the new set. For example, the spectral features of the spectral blocks within the same sub-band output from the inter-band modeling in the previous set can be arranged according to their chronological order, treating them as new time-domain sequence data. This process can be repeated to complete multiple sets of feature modeling. Based on this, classification predictions can be performed separately based on the spectral features of different spectral blocks in each sub-band to obtain the first channel type to which the spectral block belongs (e.g., using the aforementioned multi-classification network containing fully connected layers for classification prediction; see the aforementioned description for details). This approach, combining intra-band modeling in the time domain and inter-band modeling in the frequency domain during feature modeling, ensures that spectral features not only contain their own feature information but also feature information from related spectral blocks in both the time and frequency domains, thus improving the accuracy of classification predictions.

[0024] In a specific implementation scenario, after determining the first channel type of a spectrogram block, waveform reconstruction can be performed based on spectrogram blocks belonging to the same first channel type to obtain the first audio of the corresponding first channel type. Specifically, spectrogram blocks belonging to the same first channel type can be arranged and combined in order (e.g., chronological order), and then waveform reconstruction models such as ISTFT (Inverse Short-Time Fourier Transform) and vocoders can be used to recover the waveform, thus obtaining the first audio of the corresponding first channel type. For example, taking the first channel type "human voice" as an example, spectrogram blocks belonging to the first channel type "human voice" can be sorted in chronological order, and then waveform reconstruction models such as ISTFT (Inverse Short-Time Fourier Transform) and vocoders can be used to recover the waveform, thus obtaining the first audio of the first channel type "human voice". Of course, the above example is only one possible example of waveform reconstruction in practical applications, and other possible implementation methods are not limited here, nor will they be listed one by one.

[0025] In one implementation scenario, audio separation can be achieved by an audio separation model. This model could include a feature extraction network for extracting spectral features, a multi-classification network for classification prediction, etc. The network structure of the audio separation model is not limited here. To improve the separation accuracy of the audio separation model, sample mixed audio can be pre-acquired, and the sample mixed audio can be labeled with sample track audio of each first channel type. The audio separation model can then be trained and optimized based on this. Specifically, audio separation can be performed on the sample mixed audio based on the audio separation model to obtain predicted track audio of each first channel type. Then, based on the difference between the sample track audio of the same first channel type and the predicted track audio, the temporal loss of the corresponding first channel type is obtained. Based on the difference between the sample track audio of the same first channel type and the predicted track audio in the real and imaginary parts of the spectrum, the frequency domain loss of the corresponding first channel type is obtained. The temporal and frequency domain losses of the same first channel type can then be fused to obtain the total loss. Finally, the network parameters of the audio separation model can be adjusted based on the total loss. The above method combines the temporal loss in the time domain and the frequency loss in the frequency domain during model training to obtain the total loss, and adjusts the parameters of the audio separation model accordingly. This can avoid phase blurring and other problems as much as possible in the frequency domain, and avoid distortion caused by phase error during waveform recovery as much as possible in the time domain, which helps to improve the separation accuracy of the audio separation model.

[0026] In a specific implementation scenario, the detailed process of performing audio separation on mixed audio samples based on the audio separation model can be found in the aforementioned description of audio separation, and will not be repeated here. It should be noted that, in the process of performing audio separation on mixed audio samples using the audio separation model, the first channel type to which different sample spectral blocks in each sample sub-band of the mixed audio samples belong can be determined first. Then, the waveforms of spectral blocks belonging to the same first channel type are recovered to obtain the predicted track-separated audio of the corresponding first channel type.

[0027] In a specific implementation scenario, for ease of description, the total loss L obj It can be represented as:

[0028] In the above formula, S represents the complex spectrum of the sample track audio. S represents the complex spectrum of the predicted track-by-track audio. r This represents the real part of the spectrum of the sample track audio. S represents the real part of the spectrum of the predicted track-by-track audio. i This represents the imaginary part of the spectrum of the sample track audio. Let ||1| represent the imaginary part of the spectrum of the predicted track-by-track audio, ISTFT denotes the inverse Fourier transform, and ||1| denotes the L1 norm. Of course, the above example is merely one way to measure the total loss in practical applications; other possible measures are not limited here, nor will they be listed in detail.

[0029] Step S12: Based on the first channel type to which the first audio belongs, select an upmixing strategy that matches the first channel type as the target strategy for the first channel type.

[0030] Step S13: Perform upmixing on the first audio based on the target strategy of the first channel type to obtain the second audio of the second channel type under the first channel type.

[0031] It should be noted that for any first audio, an upmixing strategy that matches its first channel type can be selected firstly as the target strategy for its first channel type. Then, the first audio can be upmixed according to this target strategy to obtain the second audio of the second channel type under its first channel type.

[0032] In one implementation scenario, in response to the first audio channel being of the human voice type, the first audio can be expanded to obtain second audio with second channel types of master vocal channel and secondary vocal channel respectively (e.g., the first audio belonging to the first channel type "human voice" can be copied into three copies, one for the master vocal channel and two for the secondary vocal channel). Low-frequency attenuation is performed on the second audio with the second channel type of master vocal channel (to provide sufficient spectral space for subsequent bass and kick drum tuning), and time difference adjustment is performed on each of the second audio with the second channel type of secondary vocal channel (to enhance stereo separation and avoid phase problems as much as possible). The pitch of each audio is offset according to different cents (e.g., when there are two second audios with the second channel type of secondary vocal channel, the pitch can be offset by +5 cents and -5 cents respectively) to produce a smooth and elegant chorus effect.

[0033] In one implementation scenario, in response to the fact that the first audio channel belongs to bass, low-frequency attenuation can be performed based on the first audio (in order to reduce the overall muddiness of the mix and avoid low-frequency chaos as much as possible) to obtain a second audio with the second channel type bass.

[0034] In one implementation scenario, in response to the first audio channel being of type "drum sound", the first audio can be expanded to obtain second audio channels of type "main rhythm channel", "sub-rhythm channel", and "cymbal channel" respectively (e.g., the first audio channel of type "drum sound" can be copied into five copies, which are respectively used as one main rhythm channel, two sub-rhythm channels, and two cymbal channels). High-frequency attenuation is performed on the second audio channel of type "main rhythm channel", low-cut and high-cut operations are performed on the second audio channel of type "sub-rhythm channel" (in order to preserve the mid-frequency sound), and high-pass filtering and time difference adjustment are performed on the second audio channel of type "cymbal channel".

[0035] In one implementation scenario, in response to the first audio file belonging to a first channel type of "other," mid-low frequency attenuation, mid-high frequency attenuation, high-cut operation, and low-cut operation are performed on the first audio file to obtain second audio files with second channel types of mid-low frequency attenuation channel, mid-high frequency attenuation channel, high-cut channel, and low-cut channel, respectively. For example, the first audio file with a first channel type of "other" can be copied four times. One copy undergoes mid-low frequency attenuation to obtain a second audio file with a second channel type of "mid-low frequency attenuation channel," one copy undergoes mid-high frequency attenuation to obtain a second audio file with a second channel type of "mid-high frequency attenuation channel," one copy undergoes high-cut operation to obtain a second audio file with a second channel type of "high-cut channel," and one copy undergoes low-cut operation to obtain a second audio file with a second channel type of "low-cut channel." In addition, the above four second audio files can also undergo different degrees of pitch shift and time difference adjustment to change the tonality and rhythm and create a more spatial stereo sound field.

[0036] It should be noted that the above upmixing strategies are merely examples of possible scenarios where the first channel type includes vocals, bass, drums, and others. As mentioned earlier, when the first channel type includes vocals, the second channel type under vocals can include, but is not limited to, the main vocal channel and the secondary vocal channel; when the first channel type includes drums, the second channel type under drums can include, but is not limited to, the main rhythm channel, the secondary rhythm channel, and the cymbal channel; when the first channel type includes others, the other second channel types can include, but are not limited to, the mid-low frequency attenuation channel, the mid-high frequency attenuation channel, the high-cut channel, and the low-cut channel. Other possible upmixing strategies are not limited here, nor will they be listed one by one.

[0037] Furthermore, pitch shifting during audio upmixing can be implemented using audio processing frameworks such as SoundTouch. This allows for the stretching or shortening of the audio to be shifted in the time domain, achieving pitch and speed shifting. For ease of understanding, a brief explanation of pitch shifting operations is given below using SoundTouch as an example. For detailed implementation information, please refer to the technical details of audio processing frameworks such as SoundTouch; these will not be elaborated upon here. For instance, the audio to be shifted, x, can be divided into short frames x. m Each frame can contain N samples, H a This represents the data after framing, and can be represented as follows:

[0038] Based on this, the aforementioned short frames can be relocated to the time domain, and the input data can be actually modified by the stretching ratio α, which can be expressed as: α=H s / H a In the above formula, Hs H a These represent the overlap settings for the combined frame data and the individual frame data, respectively. Furthermore, to improve the phase continuity and amplitude fluctuation stability of frame boundaries, a Hamming window can be used to window the individual frames before reconstruction to form a combined frame. m Based on this, in order to reconstruct the modified output signal y, the adjusted composite frame y can be superimposed. m The final output signal is obtained:

[0039] It should be noted that the above description is merely one possible example of implementing pitch shifting using the SoundTouch audio processing framework. For detailed information on the SoundTouch audio processing framework, please refer to its technical specifications, which will not be elaborated upon here. Furthermore, in practical applications, the use of other audio processing frameworks for pitch shifting is excluded, and they will not be listed here.

[0040] Step S14: Select a speaker whose device properties match the type of the second channel of the second audio in the three-dimensional sound space, and use it as the target speaker for the second audio.

[0041] In this embodiment of the disclosure, the device attributes include at least the deployment location. For ease of understanding, taking an in-vehicle scenario as an example, the speaker deployment location may include, but is not limited to: center, rear, left, right, front left, front right, etc., and the speaker deployment location is not limited here. For example, for ease of naming, a speaker deployed in the "center" location can be called a center speaker, a speaker deployed in the "rear" location can be called a rear ring speaker, a speaker deployed on the "left" location can be called a left ring speaker, a speaker deployed on the "right" location can be called a right ring speaker, a speaker deployed on the "front left" location can be called a front left speaker, and a speaker deployed on the "front right" location can be called a front right speaker. In addition, the device attributes may also include the playback frequency (e.g., mid-bass, bass, treble, mid-high frequency, etc.). Taking an in-vehicle scenario as an example again, a speaker deployed in the "center" location with a playback frequency of "bass" can be called a center woofer. Of course, the above examples are only a few possible examples of in-vehicle speakers in an in-vehicle scenario, and other possible situations are not limited here, nor will they be listed one by one. To facilitate understanding, the following example, using an in-vehicle scenario, illustrates the selection of matching speakers for the second audio channels of various types in a three-dimensional sound space: In one implementation scenario, if the second channel of the second audio is the master vocal channel, the center speaker can be selected as the target speaker for the second audio. If the second channel of the second audio is the secondary vocal channel, the rear surround speakers can be selected as the target speakers for the second audio. This allows the vocal channel to present a surround sound effect, making the vocal imaging three-dimensional and enhancing the sense of depth.

[0042] In one implementation scenario, in response to the fact that the second channel of the second audio is of type bass, a center woofer can be selected as the target speaker for the second audio. This allows the bass channel, which has undergone low-frequency attenuation, to be emitted by the center woofer to provide bass support, thereby enhancing the richness and depth of the music.

[0043] In one implementation scenario, in response to the second channel type of the second audio being a main rhythm channel, a center speaker can be selected as the target speaker for the second audio. In response to the second channel type of the second audio being a sub-rhythm channel, a left loop speaker and a right loop speaker can be selected as the target speakers for the second audio. In response to the second channel type of the second audio being a cymbal channel, a sky speaker can be selected as the target speaker for the second audio.

[0044] In one implementation scenario, in response to the second audio channel being any of the following types: low-frequency attenuation channel, mid-high frequency attenuation channel, high-cut channel, or low-cut channel, the left front speaker, right front speaker, left surround speaker, and right surround speaker are selected as the target speakers.

[0045] It should be noted that the above method, through the relocalization of the sound from each channel in the three-dimensional sound space, enables the sound from each channel to have distinct depth and layers, a strong sense of lateral sound field immersion, thus creating a standard immersive space for pop music, with a clear vertical sound image perception, forming a high-intensity sound field structure. Of course, in other scenarios, the above description can be used as a reference to select the appropriate speaker for the second audio channel type. Examples of different scenarios will not be provided here.

[0046] Step S15: Play the second audio based on the target speakers of each second audio.

[0047] Specifically, after selecting a speaker with matching device attributes as the target speaker for each second audio channel type, the selected target speaker can be driven to play the second audio. Furthermore, in practical applications, when playing a real-time audio stream, the real-time audio stream can be segmented, each segment used as audio to be played, and the operations described above in this embodiment can be executed to achieve panoramic channel mapping of the real-time audio stream. For details, please refer to the foregoing overall description; further elaboration is not provided here.

[0048] The above scheme performs audio separation based on the audio to be played, obtaining several first audio channels of different types. Based on the first channel type of each first audio channel, it selects an upmixing strategy that matches the first channel type as the target strategy. The first audio channels are then upmixed using this target strategy to obtain second audio channels of different types. This allows for the selection of speakers in a three-dimensional sound space whose device attributes match the second channel type of the second audio channel, serving as the target speakers for the second audio channel. The device attributes at least include the deployment location. Finally, the second audio channel is played based on each target speaker. This approach, by separating the audio to be played, allows for the acquisition of first audio channels of different types. The audio signal is processed using an upmixing strategy that matches the first channel type of the first audio signal. This upmixing operation yields a second audio signal of a second channel type under the first channel type. Therefore, a single-channel mixed audio signal can be separated and expanded into more refined second audio signals of various second channel types under different first channel types, which helps increase the number of channels for the audio to be played. Furthermore, after obtaining the second audio signals of each second channel type, a speaker whose device properties match the second channel type of the second audio signal is selected in the three-dimensional sound space as the target speaker for playing the second audio signal. This allows for panoramic channel mapping in the three-dimensional sound space as much as possible, improving the distinction between the depth and layering of each sound element. Therefore, optimized processing of the audio to be played can be achieved, overcoming the limitations of the audio data itself to maximize the sense of space and presence.

[0049] Please see Figure 3 , Figure 3 This is a schematic diagram of a framework of an embodiment of the audio processing apparatus of this application. The audio processing apparatus 30 includes: an audio separation module 31, a strategy selection module 32, an audio upmixing module 33, a device selection module 34, and an audio playback module 35. The audio separation module 31 is used to perform audio separation based on the audio to be played, to obtain a number of first audio channels of a first type; the strategy selection module 32 is used to select an upmixing strategy that matches the first channel type based on the first channel type of the first audio, as the target strategy for the first channel type; the audio upmixing module 33 is used to perform upmixing operation on the first audio based on the target strategy of the first channel type, to obtain a second audio of a second channel type under the first channel type; the audio pitch shifting module 33 is used for; the device selection module 34 is used to select a loudspeaker whose device attributes match the second channel type of the second audio in a three-dimensional sound space, as the target loudspeaker for the second audio; wherein, the device attributes include at least the deployment location; the audio playback module 35 is used to play the second audio based on the target loudspeakers of each second audio.

[0050] In the above scheme, the audio processing device 30 performs audio separation based on the audio to be played, obtaining several first audios of different first channel types. Based on the first channel type to which the first audio belongs, it selects an upmixing strategy matching the first channel type as the target strategy for that first channel type. Based on the target strategy for the first channel type, it performs upmixing on the first audios to obtain second audios of different second channel types belonging to the first channel type. This allows it to select speakers in the three-dimensional sound space whose device attributes match the second channel type to which the second audio belongs, serving as the target speakers for the second audios. The device attributes at least include the deployment location. Then, based on each target speaker for the second audio, the second audio is played. This method, by separating the audio to be played, allows for the acquisition of different first channel types... The system firstly analyzes the audio signal and then uses an upmixing strategy that matches the first channel type of the first audio signal to perform an upmixing operation on the first audio signal to obtain a second audio signal of a second channel type under the first channel type. This allows for the separation and expansion of a single-channel mixed audio signal into more refined second audio signals of various second channel types under different first channel types, helping to increase the number of channels for the audio to be played. Furthermore, after obtaining the second audio signals of each second channel type, the system further selects speakers whose device properties match the second channel type of the second audio signal in the three-dimensional sound space as target speakers for playing the second audio signal. This enables panoramic sound channel mapping as much as possible in the three-dimensional sound space, helping to improve the distinction between the depth and layering of each sound part. Therefore, it is possible to optimize the audio to be played, overcoming the limitations of the audio data itself to maximize the sense of space and presence.

[0051] In some disclosed embodiments, the audio separation module 31 includes a data conversion submodule for performing data conversion based on the audio to be played to obtain the time spectrum of the audio to be played; the audio separation module 31 includes a spectrum segmentation submodule for segmenting based on the time spectrum to obtain several spectrum blocks belonging to different sub-frequency bands; the audio separation module 31 includes a classification prediction submodule for predicting based on different spectrum blocks in each sub-frequency band to obtain the first channel type to which the spectrum block belongs; and the audio separation module 31 includes a waveform recovery submodule for performing waveform recovery based on spectrum blocks belonging to the same first channel type to obtain the first audio of the corresponding first channel type.

[0052] In some disclosed embodiments, the classification prediction submodule includes a feature modeling unit, which is used to perform several sets of feature modeling based on several feature blocks of each sub-frequency band to obtain the spectral features of different spectral blocks in each sub-frequency band; wherein, each set of feature modeling includes: first performing intra-band modeling in the time domain dimension, and then performing inter-band modeling in the frequency domain dimension; the classification prediction submodule includes a type prediction unit, which is used to perform classification prediction based on the spectral features of different spectral blocks in each sub-frequency band to obtain the first channel type to which the spectral block belongs.

[0053] In some disclosed embodiments, the spectrum segmentation submodule includes a first segmentation unit, used to segment the time spectrum based on each preset frequency band interval to obtain sub-frequency bands of each preset frequency band interval; the spectrum segmentation submodule includes a second segmentation unit, used to divide the sub-frequency band into blocks based on a cutting interval that matches the preset frequency band interval to which the sub-frequency band belongs, to obtain several spectrum blocks of the sub-frequency band; wherein, the cutting interval is positively correlated with the representative frequency of the preset frequency band interval.

[0054] In some disclosed embodiments, audio separation is achieved by an audio separation model. The audio processing device 30 includes a sample acquisition module for acquiring sample mixed audio; wherein, the sample mixed audio is labeled with sample track audio of each first channel type; the audio processing device 30 includes a sample separation module for performing audio separation on the sample mixed audio based on the audio separation model to obtain predicted track audio of each first channel type; the audio processing device 30 includes a loss measurement module for obtaining the temporal loss of the corresponding first channel type based on the difference between the sample track audio of the same first channel type and the predicted track audio, and obtaining the frequency domain loss of the corresponding first channel type based on the difference between the sample track audio of the same first channel type and the predicted track audio in the real part and imaginary part of the spectrum, respectively; the audio processing device 30 includes a loss fusion module for fusing the temporal loss and frequency domain loss of the same first channel type to obtain the total loss; the audio processing device 30 includes a parameter adjustment module for adjusting the network parameters of the audio separation model based on the total loss.

[0055] In some disclosed embodiments, the audio upmixing module 32 includes a vocal upmixing submodule, which is used to expand the first audio based on the first audio in response to the first channel type being vocal, to obtain second audio with second channel types being the main vocal channel and the secondary vocal channel, and to perform low-frequency attenuation on the second audio with the second channel type being the main vocal channel, and to perform time difference adjustment on each of the second audio with the second channel type being the secondary vocal channel and to perform offset operation on their respective pitches according to different cents.

[0056] In some disclosed embodiments, the audio upmixing module 32 includes a bass upmixing submodule, which, in response to the first audio being of a first channel type of bass, performs low-frequency attenuation based on the first audio to obtain a second audio of a second channel type of bass.

[0057] In some disclosed embodiments, the audio upmixing module 32 includes a drum upmixing submodule, which is used to expand the first audio based on the first audio in response to the first channel type being drum sound, to obtain second audio with second channel types being main rhythm channel, sub-rhythm channel, and cymbal channel, and to perform high-frequency attenuation on the second audio with the second channel type being main rhythm channel, to perform low-cut operation and high-cut operation on the second audio with the second channel type being sub-rhythm channel, and to perform high-pass filtering and time difference adjustment on the second audio with the second channel type being cymbal channel.

[0058] In some disclosed embodiments, the audio upmixing module 32 includes other upmixing submodules, which are used to perform low-frequency attenuation, high-frequency attenuation, high-cut operation and low-cut operation respectively based on the first audio in response to the first audio belonging to the first channel type being other, to obtain a second audio with a second channel type of low-frequency attenuation channel, high-frequency attenuation channel, high-cut channel and low-cut channel respectively.

[0059] In some disclosed embodiments, the device selection module 34 includes a first selection submodule for selecting a center speaker as the target speaker for the second audio in response to the second channel type being a master vocal channel; the device selection module 34 includes a second selection submodule for selecting a rear surround speaker as the target speaker for the second audio in response to the second channel type being a secondary vocal channel; the device selection module 34 includes a third selection submodule for selecting a center woofer as the target speaker for the second audio in response to the second channel type being bass; and the device selection module 34 includes a fourth selection submodule for selecting a center speaker as the target speaker in response to the second channel type being a master rhythm channel. The device selection module 34 includes a fifth selection submodule, which selects the left ring speaker and the right ring speaker as the target speakers of the second audio in response to the second channel type being a sub-rhythm channel; the device selection module 34 includes a sixth selection submodule, which selects the sky speaker as the target speaker of the second audio in response to the second channel type being a cymbal channel; the device selection module 34 includes a seventh selection submodule, which selects the left front speaker, the right front speaker, the left ring speaker, and the right ring speaker as the target speakers in response to the second channel type being any of the mid-low frequency attenuation channel, mid-high frequency attenuation channel, high-cut channel, and low-cut channel.

[0060] In some disclosed embodiments, several first channel types include at least one of vocal, drum, bass, and other channel types; and / or, when the first channel type includes vocal, the second channel types under the vocal category include a main vocal channel and a secondary vocal channel; and / or, when the first channel type includes drum, the second channel types under the drum category include a main rhythm channel, a secondary rhythm channel, and a cymbal channel; and / or, when the first channel type includes other categories, the other second channel types include a mid-low frequency attenuation channel, a mid-high frequency attenuation channel, a high-cut channel, and a low-cut channel.

[0061] Please see Figure 4 , Figure 4 This is a schematic diagram of the framework of an embodiment of the electronic device of this application. The electronic device 40 includes at least a memory 41 and a processor 42 coupled to each other. The memory 41 stores at least program instructions, and the processor 42 is used to execute the program instructions to implement the steps in any of the above-described audio processing method embodiments. For details, please refer to the foregoing disclosed embodiments, which will not be repeated here.

[0062] Specifically, processor 42 controls itself and memory 41 to implement the steps in any of the above-described audio processing method embodiments. Processor 42 may also be referred to as a CPU (Central Processing Unit). Processor 42 may be an integrated circuit chip with signal processing capabilities. Processor 42 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 42 may be implemented using integrated circuit chips.

[0063] In the above scheme, the electronic device 40 performs audio separation based on the audio to be played, obtaining several first audio channels of different types. Based on the first channel type to which the first audio belongs, it selects an upmixing strategy that matches the first channel type as the target strategy for that first channel type. Based on the target strategy for the first channel type, it performs upmixing on the first audio to obtain second audio channels of different types. This allows it to select speakers in the three-dimensional sound space whose device attributes match the second channel type to which the second audio belongs, serving as the target speakers for the second audio. The device attributes at least include the deployment location. Then, based on each target speaker for the second audio, the second audio is played. This method, by separating the audio to be played, allows the acquisition of different first channel types. The first audio signal is processed using an upmixing strategy that matches the first channel type of the first audio signal. This upmixing operation yields a second audio signal of a second channel type under the first channel type. Therefore, a single-channel mixed audio signal can be separated and expanded into more refined second audio signals of various second channel types under different first channel types, which helps increase the number of channels for the audio to be played. Furthermore, after obtaining the second audio signals of each second channel type, a speaker whose device properties match the second channel type of the second audio signal is selected in the three-dimensional sound space as the target speaker for playing the second audio signal. This allows for panoramic channel mapping in the three-dimensional sound space as much as possible, improving the distinction between the depth and layering of each sound part. Therefore, it is possible to optimize the audio to be played, overcoming the limitations of the audio data itself to maximize the sense of space and presence.

[0064] Please see Figure 5 , Figure 5 This is a schematic diagram of a framework of an embodiment of the vehicle audio system of this application. The vehicle audio system 50 includes: a control device 51 and speakers 52 deployed at different locations in the vehicle. The control device 51 is electrically connected to each speaker 52, and the control device 51 may specifically include the electronic device in the above-described electronic device embodiments, which can be referred to in the foregoing disclosed embodiments, and will not be repeated here. Exemplarily, the speakers 52 deployed at different locations in the vehicle may include, but are not limited to: a center speaker (not shown), rear surround speakers (not shown), a center woofer (not shown), a left surround speaker (not shown), a right surround speaker (not shown), a sky speaker (not shown), a left front speaker (not shown), a right front speaker (not shown), etc. The possible deployment positions of the speakers 52 are not limited here, and will not be listed one by one.

[0065] The above-described vehicle audio system 50 includes a control device 51 and speakers 52 deployed at different locations in the vehicle. The control device 51 is electrically connected to each speaker 52. Specifically, the control device 51 may include the electronic device in the above-described electronic device embodiment. On the one hand, by performing audio separation on the audio to be played, first audio of different first channel types can be obtained. An upmixing strategy matching the first channel type of the first audio is used to upmix the first audio to obtain second audio of the second channel type under the first channel type. Therefore, a single-channel mixed audio signal can be separated and expanded into second audio of more refined second channel types under different first channel types, which helps to increase the number of channels of the audio to be played. On the other hand, after obtaining the second audio of each second channel type, a speaker 52 whose device attributes match the second channel type of the second audio is further selected in the three-dimensional sound space as the target speaker of the second audio to play the second audio. This can realize panoramic sound channel mapping as much as possible in the three-dimensional sound space, which helps to improve the distinction of the depth of each sound part. Therefore, it is possible to optimize the processing of the audio to be played, thereby overcoming the limitations of the audio data itself and enhancing the sense of space and presence as much as possible. Please see Figure 6 , Figure 6 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. The computer-readable storage medium 60 stores program instructions 61 that can be executed by a processor. The program instructions 61 are used to implement the steps in any of the above-described audio processing method embodiments.

[0066] In the above scheme, the computer-readable storage medium 60 performs audio separation based on the audio to be played, obtaining several first audio channels of different types. Based on the first channel type to which the first audio belongs, it selects an upmixing strategy matching the first channel type as the target strategy for that first channel type. Based on the target strategy for the first channel type, it performs upmixing on the first audio to obtain second audio channels of different types. This allows it to select speakers in the three-dimensional sound space whose device attributes match the second channel type to which the second audio belongs, serving as the target speakers for the second audio. The device attributes at least include the deployment location. Then, based on each target speaker for the second audio, the second audio is played. This method, by separating the audio to be played, allows for the acquisition of different first channel types. The system firstly analyzes the first audio signal of a certain channel type and then uses an upmixing strategy that matches the first channel type to perform upmixing. This upmixing process yields second audio signals of different second channel types under the first channel type. This allows for the separation and expansion of a single-channel mixed audio signal into more refined second audio signals of various second channel types under different first channel types, increasing the number of channels for the audio to be played. Furthermore, after obtaining the second audio signals of each second channel type, the system further selects speakers whose device properties match the second channel type of the second audio signal in a three-dimensional sound space. These speakers serve as the target speakers for the second audio signal, enabling the playback of the second audio signal. This allows for the realization of panoramic sound channel mapping in the three-dimensional sound space, improving the distinction between the depth and layering of each sound element. Therefore, it enables optimized processing of the audio to be played, overcoming the limitations of the audio data itself to maximize the sense of space and presence.

[0067] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0068] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0069] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0070] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0071] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0072] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0073] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

Claims

1. An audio processing method, characterized in that, include: Based on the audio to be played, audio separation is performed to obtain several first audio channels of the first audio type; Based on the first channel type to which the first audio belongs, select an upmixing strategy that matches the first channel type as the target strategy for the first channel type; Based on the target strategy of the first channel type, the first audio is upmixed to obtain the second audio of the second channel type under the first channel type; In a three-dimensional acoustic space, a loudspeaker whose device attributes match the type of the second channel to which the second audio belongs is selected as the target loudspeaker for the second audio; wherein, the device attributes include at least the deployment location; The second audio is played based on each target speaker of the second audio.

2. The method according to claim 1, characterized in that, The audio separation based on the audio to be played yields several first audio channels, including: Data conversion is performed based on the audio to be played to obtain the time spectrum of the audio to be played; Based on the time spectrum, several spectral blocks belonging to different sub-bands are obtained; Based on the prediction of different spectral blocks in each of the sub-frequency bands, the first channel type to which the spectral block belongs is obtained; Waveform recovery is performed based on spectral blocks belonging to the same first channel type to obtain the first audio corresponding to the first channel type.

3. The method according to claim 2, characterized in that, The step of predicting the first channel type to which the spectral block belongs based on different spectral blocks in each of the sub-frequency bands includes: Based on several feature blocks of each sub-frequency band, several sets of feature modeling are performed to obtain the spectral features of different spectral blocks in each sub-frequency band; wherein, each set of feature modeling includes: first performing intra-band modeling in the time domain dimension, and then performing inter-band modeling in the frequency domain dimension; Based on the spectral features of different spectral blocks in each of the sub-bands, classification and prediction are performed to obtain the first channel type to which the spectral block belongs.

4. The method according to claim 2, characterized in that, The segmentation based on the time spectrum yields several spectral blocks belonging to different sub-bands, including: The time spectrum is divided based on each preset frequency band interval to obtain sub-frequency bands for each preset frequency band interval; Based on a cutting interval that matches the preset frequency band interval to which the sub-frequency band belongs, the sub-frequency band is divided into blocks to obtain several spectral blocks of the sub-frequency band; wherein the cutting interval is positively correlated with the representative frequency of the preset frequency band interval.

5. The method according to claim 1, characterized in that, The audio separation is achieved by an audio separation model, and the training steps of the audio separation model include: Acquire sample mixed audio; wherein, the sample mixed audio is labeled with sample track audio of each of the first channel types; Based on the audio separation model, the audio separation is performed on the sample mixed audio to obtain the predicted track-by-track audio of each of the first channel types; Based on the difference between the sample track-segmented audio and the predicted track-segmented audio of the same first channel type, the temporal loss corresponding to the first channel type is obtained, and based on the difference between the sample track-segmented audio and the predicted track-segmented audio of the same first channel type in the real part and the imaginary part of the spectrum, respectively, the frequency domain loss corresponding to the first channel type is obtained. The total loss is obtained by fusing the time-domain loss and frequency-domain loss of the same first channel type. Based on the total loss, the network parameters of the audio separation model are adjusted.

6. The method according to claim 1, characterized in that, The upmixing operation performed on the first audio based on the target strategy of the first channel type to obtain the second audio of the second channel type under the first channel type includes: In response to the fact that the first audio belongs to the first channel type of human voice, the first audio is expanded to obtain second audio with the second channel type of master voice channel and sub-human voice channel respectively. Low frequency attenuation is performed on the second audio with the second channel type of master voice channel, and time difference adjustment is performed on each of the second audio with the second channel type of sub-human voice channel and the pitch of each audio is offset according to different cents. In response to the fact that the first audio belongs to the first channel type of bass, low-frequency attenuation is performed based on the first audio to obtain a second audio of the second channel type of bass; In response to the fact that the first channel type of the first audio is drum sound, the first audio is expanded to obtain second audio with second channel types of main rhythm channel, sub-rhythm channel and cymbal channel respectively. High frequency attenuation is performed on the second audio with the second channel type of the main rhythm channel, low-cut operation and high-cut operation are performed on the second audio with the second channel type of the sub-rhythm channel, and high-pass filtering and time difference adjustment are performed on the second audio with the second channel type of the cymbal channel. In response to the first audio being of a different channel type, mid-low frequency attenuation, mid-high frequency attenuation, high-cut operation, and low-cut operation are performed on the first audio to obtain a second audio with a second channel type of mid-low frequency attenuation channel, mid-high frequency attenuation channel, high-cut channel, and low-cut channel, respectively.

7. The method according to claim 1, characterized in that, The step of selecting a loudspeaker whose device attributes match the second channel type of the second audio in the three-dimensional sound space as the target loudspeaker for the second audio includes: In response to the fact that the second channel type of the second audio is the master voice channel, the center speaker is selected as the target speaker for the second audio. In response to the fact that the second channel type of the second audio is a secondary vocal channel, the rear surround speaker is selected as the target speaker for the second audio. In response to the fact that the second channel type of the second audio is bass, the center woofer is selected as the target speaker for the second audio. In response to the fact that the second channel type to which the second audio belongs is the main rhythm channel, the center speaker is selected as the target speaker for the second audio. In response to the fact that the second channel type of the second audio is a sub-rhythm channel, the left loop speaker and the right loop speaker are selected as the target speakers for the second audio. In response to the fact that the second channel type of the second audio is a cymbal channel, the sky speaker is selected as the target speaker for the second audio; In response to the second audio channel being of any one of the following types: low-frequency attenuation channel, mid-high frequency attenuation channel, high-cut channel, or low-cut channel, the left front speaker, right front speaker, left surround speaker, and right surround speaker are selected as the target speakers.

8. The method according to any one of claims 1 to 7, characterized in that, The first channel types include at least one of the following: vocals, drums, bass, and others; And / or, if the first channel type includes human voice, the second channel type under the human voice includes a main human voice channel and a secondary human voice channel; And / or, if the first channel type includes drum sounds, the second channel type under the drum sounds includes a main rhythm channel, a secondary rhythm channel, and a cymbal channel; And / or, where the first channel type includes others, the other subordinate second channel types include low-to-medium frequency attenuation channels, high-to-medium frequency attenuation channels, high-cut channels, and low-cut channels.

9. An audio processing device, characterized in that, include: The audio separation module is used to separate audio based on the audio to be played, and obtain several first audio channels of the first audio type. The strategy selection module is used to select an upmixing strategy that matches the first channel type based on the first channel type to which the first audio belongs, and use it as the target strategy for the first channel type. An audio upmixing module is used to perform upmixing operations on the first audio based on a target strategy of the first channel type to obtain a second audio of a second channel type under the first channel type. A device selection module is used to select, in a three-dimensional acoustic space, a loudspeaker whose device attributes match the type of the second channel to which the second audio belongs, as the target loudspeaker for the second audio; wherein, the device attributes include at least the deployment location; An audio playback module is used to play the second audio based on each of the target speakers of the second audio.

10. An electronic device, characterized in that, It includes at least a memory and a processor coupled to each other, wherein the memory stores at least program instructions, and the processor is used to execute the program instructions to implement the audio processing method according to any one of claims 1 to 8.

11. A vehicle audio system, characterized in that, It includes: a control device and speakers deployed at different locations in the vehicle, the control device being electrically connected to each of the speakers, and the control device including the electronic device as described in claim 10.

12. A computer-readable storage medium, characterized in that, The device stores program instructions that can be executed by a processor, the program instructions being used to implement the audio processing method according to any one of claims 1 to 8.