Audio up-mixing method and audio equipment

By combining feature extraction and independent audio output channels, the problems of poor spatial sense and high channel correlation in traditional audio upmixing technology are solved, achieving better audio upmixing effects.

CN120980438APending Publication Date: 2025-11-18ANKER INNOVATIONS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410613182.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In traditional audio upmixing techniques, the audio upmixed signal generated by decoding stereo audio signals suffers from poor spatial sense and high channel correlation, which affects the audio upmixing effect.

Method used

By acquiring stereo audio signals, feature extraction is performed, and the left and right channel audio signals are extracted using multiple independent audio output channels. These signals are then fused separately to generate the target channel audio signal, and finally, the target format audio upmix signal is output.

Benefits of technology

It improves the spatial sense and positioning accuracy of the audio upmix signal, reduces the correlation between channels, and enhances the audio upmix effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980438A_ABST
    Figure CN120980438A_ABST
Patent Text Reader

Abstract

The invention relates to an audio upmixing method and audio equipment. The method comprises the following steps: obtaining a stereo audio signal, and carrying out feature extraction on the stereo audio signal to obtain stereo audio features; according to a plurality of audio output channels, a plurality of left channel audio signals and a plurality of right channel audio signals are extracted from the stereo audio features, the plurality of audio output channels are mutually independent, and each audio output channel is used for outputting the corresponding left channel audio signal or the corresponding right channel audio signal; the left channel audio signals and the right channel audio signals corresponding to the same target channel are fused to obtain first audio signals of a plurality of target channels, and each target channel corresponds to two audio output channels; and according to the first audio signals of the plurality of target sound channels, outputting an audio upmix signal in a target format, the target format corresponding to the plurality of target sound channels. By adopting the method, the audio upmixing effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound processing, and in particular to an audio upmix method and device, an audio equipment and a storage medium. BACKGROUND

[0002] With the development of sound processing technology, audio upmix technology appears, and the audio upmix technology can convert a stereo audio signal into a required audio upmix signal, such as a 5.1 channel signal, a 7.1 channel signal or a 7.1.2 channel signal, etc.

[0003] In the traditional technology, the stereo audio signal is usually decoded to generate the required audio upmix signal, for example, the stereo audio signal can be decoded by Dolby directional logic to generate a 5.1 channel signal.

[0004] However, for the audio upmix signal generated by decoding the stereo audio signal, on the one hand, the positioning accuracy of the audio signals of different channels is questionable, which will result in poor spatial sense of the audio upmix signal, thereby making the audio upmix effect worse, on the other hand, there may be a case that there is a high correlation between the audio signals of different channels, which will further affect the audio upmix effect. SUMMARY

[0005] Therefore, it is necessary to provide an audio upmix method, device, audio equipment and computer readable storage medium capable of improving the audio upmix effect in view of the above technical problems.

[0006] In a first aspect, the present application provides an audio upmix method. The method comprises:

[0007] obtaining a stereo audio signal and performing feature extraction on the stereo audio signal to obtain a stereo audio feature;

[0008] extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio feature according to a plurality of audio output channels, wherein the plurality of audio output channels are independent of each other, and each audio output channel is used to output a corresponding left channel audio signal or right channel audio signal;

[0009] fusing left channel audio signals and right channel audio signals corresponding to the same target channel respectively to obtain a plurality of first audio signals of target channels, wherein each target channel corresponds to two audio output channels;

[0010] outputting an audio upmix signal of a target format according to the plurality of first audio signals of target channels, wherein the target format corresponds to the plurality of target channels.

[0011] In one of the embodiments, the extracting, from the stereo audio feature, a plurality of left-channel audio signals and a plurality of right-channel audio signals according to the plurality of audio output channels comprises:

[0012] extracting, from the stereo audio feature, a position audio feature, wherein the position audio feature is used to represent different position sound source signal features of the stereo audio signal; and extracting, from the position audio feature, a plurality of left-channel audio signals and a plurality of right-channel audio signals according to the plurality of audio output channels.

[0013] In one of the embodiments, the position audio feature comprises a first position sound source signal feature and a second position sound source signal feature of the stereo audio signal, and the plurality of audio output channels comprises a plurality of first fully connected networks and a plurality of second fully connected networks which are different from each other; and the extracting, from the position audio feature, a plurality of left-channel audio signals and a plurality of right-channel audio signals according to the plurality of audio output channels comprises:

[0014] inputting the first position sound source signal feature into the plurality of first fully connected networks to output corresponding left-channel audio signals and right-channel audio signals, wherein each of the first fully connected networks is used to output a left-channel audio signal or a right-channel audio signal; and inputting the second position sound source signal feature into the plurality of second fully connected networks to output corresponding left-channel audio signals and right-channel audio signals, wherein each of the second fully connected networks is used to output a left-channel audio signal or a right-channel audio signal.

[0015] In one of the embodiments, the first audio signals of the plurality of target channels comprise a first front-left-channel signal of a front-left channel, a first front-right-channel signal of a front-right channel, a first back-left-channel signal of a back-left channel, a first back-right-channel signal of a back-right channel, and a first center-channel signal of a center channel; and the fusing the left-channel audio signals and the right-channel audio signals corresponding to the same target channel to obtain the first audio signals of the plurality of target channels comprises:

[0016] fusing the left channel audio signal and the right channel audio signal output by the first full connection network corresponding to the front left channel to obtain the first front left channel signal; fusing the left channel audio signal and the right channel audio signal output by the first full connection network corresponding to the front right channel to obtain the first front right channel signal; fusing the left channel audio signal and the right channel audio signal output by the first full connection network corresponding to the center channel to obtain the first center channel signal; fusing the left channel audio signal and the right channel audio signal output by the second full connection network corresponding to the back left channel to obtain the first back left channel signal; and fusing the left channel audio signal and the right channel audio signal output by the second full connection network corresponding to the back right channel to obtain the first back right channel signal.

[0017] In one of the embodiments, the obtaining the stereo audio signal comprises:

[0018] obtaining an original stereo signal, performing a vocal separation on the original stereo signal to obtain a non-vocal signal and a vocal signal, and taking the non-vocal signal as the stereo audio signal.

[0019] In one of the embodiments, the outputting the audio up-mix signal in the target format according to the first audio signals of the multiple target channels comprises:

[0020] merging the vocal signal into front left channel signals, front right channel signals and center channel signals in the first audio signals of the multiple target channels to obtain second audio signals of the multiple target channels, and outputting the audio up-mix signal in the target format according to the second audio signals of the multiple target channels.

[0021] In one of the embodiments, the first audio signals of the multiple target channels include first front left channel signals, first front right channel signals, first back left channel signals, first back right channel signals and first center channel signals, the vocal signal includes left channel vocal signals and right channel vocal signals, the second audio signals of the multiple target channels include second front left channel signals, second front right channel signals, second back left channel signals, second back right channel signals and second center channel signals, and the merging the vocal signal into the front channel signals and the center channel signals in the first audio signals of the multiple target channels to obtain the second audio signals of the multiple target channels comprises:

[0022] The first front left channel signal and the left vocal signal are combined by weighting to obtain the second front left channel signal; the first front right channel signal and the right vocal signal are combined by weighting to obtain the second front right channel signal; the first rear left channel signal is weighted to be the second rear left channel signal, and the first rear right channel signal is weighted to be the second rear right channel signal; the left vocal signal, the right vocal signal and the first center channel signal are combined by weighting to obtain the second center channel signal.

[0023] In one of the embodiments, the audio upmix method is performed by an audio upmix model, and the method further comprises:

[0024] 5.1 channel source signals are obtained, target source signals are selected from the 5.1 channel source signals, and 5-channel target audio signals are extracted from the target source signals; the 5-channel target audio signals are downmixed to obtain stereo training audio signals; 5-channel left audio signals and 5-channel right audio signals are extracted from the stereo training audio signals according to the audio upmix model, and the 5-channel left audio signals and the 5-channel right audio signals are combined to obtain 5-channel output audio signals; and the audio upmix model is optimized according to the difference between the 5-channel target audio signals and the 5-channel output audio signals.

[0025] In one of the embodiments, the audio upmix model is optimized according to the difference between the 5-channel target audio signals and the 5-channel output audio signals, which comprises:

[0026] A first model loss is generated according to the overall difference between the 5-channel target audio signals and the 5-channel output audio signals; a second model loss is generated according to the first volume difference of different channel audio signals in the 5-channel target audio signals and the second volume difference of different channel audio signals in the 5-channel output audio signals; and the audio upmix model is optimized according to the first model loss and the second model loss.

[0027] In a second aspect, the present application further provides an audio device. The audio device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0028] The stereo audio signal is acquired, and feature extraction is performed on the stereo audio signal to obtain stereo audio features. A plurality of left-channel audio signals and a plurality of right-channel audio signals are extracted from the stereo audio features according to a plurality of audio output channels, wherein the plurality of audio output channels are independent of each other, and each audio output channel is configured to output a corresponding left-channel audio signal or right-channel audio signal. Left-channel audio signals and right-channel audio signals corresponding to a same target channel are fused respectively to obtain a plurality of first audio signals of target channels, wherein each target channel corresponds to two audio output channels. An audio upmix signal in a target format is output according to the plurality of first audio signals of target channels, wherein the target format corresponds to the plurality of target channels.

[0029] The audio upmix method described above first acquires a stereo audio signal, and performs feature extraction on the stereo audio signal to obtain stereo audio features. Then, a plurality of left-channel audio signals and a plurality of right-channel audio signals are extracted from the stereo audio features according to a plurality of audio output channels. Each audio output channel is configured to output a corresponding left-channel audio signal or right-channel audio signal. Since the audio output channels are independent of each other, the left-channel audio signals output by the plurality of audio output channels are different from each other and do not interfere with each other. Similarly, the right-channel audio signals output by the plurality of audio output channels are different from each other and do not interfere with each other, which helps to reduce the correlation between the left-channel audio signals output by the plurality of audio output channels and the correlation between the right-channel audio signals output by the plurality of audio output channels. Furthermore, since each target channel corresponds to two audio output channels, the left-channel audio signals and the right-channel audio signals corresponding to a same target channel can be fused respectively to obtain a plurality of first audio signals of target channels with low correlation. According to the plurality of first audio signals of target channels, an audio upmix signal in a target format corresponding to the plurality of target channels can be output. On the one hand, since the left-channel audio signals output by the plurality of audio output channels are different from each other and do not interfere with each other, and the right-channel audio signals output by the plurality of audio output channels are different from each other and do not interfere with each other, the correlation between the first audio signals of the plurality of target channels can be reduced, and the effect of audio upmixing can be improved. On the other hand, since the audio output channels are independent of each other, the first audio signal corresponding to each target channel is generated separately in this application. The first audio signals of the plurality of target channels do not interfere with each other, the positioning accuracy is higher, and the spatial sense of the output audio upmix signal in the target format is better. Therefore, the effect of audio upmixing can be further improved. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 A flowchart of an audio upmix method according to an embodiment of the present application is shown in the figure.

[0031] Figure 2 a visual flowchart of feature extraction of a stereo audio signal in one embodiment of the present application;

[0032] Figure 3 a visual flowchart of feature extraction of a stereo audio signal in one embodiment of the present application;

[0033] Figure 4 a visual flowchart of feature extraction of a stereo audio signal in one embodiment of the present application;

[0034] Figure 5 a visual flowchart of feature extraction of a stereo audio signal in one embodiment of the present application;

[0035] Figure 6 a visual flowchart of feature extraction of a stereo audio signal in one embodiment of the present application;

[0036] Figure 7 a visual flowchart of feature extraction of a stereo audio signal in one embodiment of the present application;

[0037] Figure 8 a visual flowchart of feature extraction of a stereo audio signal in one embodiment of the present application; DETAILED DESCRIPTION

[0038] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0039] It should be noted that in order to improve the spatial sense of playing, the stereo audio signal is usually converted into an audio upmix signal, such as a 5.1 channel audio signal, etc. At present, the stereo audio signal is usually directly decoded into a 5.1 channel audio signal by using a Dolby directional logic decoding method. For the present method, on the one hand, the positioning accuracy of the audio signals of different channels is questionable, which will result in poor spatial sense of the audio upmix signal, thereby making the effect of audio upmix poor. On the other hand, there may be a case that the audio signals of different channels have a high correlation, which will further affect the effect of audio upmix.

[0040] It should be noted that the generated audio upmix signal in the present application is not limited to a 5.1 channel audio signal, but can also be a 7.1 channel audio signal or a 7.1.2 channel audio signal, etc., and the user can select according to actual needs; the audio upmix method of the present application can be applied to an audio device, which can be a headset, a loudspeaker, a home theater type audio device or a hearing aid device, etc., which is not limited here.

[0041] In one embodiment, as shown in Figure 1 , an audio upmix method is provided, comprising the following steps:

[0042] Step 202, obtaining a stereo audio signal, and performing feature extraction on the stereo audio signal to obtain a stereo audio feature.

[0043] Among them, the stereo audio signal is an audio signal with a sense of stereo, which is usually composed of a left channel audio signal and a right channel audio signal; the stereo audio feature is a high-dimensional feature representing the signal feature of the stereo audio signal, which can be a high-dimensional feature matrix, etc.; the audio device can request the external device to obtain the stereo audio signal, so that the external device transmits the stereo audio signal to the audio device, and the audio device can also directly obtain the stereo audio signal from the locally stored audio data.

[0044] As an example, the stereo audio signal can be an original stereo signal collected, which is composed of a vocal signal and a non-vocal signal; the stereo audio signal can also be the original stereo signal after separating the vocal signal, that is, the non-vocal signal in the original stereo signal.

[0045] As an example, step 202 includes: obtaining a stereo audio signal; converting the stereo audio signal from time domain to frequency domain to obtain a stereo frequency domain signal; performing frequency division on the stereo frequency domain signal to obtain a plurality of frequency division signals; then performing feature extraction on the plurality of frequency division signals respectively to obtain a plurality of frequency division signal features; and fusing the plurality of frequency division signal features according to the frequency band bandwidth to obtain a stereo audio feature. In this way, since the feature extraction is performed after the frequency division of the stereo frequency domain signal, the granularity of the feature extraction is finer and the accuracy of the feature extraction is higher, so the accuracy of the final stereo audio feature is also higher.

[0046] As an example, step 202 includes: converting the stereo audio signal from time domain to frequency domain to obtain a stereo frequency domain signal; performing feature extraction on the stereo frequency domain signal to obtain a stereo audio feature.

[0047] As an example, a series of GRU modules can be used as a feature extraction module to realize the feature extraction function in the present embodiment, which is described in detail in Figure 2 ,Figure 2 For feature extraction of a stereo audio signal in an embodiment, a visual flow diagram for obtaining stereo audio features is shown. In this embodiment, after converting the stereo audio signal from the time domain to the frequency domain to obtain a stereo frequency domain signal, the stereo frequency domain signal is first input into a frequency division module to obtain a plurality of frequency division signals under a plurality of frequency division bands, and the plurality of frequency division signals under the plurality of frequency division bands are respectively input into corresponding complex time-frequency encoders for feature encoding to obtain frequency division signal encoding features under the plurality of frequency division bands. Then, the frequency division signal encoding features under the plurality of frequency division bands are respectively input into a series of GRU modules (feature extraction modules) to obtain extracted frequency division features under the plurality of frequency division bands. Then, the frequency division features under the plurality of frequency division bands are input into a Merge module (feature fusion module) for feature fusion to obtain fused features. Finally, the fused features are input into a GRU module to map the fused features to a preset feature dimension, and the final stereo audio features are obtained.

[0048] In step 204, a plurality of left channel audio signals and a plurality of right channel audio signals are extracted from the stereo audio features according to a plurality of audio output channels, wherein the plurality of audio output channels are independent of each other, and each audio output channel is used to output a corresponding left channel audio signal or right channel audio signal.

[0049] In this embodiment, a plurality of audio output channels are provided, which are independent of each other and can output left channel audio signals and right channel audio signals with different weights. One audio output channel can output a corresponding left channel audio signal or right channel audio signal. Since the audio output channels are independent of each other, the left channel audio signals and right channel audio signals output by each audio output channel are different and have low correlation.

[0050] As an example, step 204 includes: outputting a plurality of left channel audio signals and a plurality of right channel audio signals by inputting the stereo audio features into a plurality of audio output channels, respectively, wherein each audio output channel is used to map the stereo audio features to a corresponding left channel audio signal or right channel audio signal, and the mapping parameters of each audio output channel are different, so that the plurality of left channel audio signals and the plurality of right channel audio signals have low correlation.

[0051] In step 206, left channel audio signals and right channel audio signals corresponding to the same target channel are fused respectively to obtain a plurality of first audio signals of a plurality of target channels, wherein each target channel corresponds to two audio output channels.

[0052] The number and type of the plurality of target channels are determined by a target format of the generated audio upmix signal. For example, the plurality of target channels are front left, front right, back left, back right and center channels, and the first audio signals of the plurality of target channels are first front left, first front right, first back left, first back right and first center channel signals.

[0053] It should be noted that the plurality of target channels and the plurality of audio output channels have a corresponding relationship in this embodiment. Each two audio output channels correspond to one target channel, and the two audio output channels corresponding to each target channel output left and right channel audio signals respectively. Therefore, each target channel corresponds to one left channel audio signal and one right channel audio signal.

[0054] As an example, step 206 includes: for each target channel, obtaining the left and right channel audio signals output by the two audio output channels corresponding to the target channel, weighting the obtained left and right channel audio signals respectively, and fusing the weighted left and right channel audio signals to obtain the first audio signal of the target channel. In this embodiment, the corresponding relationship between the target channels and the audio output channels is set, and the left and right channel audio signals output by different audio output channels are combined into the first audio signal of the target channel. Since the audio output channels are independent of each other and do not interfere with each other, the correlation between the first audio signals of the target channels obtained by combination is low, and the positioning is more accurate. Therefore, it is helpful to improve the effect of audio upmixing.

[0055] Step 208 outputs an audio upmix signal of a target format according to the first audio signals of the plurality of target channels, wherein the target format corresponds to the plurality of target channels.

[0056] As an example, step 208 includes: combining the first audio signals of the plurality of target channels to obtain the audio upmix signal of the target format.

[0057] As an example, the audio upmix signal of the target format is a 5.1 channel audio signal, and step 208 includes: low-pass filtering the original stereo signal corresponding to the stereo audio signal to obtain a bass channel signal; and combining the bass channel signal, the first front left channel signal, the first front right channel signal, the first back left channel signal, the first back right channel signal and the first center channel signal to obtain the 5.1 channel audio signal.

[0058] In the above audio upmix method, first, a stereo audio signal is obtained, and feature extraction is performed on the stereo audio signal to obtain a stereo audio feature. Then, a plurality of left channel audio signals and a plurality of right channel audio signals are extracted from the stereo audio feature according to a plurality of audio output channels. Each audio output channel is used to output a corresponding left channel audio signal or right channel audio signal. Since the audio output channels are independent of each other, the left channel audio signals output by the plurality of audio output channels are different and do not interfere with each other. Similarly, the right channel audio signals output by the plurality of audio output channels are different and do not interfere with each other. This helps to reduce the correlation between the left channel audio signals output by the plurality of audio output channels and the correlation between the right channel audio signals output by the plurality of audio output channels. Furthermore, since each target channel corresponds to two audio output channels, the left channel audio signals and the right channel audio signals corresponding to the same target channel can be fused respectively to obtain a plurality of first audio signals of target channels with low correlation. According to the first audio signals of the plurality of target channels, an audio upmix signal of a target format corresponding to the plurality of target channels can be output. On the one hand, since the left channel audio signals output by the plurality of audio output channels are different and do not interfere with each other, and the right channel audio signals output by the plurality of audio output channels are different and do not interfere with each other, the correlation between the first audio signals of the plurality of target channels can be reduced, and the effect of audio upmix can be improved. On the other hand, since the audio output channels are independent of each other, the first audio signal corresponding to each target channel is generated separately in this application. The first audio signals of the plurality of target channels do not interfere with each other, have higher positioning accuracy, and the spatial sense of the output audio upmix signal of the target format is better. Therefore, the effect of audio upmix can be further improved.

[0059] In one embodiment, as shown in Figure 3 extracting the left channel audio signals output by the plurality of audio output channels and the right channel audio signals output by the plurality of audio output channels from the stereo audio feature also includes:

[0060] Step 302, extracting a position audio feature from the stereo audio feature, wherein the position audio feature is used to represent the sound source signal features of different positions of the stereo audio signal;

[0061] The position audio feature corresponding to the stereo audio feature can be one or multiple, which is not limited here. For example, the position audio feature can be a first position sound source signal feature representing a front position audio feature, or a second position audio signal feature representing a rear position audio feature.

[0062] As an example, step 302 comprises: extracting, according to at least one preset orientation feature extraction module, a corresponding orientation audio feature from the stereo audio feature, wherein the preset orientation feature extraction module is configured to extract an orientation audio feature representing a sound source signal feature at a preset orientation from the stereo audio feature.

[0063] As an example, the preset orientation feature extraction module can be a GRU (Gated Recurrent Unit) module.

[0064] Step 304: extracting, according to a plurality of audio output channels, a plurality of left channel audio signals and a plurality of right channel audio signals from the orientation audio features.

[0065] As an example, step 304 comprises: for each orientation audio feature, mapping the orientation audio feature to a corresponding left channel audio signal and a right channel audio signal through an audio output channel corresponding to the orientation audio feature, with the orientation audio feature as an input of the corresponding audio output channel.

[0066] As an example, if the finally generated audio upmix signal is a 5.1 channel audio signal, the different orientations can be a front left orientation (corresponding to a front left channel in the 5.1 channel audio signal), a front right orientation (corresponding to a front right channel in the 5.1 channel audio signal), a rear left orientation (corresponding to a rear left channel in the 5.1 channel audio signal), a rear right orientation (corresponding to a rear right channel in the 5.1 channel audio signal), and a center orientation (corresponding to a center channel in the 5.1 channel audio signal).

[0067] In this way, in the embodiment, the left channel audio signals and the right channel audio signals are generated by dividing the stereo audio feature into at least one orientation audio feature, so that the sound source signal features at different orientations in the stereo audio feature do not interfere with each other, which helps to improve the accuracy of the finally generated left channel audio signals and right channel audio signals.

[0068] In one embodiment, the orientation audio features comprise a first orientation sound source signal feature and a second orientation sound source signal feature of the stereo audio signal, and each preset orientation feature extraction module comprises a first orientation feature extraction module and a second orientation feature extraction module; the step of extracting, according to a plurality of audio output channels, a plurality of left channel audio signals and a plurality of right channel audio signals from the orientation audio features comprises:

[0069] extracting, according to the first orientation feature extraction module, the first orientation sound source signal feature from the stereo audio feature, and extracting, according to the second orientation feature extraction module, the second orientation sound source signal feature from the stereo audio feature.

[0070] As an example, the first directional sound signal feature can be a directional signal feature representing a front sound source signal feature of the stereo audio signal; the second directional sound signal feature can be a directional signal feature representing a back sound source signal feature of the stereo audio signal.

[0071] As an example, the first directional sound signal feature can be a directional signal feature representing a front left sound source signal feature of the stereo audio signal; the second directional sound signal feature can be a directional signal feature representing a back right sound source signal feature of the stereo audio signal.

[0072] The first directional sound source signal feature and the second directional sound source signal feature described above are specifically used to represent sound source signal features of which direction of the stereo audio signal, which is not limited here and can be selected according to actual needs.

[0073] In an embodiment, the directional audio feature includes a first directional sound source signal feature and a second directional sound source signal feature of the stereo audio signal, and the plurality of audio output channels includes a plurality of first fully connected networks and a plurality of second fully connected networks which are different from each other; the plurality of left channel audio signals and the plurality of right channel audio signals are extracted from the directional audio feature according to the plurality of audio output channels, including:

[0074] The first directional sound source signal feature is inputted into the plurality of first fully connected networks, and corresponding left channel audio signals and right channel audio signals are outputted through the plurality of first fully connected networks, wherein each first fully connected network is used to output a left channel audio signal or a right channel audio signal; the second directional sound source signal feature is inputted into the plurality of second fully connected networks, and corresponding left channel audio signals and right channel audio signals are outputted through the plurality of second fully connected networks, wherein each second fully connected network is used to output a left channel audio signal or a right channel audio signal.

[0075] Each audio output channel can include one first fully connected network or one second fully connected network, and all first fully connected networks and all second fully connected networks are different from each other.

[0076] Specifically, for each first fully connected network, the first directional sound source signal feature is inputted into the first fully connected network, the first directional sound source signal feature is fully connected through the first fully connected network, and a left channel audio signal or a right channel audio signal outputted by the first fully connected network is obtained through a preset activation function after full connection; for each second fully connected network, the second directional sound source signal feature is inputted into the second fully connected network, the second directional sound source signal feature is fully connected through the second fully connected network, and a left channel audio signal or a right channel audio signal outputted by the second fully connected network is obtained through a preset activation function after full connection.

[0077] In the embodiment, first, the first azimuth sound source signal feature and the second azimuth sound source signal feature of the stereo audio signal are separated from the stereo audio feature; then, the left channel audio signal and the right channel audio signal output by the plurality of audio output channels are extracted for the first azimuth sound source signal feature, and the left channel audio signal and the right channel audio signal output by the plurality of audio output channels are extracted for the second azimuth sound source signal feature. In this way, after the different azimuth sound source signal features in the stereo audio feature are separated, the left channel audio signal and the right channel audio signal are generated. During the generation of the left channel audio signal and the right channel signal, the different azimuth sound source signal features in the stereo audio feature do not interfere with each other, so that the accuracy of the generated left channel audio signal and right channel audio signal can be improved.

[0078] In addition, it should be noted that the more the number of azimuth audio features separated from the stereo audio feature, the higher the accuracy of the finally generated left channel audio signal and right channel audio signal. However, the more the number of azimuth audio features separated from the stereo audio feature, the lower the efficiency of the generated left channel audio signal and right channel audio signal. In the embodiment, the first azimuth sound source signal feature and the second azimuth sound source signal feature outputting two azimuths are separated from the stereo audio feature, and the first azimuth sound source signal feature and the second azimuth sound source signal feature are used to generate the left channel audio signal and the right channel audio signal output by the plurality of audio output channels. The generation efficiency and generation accuracy of the left channel audio signal and the right channel audio signal can be considered.

[0079] In one embodiment, taking the generation of 5.1 channel audio signals as an example, the first audio signal of the plurality of target channels includes the first front left channel signal of the front left channel, the first front right channel signal of the front right channel, the first rear left channel signal of the rear left channel, the first rear right channel signal of the rear right channel, and the first center channel signal of the center channel; the left channel audio signal and the right channel audio signal corresponding to the same target channel are fused to obtain the first audio signal of the plurality of target channels, including:

[0080] The left channel audio signal and the right channel audio signal output by the first full connection network corresponding to the front left channel are fused to obtain a first front left channel signal; the left channel audio signal and the right channel audio signal output by the first full connection network corresponding to the front right channel are fused to obtain a first front right channel signal; the left channel audio signal and the right channel audio signal output by the first full connection network corresponding to the center channel are fused to obtain a first center channel signal; the left channel audio signal and the right channel audio signal output by the second full connection network corresponding to the rear left channel are fused to obtain a first rear left channel signal; and the left channel audio signal and the right channel audio signal output by the second full connection network corresponding to the rear right channel are fused to obtain a first rear right channel signal.

[0081] Specifically, according to the correspondence between the target channel and the audio output channel, the left channel audio signal and the right channel audio signal corresponding to the front left channel, the left channel audio signal and the right channel audio signal corresponding to the front right channel, and the left channel audio signal and the right channel audio signal corresponding to the center channel are determined in the left channel audio signals and the right channel audio signals output by the different first full connection networks; the left channel audio signal and the right channel audio signal corresponding to the front left channel are weighted and fused to obtain a first front left channel signal; the left channel audio signal and the right channel audio signal corresponding to the front right channel are weighted and fused to obtain a first front right channel signal; and the left channel audio signal and the right channel audio signal corresponding to the center channel are weighted and fused to obtain a first center channel signal; according to the correspondence between the target channel and the audio output channel, the left channel audio signal and the right channel audio signal corresponding to the rear left channel and the left channel audio signal and the right channel audio signal corresponding to the rear right channel are determined in the left channel audio signals and the right channel audio signals output by the different second full connection networks; the left channel audio signal and the right channel audio signal corresponding to the rear left channel are weighted and fused to obtain a first rear left channel signal; and the left channel audio signal and the right channel audio signal corresponding to the rear right channel are weighted and fused to obtain a first rear right channel signal.

[0082] In this embodiment, the left channel audio signal and the right channel audio signal that are different from each other and do not interfere with each other are generated by setting different first full connection networks and different second full connection networks. In this way, the first front left channel signal, the first front right channel signal and the first center channel signal are generated by respectively weighting and fusing the left channel audio signal and the right channel audio signal output by the different first full connection networks, and the first rear left channel signal and the first rear right channel signal are generated by respectively weighting and fusing the left channel audio signal and the right channel audio signal output by the different second full connection networks. By combining the left channel audio signal and the right channel audio signal that are different from each other and do not interfere with each other, it can be further ensured that the first front left channel signal, the first front right channel signal and the first center channel signal, the left channel audio signal and the right channel audio signal obtained after combination are also different from each other and do not interfere with each other, so that the positioning accuracy of the spatial direction of the 5.1 channel audio signal can be improved, and the correlation between the audio signals of different channels is reduced, which helps to improve the effect of audio upmixing.

[0083] In one embodiment, referring to Figure 4 , Figure 4 FIG. 1 is a visual flow diagram for generating audio signals of each target channel in a 5.1 channel audio signal according to characteristics of first directional sound source signals and characteristics of second directional sound source signals in one embodiment. In FIG. 1, GRU_F is a first directional feature extraction module for extracting first directional sound source signal characteristics of a stereo audio signal from stereo audio characteristics, GRU_S is a second directional feature extraction module for extracting second directional sound source signal characteristics of the stereo audio signal from the stereo audio characteristics, the FC module connected to GRU_F is each first full connection network, and each first full connection network is connected to an activation function Sigmoid, the FC module connected to GRU_S is each second full connection network, and each second full connection network is connected to an activation function Sigmoid, L represents an output left channel audio signal, R represents an output right channel audio signal, FL represents a first front left channel signal, FR represents a first front right channel signal, C represents a first center channel signal, SL represents a first rear left channel signal, and SR represents a first rear right channel signal. represents weighting and fusing. As an example, the way of weighting and fusing can be that the output left channel signal and the output right channel signal are respectively multiplied by corresponding masks and then added.

[0084] In one embodiment, the stereo audio signal is obtained, including:

[0085] An original stereo signal is obtained, and a vocal separation is performed on the original stereo signal to obtain a non-vocal signal and a vocal signal. The non-vocal signal is taken as the stereo audio signal.

[0086] In order to keep the integrity of the human voice, the human voice signal in the original stereo signal can be separated before the audio downmixing in the embodiment, so that the stereo audio signal without the human voice signal can be obtained.

[0087] As an example, the original stereo signal can also be directly used as the stereo audio signal in the embodiment, so that the human voice signal exists in the stereo audio signal.

[0088] In one embodiment, the audio upmix signal in the target format is output according to the first audio signal of the plurality of target channels, including:

[0089] The human voice signal is merged into the front left channel signal, the front right channel signal and the center channel signal in the first audio signal of the plurality of target channels to obtain the second audio signal of the plurality of target channels; and the audio upmix signal in the target format is output according to the second audio signal of the plurality of target channels.

[0090] For the audio upmix signal, the human voice signal usually exists in the front left channel, the front right channel and the center channel, so that the user can have a better auditory experience. For example, the human voice signal usually exists in the front left channel, the front right channel and the center channel in the 5.1 channel audio signal, and the human voice signal usually does not exist in the rear left channel and the rear right channel, so that the auditory experience of the 5.1 channel audio signal is better. Therefore, after obtaining the first audio signal of the plurality of target channels and before finally generating the audio upmix signal in the target format, the human voice signal is merged into the first audio signal of the plurality of target channels, so that the audio upmix signal in the target format finally generated has complete human voice signal. Since the human voice signal does not follow the stereo audio signal to perform a series of audio upmix operations, the distortion of the human voice signal in the audio upmix process can be avoided, the integrity and authenticity of the human voice signal are ensured, and therefore the effect of the audio upmix can be further improved.

[0091] Specifically, the human voice signal is weighted and merged into the front left channel signal, the front right channel signal and the center channel signal in the first audio signal of the plurality of target channels to obtain the second audio signal of the plurality of target channels; and the second audio signal of the plurality of target channels is merged into the audio upmix signal in the target format.

[0092] In the embodiment, after separating the vocal signal from the original stereo signal and generating the first audio signals of the multiple target channels according to the stereo audio signal obtained from the separated vocal signal, the selective audio rendering of the first audio signals of the multiple target channels is realized by merging the vocal signal into the front left channel signal, the front right channel signal and the center channel signal in the first audio signals of the multiple target channels only, so that the generated second audio signals of the multiple target channels are more in line with the hearing habits of the user and the hearing experience is better, and thus the effect of the audio upmixing can be improved.

[0093] In one embodiment, the audio upmixing signal in the target format is taken as a 5.1 channel audio signal for example, the first audio signals of the multiple target channels include a first front left channel signal, a first front right channel signal, a first center channel signal, a first rear left channel signal and a first rear right channel signal, and the vocal signal is also a stereo signal, so the vocal signal includes a left channel vocal signal and a right channel vocal signal, and the second audio signals of the multiple target channels include a second front left channel signal, a second front right channel signal, a second center channel signal, a second rear left channel signal and a second rear right channel signal; the vocal signal is merged into the front channel signal and the center channel signal in the first audio signals of the multiple target channels to obtain the second audio signals of the multiple target channels, which include:

[0094] The first front left channel signal and the left channel vocal signal are weighted and merged to obtain the second front left channel signal; the first front right channel signal and the right channel vocal signal are weighted and merged to obtain the second front right channel signal; the first rear left channel signal is weighted to obtain the second rear left channel signal, and the first rear right channel signal is weighted to obtain the second rear right channel signal; the left channel vocal signal, the right channel vocal signal and the first center channel signal are weighted and merged to obtain the second center channel signal.

[0095] Specifically, the first front left channel signal is weighted according to a first preset weight, and the left channel vocal signal is weighted according to a second preset weight, and the weighted first front left channel signal and the weighted left channel vocal signal are added to obtain a second front left channel signal; the first front right channel signal is weighted according to the first preset weight, and the right channel vocal signal is weighted according to the second preset weight, and the weighted first front right channel signal and the weighted right channel vocal signal are added to obtain a second front right channel signal; the first rear left channel signal and the first rear right channel signal are weighted according to a third preset weight to obtain a second rear left channel signal corresponding to the first rear left channel signal and a second rear right channel signal corresponding to the first rear right channel signal; the first center channel signal is weighted according to a fourth preset weight, and the left channel vocal signal and the right channel vocal signal are weighted according to a fifth preset weight, and the weighted first center channel signal, the weighted left channel vocal signal and the weighted right channel vocal signal are added to obtain a second center channel signal.

[0096] Further, the original stereo signal can be low-pass filtered, and the low-pass filtered original stereo signal is weighted according to a sixth preset weight to obtain a bass channel signal, so that the second front left channel signal, the second front right channel signal, the second center channel signal, the second rear left channel signal, the second rear right channel signal and the bass channel signal together serve as a 5.1 channel audio signal.

[0097] As an example, the first preset weight is used to represent the importance of non-vocal sound in the front surround channel (including the front left channel and the front right channel) to the 5.1 channel audio signal, the more important the non-vocal sound in the front surround channel, the higher the first preset weight; the second preset weight is used to represent the importance of vocal sound in the front surround channel to the 5.1 channel audio signal, the more important the vocal sound in the front surround channel, the higher the second preset weight; the third preset weight is used to represent the importance of non-vocal sound in the rear surround channel (including the rear left channel and the rear right channel) to the 5.1 channel audio signal, the more important the non-vocal sound in the rear surround channel, the higher the third preset weight; the fourth preset weight is used to represent the importance of non-vocal sound in the center channel to the 5.1 channel audio signal, the more important the non-vocal sound in the center channel, the higher the fourth preset weight; the fifth preset weight is used to represent the importance of vocal sound in the center channel to the 5.1 channel audio signal, the more important the vocal sound in the center channel, the higher the fifth preset weight; the sixth preset weight is used to represent the importance of the bass channel signal to the 5.1 channel audio signal, the more important the bass channel signal, the higher the sixth preset weight.

[0098] In one embodiment, the finally generated target format audio upmix signal is taken as an example of a 5.1 channel audio signal, and the following description is made with reference toFigure 5 , Figure 5 is a visual flowchart for generating 5.1 channel audio signals from original stereo signals in an embodiment, wherein stereo L, R are original stereo signals, vocal L, R are left and right channel vocal signals, non-vocal is stereo audio signal, AI upmix module is used to generate first audio signals of multiple target channels according to non-vocal, the first audio signals of multiple target channels include first front left channel signal O FL, first front right channel signal O FR, first rear left channel signal O SL, first rear right channel signal O RL, first center channel signal O C, LPF is a low-pass filter, F Gain is a first preset weight, V Gain is a second preset weight, S Gain is a third preset weight, C Gain1 is a fourth preset weight, C Gain2 is a fifth preset weight, Bass Gain is a sixth preset weight, represents addition, FL is a second front left channel signal, FR is a second front right channel signal, SL is a second rear left channel signal, RL is a second rear right channel signal, C is a second center channel signal, and Bass is a bass channel signal, so that FL, FR, SL, RL, C and Bass together constitute 5.1 channel audio signals.

[0099] In an embodiment, the finally generated audio upmix signal in the target format is taken as an example of 5.1 channel audio signals, and reference is made to Figure 6 , Figure 6For another embodiment, a visual flowchart for generating a 5.1 channel audio signal from an original stereo signal is shown, wherein the original stereo signal can not be separated into a vocal signal and a music signal, and the original stereo signal is directly used as a stereo audio signal, the stereo L and R are the stereo audio signal, an AI upmix module is used to generate a plurality of target channel first audio signals from the stereo audio signal, the plurality of target channel first audio signals include a first front left channel signal O_FL, a first front right channel signal O_FR, a first rear left channel signal O_SL, a first rear right channel signal O_RL, and a first center channel signal O_C, a LPF is a low pass filter, the first front left channel signal O_FL is directly weighted into a second front left channel signal FL according to a first preset weight F_Gain, and the first front right channel signal O_FR is directly weighted into FR according to the first preset weight F_Gain; the first rear left channel signal O_SL is directly weighted into a second rear left channel signal SL according to a third preset weight S_Gain, and the first rear right channel signal O_RL is directly weighted into a second rear right channel signal RL according to the third preset weight S_Gain; the first center channel signal O_C is directly weighted into a second center channel signal C according to a fourth preset weight C_Gain1; the low pass filtered stereo L and R are weighted into a bass channel signal Bass according to a sixth preset weight Bass_Gain, so that FL, FR, SL, RL, C and Bass jointly form a 5.1 channel audio signal.

[0100] In one embodiment, the audio upmix method is executed by an audio upmix model, and the audio upmix method further includes:

[0101] In one embodiment, the audio upmix method is executed by an audio upmix model, and the audio upmix method further includes:

[0102] In one embodiment, the audio upmix method is executed by an audio upmix model, and the audio upmix method further includes:

[0103] Specifically, each 5.1 channel audio source signal is acquired, a target audio source signal is screened from each 5.1 channel audio source signal, and a 5-channel target audio signal is extracted from the target audio source signal; a preset downmix matrix is used to downmix the 5-channel target audio signal, and a stereo signal obtained by downmixing is used as a stereo training audio signal; based on an audio upmix model, the following audio upmix process is performed: the stereo training audio signal is converted from a time domain to a frequency domain to obtain a stereo frequency domain training signal, the stereo frequency domain training signal is frequency-division processed to obtain a plurality of frequency-division training signals, feature extraction is performed on the plurality of frequency-division training signals to obtain a plurality of frequency-division training signal features, the plurality of frequency-division training signal features are fused according to frequency band bandwidth to obtain a stereo training audio feature; the stereo training audio feature is input into 10 audio output channels to output a 5-channel left channel audio signal and a 5-channel right channel audio signal, and the 5-channel left channel audio signal and the 5-channel right channel audio signal are weighted and fused into a 5-channel output audio signal; the audio upmix model is iteratively updated and optimized according to a model loss calculated based on a difference between the 5-channel target audio signal and the 5-channel output audio signal.

[0104] It should be noted that the audio upmix process performed by the audio upmix model can refer to the audio upmix process in the audio upmix method embodiment, which will not be described in detail here.

[0105] As an example, the 5-channel target audio signal is extracted from the target audio source signal, including:

[0106] The bass channel signal is removed from the target audio source signal to obtain the 5-channel target audio signal.

[0107] As an example, the 5-channel target audio signal is extracted from the target audio source signal, including:

[0108] The bass channel signal and the vocal signal are removed from the target audio source signal to obtain the 5-channel target audio signal.

[0109] It should be noted that the 5-channel target audio signal can be a front left channel signal, a front right channel signal, a rear left channel signal, a rear right channel signal, and a center channel signal in a 5.1 audio signal.

[0110] In one embodiment, the audio upmix model is optimized according to a difference between the 5-channel target audio signal and the 5-channel output audio signal, including:

[0111] A first model loss is generated according to a signal overall difference between the 5-channel target audio signal and the 5-channel output audio signal, a second model loss is generated according to a volume difference of different channel audio signals in the 5-channel output audio signal, and the audio upmix model is optimized according to the first model loss and the second model loss.

[0112] Specifically, the first model loss is calculated according to the signal overall difference between the 5-channel target audio signal and the 5-channel output audio signal; the second model loss is calculated according to the first volume difference of different channel audio signals in the 5-channel target audio signal and the second volume difference of different channel audio signals in the 5-channel output audio signal; the first model loss and the second model loss are added to obtain the total model loss; and the gradient information calculated according to the total model loss is used to iteratively update the optimized audio up-mix model.

[0113] As an example, the calculation formula of the first model loss according to the signal overall difference between the 5-channel target audio signal and the 5-channel output audio signal is as follows:

[0114]

[0115] wherein, the first model loss is L1, the 5-channel target audio signal is S, the 5-channel output audio signal is Y.

[0116] As an example, the second model loss is calculated according to the first volume difference of different channel audio signals in the 5-channel target audio signal and the second volume difference of different channel audio signals in the 5-channel output audio signal, including:

[0117] The first volume difference values between different channel audio signals in the 5-channel target audio signal and the second volume difference values between different channel audio signals in the 5-channel output audio signal are calculated; and the accumulated result of the difference values of each group of first volume difference values and second volume difference values corresponding to the same channel audio signal is calculated, and the accumulated result is taken as the second model loss.

[0118] As an example, the way of screening the target audio source signal from each 5.1 channel audio source signal includes at least one of the following ways:

[0119] The first way: the target audio source signal is screened from each 5.1 channel audio source signal according to the correlation between the front left channel signal and the rear left channel signal and the correlation between the front right channel signal and the rear right channel signal in each 5.1 channel audio source signal.

[0120] For the 5.1 channel audio source signal, if the correlation between the front left channel signal and the rear left channel signal is too high or the correlation between the front right channel signal and the rear right channel signal is too high, the hearing effect of the 5.1 channel audio source signal will be poor.

[0121] Specifically, for each 5.1 channel sound source signal, a first correlation value between the front left channel signal and the rear left channel signal in the 5.1 channel sound source signal is calculated, and a second correlation value between the front right channel signal and the rear right channel signal in the 5.1 channel sound source signal is calculated; if both the first correlation value and the second correlation value are less than a preset correlation threshold, the 5.1 channel sound source signal is taken as a target sound source signal; if both the first correlation value and the second correlation value are not less than the preset correlation threshold, the 5.1 channel sound source signal is not taken as the target sound source signal.

[0122] The second mode is to screen a target sound source signal from each 5.1 channel sound source signal according to a signal energy of the rear surround channel signal in each 5.1 channel sound source signal.

[0123] For the 5.1 channel sound source signal, if the signal energy of the rear surround channel signal is relatively small, it is easy to cause the 5.1 channel sound source signal to have no spatial listening of the rear surround, and cause the hearing effect to be poor.

[0124] Specifically, for each 5.1 channel sound source signal, a signal energy of the rear surround channel signal in the 5.1 channel sound source signal is determined, if the signal energy is greater than a preset signal energy threshold, the 5.1 channel sound source signal is taken as a target sound source signal; if the signal energy is not greater than the preset signal energy threshold, the 5.1 channel sound source signal is not taken as the target sound source signal.

[0125] In the embodiment, the audio up-mix model is trained specifically, the audio up-mix model used to perform the audio up-mix process can be trained, on the one hand, the model loss can be reasonably set according to the signal overall difference between the 5-channel target audio signal and the 5-channel output audio signal, the first volume difference of different channel audio signals in the 5-channel target audio signal, and the second volume difference of different channel audio signals in the 5-channel output audio signal, so that the audio up-mix model can learn a more accurate 5-channel audio signal, and finally a more accurate 5.1 channel audio signal can be generated, on the other hand, the 5.1 channel sound source signal can also be selected specifically, so that the correlation between the front left channel signal and the rear left channel signal in the selected 5.1 channel sound source signal is low, and the correlation between the front right channel signal and the rear right channel signal is low, or the signal energy of the rear surround channel signal in the selected 5.1 channel sound source signal is large, so that the training sample quality of the audio up-mix model can be improved, the audio up-mix model trained is more accurate, and a foundation for improving the audio up-mix effect is laid.

[0126] It should be understood that although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time but can be executed at different times, and the execution of the steps or stages is not necessarily sequential but can be executed alternately or alternately with at least part of other steps or stages.

[0127] Based on the same inventive concept, the embodiments of the present application also provide an audio upmixing device for implementing the audio upmixing method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more audio upmixing device embodiments provided below can refer to the limitations of the audio upmixing method described above, and will not be repeated here.

[0128] In one embodiment, as shown in Figure 7 An audio upmixing device is provided, comprising: a stereo feature extraction module 702, a signal extraction module 704, a fusion module 706, and a signal generation module 708, wherein:

[0129] The stereo feature extraction module 702 is configured to obtain a stereo audio signal and perform feature extraction on the stereo audio signal to obtain a stereo audio feature.

[0130] The signal extraction module 704 is configured to extract a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio feature according to a plurality of audio output channels, wherein the plurality of audio output channels are independent of each other, and each audio output channel is configured to output a corresponding left channel audio signal or right channel audio signal.

[0131] The fusion module 706 is configured to fuse left channel audio signals and right channel audio signals corresponding to the same target channel respectively to obtain a plurality of first audio signals of target channels, wherein each target channel corresponds to two audio output channels.

[0132] The signal generation module 708 is configured to output an audio upmixing signal of a target format according to the plurality of first audio signals of target channels, wherein the target format corresponds to the plurality of target channels.

[0133] In one embodiment, the signal extraction module is further configured to:

[0134] extracting a position audio feature from the stereo audio feature, wherein the position audio feature is used to represent different position sound source signal features of the stereo audio signal; and extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the position audio feature according to the plurality of audio output channels.

[0135] In one of the embodiments, the position audio feature includes a first position sound source signal feature and a second position sound source signal feature of the stereo audio signal, and the plurality of audio output channels includes a plurality of first full connection networks and a plurality of second full connection networks which are different from each other; and the signal extraction module is further configured to:

[0136] inputting the first position sound source signal feature to output corresponding left channel audio signals and right channel audio signals through the plurality of first full connection networks, wherein each of the first full connection networks is configured to output left channel audio signals or right channel audio signals; and inputting the second position sound source signal feature to output corresponding left channel audio signals and right channel audio signals through the plurality of second full connection networks, wherein each of the second full connection networks is configured to output left channel audio signals or right channel audio signals.

[0137] In one of the embodiments, the first audio signals of the plurality of target channels include a first front left channel signal of a front left channel, a first front right channel signal of a front right channel, a first back left channel signal of a back left channel, a first back right channel signal of a back right channel, and a first center channel signal of a center channel; and the signal extraction module is further configured to:

[0138] fusing left channel audio signals and right channel audio signals output by the first full connection network corresponding to the front left channel to obtain the first front left channel signal; fusing left channel audio signals and right channel audio signals output by the first full connection network corresponding to the front right channel to obtain the first front right channel signal; fusing left channel audio signals and right channel audio signals output by the first full connection network corresponding to the center channel to obtain the first center channel signal; fusing left channel audio signals and right channel audio signals output by the second full connection network corresponding to the back left channel to obtain the first back left channel signal; and fusing left channel audio signals and right channel audio signals output by the second full connection network corresponding to the back right channel to obtain the first back right channel signal.

[0139] In one of the embodiments, the stereo feature extraction module is further configured to:

[0140] obtaining an original stereo signal, performing voice separation on the original stereo signal to obtain a non-voice signal and a voice signal, and taking the non-voice signal as the stereo audio signal.

[0141] In one of the embodiments, the signal generation module is further configured to:

[0142] merge the vocal signal into a front left channel signal, a front right channel signal and a center channel signal of the first audio signals of the plurality of target channels to obtain second audio signals of the plurality of target channels; and output an audio upmix signal in a target format according to the second audio signals of the plurality of target channels.

[0143] In one of the embodiments, the first audio signals of the plurality of target channels include a first front left channel signal, a first front right channel signal, a first back left channel signal, a first back right channel signal and a first center channel signal, the vocal signal includes a left channel vocal signal and a right channel vocal signal, and the second audio signals of the plurality of target channels include a second front left channel signal, a second front right channel signal, a second back left channel signal, a second back right channel signal and a second center channel signal; the signal generation module is further configured to:

[0144] weight-merge the first front left channel signal and the left channel vocal signal to obtain the second front left channel signal, weight-merge the first front right channel signal and the right channel vocal signal to obtain the second front right channel signal, weight the first back left channel signal to the second back left channel signal and weight the first back right channel signal to the second back right channel signal, and weight-merge the left channel vocal signal, the right channel vocal signal and the first center channel signal to obtain the second center channel signal.

[0145] In one of the embodiments, the audio upmix process is performed by an audio upmix model, and the audio upmix device further includes:

[0146] a training module configured to obtain each 5.1 channel source signal, select a target source signal from each of the 5.1 channel source signals, and extract a 5-channel target audio signal from the target source signal; downmix the 5-channel target audio signal to obtain a stereo training audio signal; extract a 5-channel left channel audio signal and a 5-channel right channel audio signal from the stereo training audio signal according to the audio upmix model, and combine the 5-channel left channel audio signal and the 5-channel right channel audio signal into a 5-channel output audio signal; and optimize the audio upmix model according to a difference between the 5-channel target audio signal and the 5-channel output audio signal.

[0147] In one of the embodiments, the training module is further configured to:

[0148] The first model loss is generated according to a signal overall difference between the 5-channel target audio signal and the 5-channel output audio signal; the second model loss is generated according to a first volume difference of different channel audio signals in the 5-channel target audio signal and a second volume difference of different channel audio signals in the 5-channel output audio signal; and the audio up-mixing model is optimized according to the first model loss and the second model loss.

[0149] The modules in the audio up-mixing device can be implemented by software, hardware, or a combination thereof. The modules can be embedded in or independent of a processor in an audio device in hardware form, or stored in a memory in the audio device in software form, so as to be invoked and executed by the processor to perform operations corresponding to the modules.

[0150] In one embodiment, an audio device, which can be a terminal, has an internal structure as shown in Figure 7 The audio device includes a processor, a memory, a communication interface, a display screen, and an input device connected by a system bus. The processor of the audio device is configured to provide computing and control capabilities. The memory of the audio device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the audio device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved by WIFI, mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program is executed by the processor to implement an audio up-mixing method.

[0151] Those skilled in the art can understand that Figure 7 The structure shown in the above figure is only a block diagram of part of the structure related to the scheme of the present application, and does not limit the audio device to which the scheme of the present application is applied. The specific audio device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0152] In one embodiment, an audio device is also provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0153] In one embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.

[0154] In an embodiment, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of any of the above method embodiments.

[0155] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0156] The technical features of the above embodiments can be combined arbitrarily. In order to make the description simple, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist, they should be considered as the scope of the present disclosure.

[0157] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. An audio upmixing method, characterized in that, The method includes: Acquire stereo audio signals and extract features from the stereo audio signals to obtain stereo audio features; Based on multiple audio output channels, multiple left channel audio signals and multiple right channel audio signals are extracted from the stereo audio features, wherein the multiple audio output channels are independent of each other, and each audio output channel is used to output the corresponding left channel audio signal or right channel audio signal. The left and right channel audio signals corresponding to the same target channel are fused to obtain the first audio signal of multiple target channels, wherein each target channel corresponds to two audio output channels. Based on the first audio signals of the plurality of target channels, an audio upmix signal in a target format is output, wherein the target format corresponds to the plurality of target channels.

2. The method according to claim 1, characterized in that, The step of extracting multiple left-channel audio signals and multiple right-channel audio signals from the stereo audio features based on multiple audio output channels includes: The azimuth audio features are extracted from the stereo audio features, wherein the azimuth audio features are used to represent the sound source signal features of different directions of the stereo audio signal; Based on the multiple audio output channels, multiple left channel audio signals and multiple right channel audio signals are extracted from the directional audio features.

3. The method according to claim 2, characterized in that, The directional audio features include the first directional sound source signal features and the second directional sound source signal features of the stereo audio signal, and the multiple audio output channels include multiple first fully connected networks and multiple second fully connected networks, each of which is different. The step of extracting multiple left-channel audio signals and multiple right-channel audio signals from the directional audio features based on the multiple audio output channels includes: Using the first directional sound source signal characteristics as input, the corresponding left channel audio signal and right channel audio signal are output through the plurality of first fully connected networks, wherein each of the first fully connected networks is used to output the left channel audio signal or the right channel audio signal. Using the second directional sound source signal characteristics as input, the corresponding left channel audio signal and right channel audio signal are output through the plurality of second fully connected networks, wherein each second fully connected network is used to output the left channel audio signal or the right channel audio signal.

4. The method according to claim 3, characterized in that, The first audio signals of the plurality of target channels include the first front left channel signal of the front left channel, the first front right channel signal of the front right channel, the first rear left channel signal of the rear left channel, the first rear right channel signal of the rear right channel, and the first center channel signal of the center channel. The process of fusing the left and right channel audio signals corresponding to the same target channel to obtain a first audio signal for multiple target channels includes: The left channel audio signal and the right channel audio signal output from the first fully connected network corresponding to the front left channel are fused to obtain the first front left channel signal. The left channel audio signal and the right channel audio signal output by the first fully connected network corresponding to the front right channel are fused to obtain the first front right channel signal. The left channel audio signal and the right channel audio signal output from the first fully connected network corresponding to the center channel are fused to obtain the first center channel signal. The left channel audio signal and the right channel audio signal output by the second fully connected network corresponding to the rear left channel are fused to obtain the first rear left channel signal; The left channel audio signal and the right channel audio signal output from the second fully connected network corresponding to the rear right channel are fused to obtain the first rear right channel signal.

5. The method according to claim 1, characterized in that, The acquisition of stereo audio signals includes: The original stereo signal is acquired, and human voice separation is performed on the original stereo signal to obtain non-human voice signal and human voice signal; The non-human voice signal is used as the stereo audio signal.

6. The method according to claim 5, characterized in that, The step of outputting an audio upmix signal in a target format based on the first audio signals of the plurality of target channels includes: The human voice signal is merged into the front left channel signal, the front right channel signal and the center channel signal of the first audio signal of the multiple target channels to obtain the second audio signal of the multiple target channels. Based on the second audio signals of the multiple target channels, output an audio upmix signal in the target format.

7. The method according to claim 6, characterized in that, The first audio signals of the multiple target channels include a first front left channel signal, a first front right channel signal, a first rear left channel signal, a first rear right channel signal, and a first center channel signal; the human voice signal includes a left channel human voice signal and a right channel human voice signal; the second audio signals of the multiple target channels include a second front left channel signal, a second front right channel signal, a second rear left channel signal, a second rear right channel signal, and a second center channel signal; merging the human voice signal into the front channel signal and the center channel signal of the first audio signals of the multiple target channels to obtain the second audio signals of the multiple target channels includes: The first front left channel signal and the left channel human voice signal are weighted and merged to obtain the second front left channel signal; The first front right channel signal and the right channel voice signal are weighted and merged to obtain the second front right channel signal; The first rear left channel signal is weighted to obtain the second rear left channel signal, and the first rear right channel signal is weighted to obtain the second rear right channel signal. The left channel voice signal, the right channel voice signal, and the first center channel signal are weighted and combined to obtain the second center channel signal.

8. The method according to claim 1, characterized in that, The audio upmixing method is executed by an audio upmixing model, and the method further includes: Acquire each 5.1 channel audio source signal, filter target audio source signals from each of the 5.1 channel audio source signals, and extract 5-channel target audio signals from the target audio source signals; The five target audio signals are downmixed to obtain stereo training audio signals; According to the audio upmixing model, five channels of left channel audio signal and five channels of right channel audio signal are extracted from the stereo training audio signal, and the five channels of left channel audio signal and the five channels of right channel audio signal are combined into a five-channel output audio signal. The audio upmixing model is optimized based on the difference between the 5-channel target audio signal and the 5-channel output audio signal.

9. The method according to claim 8, characterized in that, The step of optimizing the audio upmixing model based on the difference between the 5-channel target audio signal and the 5-channel output audio signal includes: A first model loss is generated based on the overall signal difference between the 5-channel target audio signal and the 5-channel output audio signal. A second model loss is generated based on the first volume difference of the audio signals in different channels of the 5-channel target audio signal and the second volume difference of the audio signals in different channels of the 5-channel output audio signal; The audio upmixing model is optimized based on the first model loss and the second model loss.

10. An audio device comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.