Audio upmixing method and audio device

The audio upmix method addresses the issues of reduced localization accuracy and channel correlation by independently extracting and combining left and right channel audio signals, resulting in improved spatialization and localization accuracy.

JP2025174906APending Publication Date: 2025-11-28SHENZHEN OCEANWING SMART INNOVATIONS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025080574
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-16
Filing Date
2025-05-13
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing audio upmixing methods, such as Dolby Pro Logic decoding, result in reduced localization accuracy and increased correlation between audio signals of different channels, affecting the spatialization effect of the audio up-mix signal.

Method used

An audio upmix method that involves feature extraction on a stereo audio signal to obtain stereo audio features, and then extracts a plurality of left and right channel audio signals from the stereo audio characteristics based on independent audio output channels, combining these signals to generate audio upmix signals with reduced correlation and improved localization accuracy.

Benefits of technology

The method enhances the spatialization effect of audio upmix signals by reducing correlation between channels and improving localization accuracy, thereby improving the overall effect of the audio upmixing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025174906000001_ABST
    Figure 2025174906000001_ABST
Patent Text Reader

Abstract

To provide an audio upmixing method, an apparatus, an audio device, and a computer readable storage medium capable of improving an audio upmixing effect.SOLUTION: An audio upmixing method includes steps of: obtaining a stereo audio feature; extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio feature based on a plurality of audio output channels; combining the left channel audio signal and the right channel audio signal respectively corresponding to a same target channel to obtain a first audio signal of a plurality of target channels, each target channel corresponding to two audio output channels; and outputting, based on the first audio signal of the plurality of target channels, an audio upmix signal of a target format corresponding to the plurality of target channels.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the technical field of audio processing, and in particular to an audio upmix method, device, audio equipment and storage medium. [Background technology]

[0002] With the development of audio processing technology, audio upmix technology has emerged, which can be used to convert a stereo audio signal into a desired audio upmix signal, such as a 5.1-channel signal, a 7.1-channel signal, or a 7.1.2-channel signal.

[0003] In the prior art, a desired audio upmix signal is typically generated by decoding a stereo audio signal, for example, a 5.1 channel signal can be generated by performing Dolby Pro Logic decoding on the stereo audio signal.

[0004] However, in an audio up-mix signal generated by decoding a stereo audio signal, the localization accuracy of audio signals of different channels is debatable, which may reduce the spatialization effect of the audio up-mix signal, thereby reducing the effect of the audio up-mix; meanwhile, there may be high correlation between audio signals of different channels, which may further affect the effect of the audio up-mix. Summary of the Invention

[0005] In view of the above, there is a need to provide an audio upmixing method, apparatus, audio device and computer-readable storage medium that can improve the audio upmixing effect in response to the above technical problems.

[0006] According to a first aspect, the present application provides an audio upmix method, said method comprising: obtaining a stereo audio signal and performing feature extraction on the stereo audio signal to obtain stereo audio features; extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio characteristics based on a plurality of audio output channels, each for outputting a corresponding left channel audio signal or a right channel audio signal independently of each other; respectively combining a left channel audio signal and a right channel audio signal corresponding to the same target channel to obtain a first audio signal of a plurality of target channels, each of the target channels corresponding to two of the audio output channels; and outputting an audio upmix signal in a target format corresponding to the plurality of target channels based on the first audio signals of the plurality of target channels.

[0007] In one embodiment, extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio characteristic based on a plurality of the audio output channels comprises: The method includes extracting azimuth audio features representing sound source signal features of different azimuths of the stereo audio signal from the stereo audio features, and extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the azimuth audio features based on a plurality of the audio output channels.

[0008] In one embodiment, the orientation audio features include first orientation sound source signal features and second orientation sound source signal features of the stereo audio signal, and the plurality of audio output channels include a plurality of first fully-connected networks and a plurality of second fully-connected networks, respectively, and extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the orientation audio features based on the plurality of audio output channels includes: The first direction sound source signal features are input, and the first fully-connected networks are configured to output left channel audio signals or right channel audio signals, thereby outputting corresponding left channel audio signals and right channel audio signals, and the second direction sound source signal features are input, and the second fully-connected networks are configured to output left channel audio signals or right channel audio signals, thereby outputting corresponding left channel audio signals and right channel audio signals.

[0009] In one embodiment, the first audio signals of the plurality of target channels include a first front left channel signal of a front left channel, a first front right channel signal of a front right channel, a first rear left channel signal of a rear left channel, a first rear right channel signal of a rear right channel, and a first center channel signal of a center channel, and combining the left channel audio signal and the right channel audio signal corresponding to the same target channel to obtain the first audio signals of the plurality of target channels includes: The audio processing circuit includes combining the left channel audio signal and the right channel audio signal output from a first fully-connected network corresponding to the front left channel to obtain the first front left channel signal, combining the left channel audio signal and the right channel audio signal output from a first fully-connected network corresponding to the front right channel to obtain the first front right channel signal, combining the left channel audio signal and the right channel audio signal output from a first fully-connected network corresponding to the center channel to obtain the first center channel signal, combining the left channel audio signal and the right channel audio signal output from a second fully-connected network corresponding to the rear left channel to obtain the first rear left channel signal, and combining the left channel audio signal and the right channel audio signal output from a second fully-connected network corresponding to the rear right channel to obtain the first rear right channel signal.

[0010] In one embodiment, obtaining the stereo audio signal comprises: The method includes obtaining an original stereo signal, performing voice separation on the original stereo signal to obtain a non-voice signal and a voice signal, and using the non-voice signal as the stereo audio signal.

[0011] In one embodiment, outputting an audio up-mix signal in a target format based on the first audio signals of the plurality of target channels includes: The method includes combining the voice signal with a front left channel signal, a front right channel signal, and a center channel signal among the first audio signals of the plurality of target channels to obtain second audio signals of the plurality of target channels, and outputting an audio up-mix signal in a target format based on the second audio signals of the plurality of target channels.

[0012] In one embodiment, the first audio signals of the plurality of target channels include a first front left channel signal, a first front right channel signal, a first rear left channel signal, a first rear right channel signal, and a first center channel signal, the voice signal includes a left channel voice signal and a right channel voice signal, the second audio signals of the plurality of target channels include a second front left channel signal, a second front right channel signal, a second rear left channel signal, a second rear right channel signal, and a second center channel signal, and combining the voice signal with the front channel signal and the center channel signal of the first audio signals of the plurality of target channels to obtain the second audio signals of the plurality of target channels includes: The method includes weighting and combining the first front left channel signal and the left channel voice signal to obtain the second front left channel signal, weighting and combining the first front right channel signal and the right channel voice signal to obtain the second front right channel signal, weighting the first rear left channel signal to obtain the second rear left channel signal, weighting the first rear right channel signal to obtain the second rear right channel signal, and weighting and combining the left channel voice signal, the right channel voice signal, and the first center channel signal to obtain the second center channel signal.

[0013] In one embodiment, the audio upmix method is performed by an audio upmix model, the method comprising: The method further includes obtaining each 5.1-channel sound source signal, selecting a target sound source signal from each of the 5.1-channel sound source signals, extracting 5-channel target audio signals from the target sound source signals, downmixing the 5-channel target audio signals to obtain stereo training audio signals, extracting 5-channel left channel audio signals and 5-channel right channel audio signals from the stereo training audio signals using the audio up-mix model, combining the 5-channel left channel audio signals and the 5-channel right channel audio signals into a 5-channel output audio signal, and optimizing the audio up-mix model based on a difference between the 5-channel target audio signals and the 5-channel output audio signals.

[0014] In one embodiment, optimizing the audio up-mix model based on a difference between the 5-channel target audio signal and the 5-channel output audio signal comprises: generating a first model loss based on an overall signal difference between the five-channel target audio signal and the five-channel output audio signal; generating a second model loss based on a first volume difference between different channel audio signals of the five-channel target audio signal and a second volume difference between different channel audio signals of the five-channel output audio signal; and optimizing the audio up-mix model based on the first model loss and the second model loss.

[0015] According to a second aspect, the present application further provides an audio device, the audio device including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, The method includes the steps of: acquiring a stereo audio signal, performing feature extraction on the stereo audio signal, and acquiring a stereo audio feature; extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio feature based on a plurality of audio output channels for independently outputting corresponding left channel audio signals or right channel audio signals; combining the left channel audio signal and the right channel audio signal corresponding to the same target channel, respectively, to acquire first audio signals of a plurality of target channels, where each of the target channels corresponds to two of the audio output channels; and outputting an audio up-mix signal in a target format corresponding to the plurality of target channels based on the first audio signals of the plurality of target channels.

[0016] The audio upmixing method first obtains a stereo audio signal, performs feature extraction on the stereo audio signal to obtain stereo audio features, and then extracts multiple left channel audio signals and multiple right channel audio signals from the stereo audio features based on multiple audio output channels. Each audio output channel is used to output a corresponding left channel audio signal or right channel audio signal. Because the audio output channels are independent of each other, the left channel audio signals output from the multiple audio output channels are different and do not interfere with each other. Similarly, the right channel audio signals output from the multiple audio output channels are different and do not interfere with each other. This contributes to reducing the correlation between the left channel audio signals output from the multiple audio output channels and the right channel audio signals output from the multiple audio output channels. Furthermore, because each target channel corresponds to two audio output channels, first audio signals of multiple target channels with lower correlation can be obtained by combining the left channel audio signal and the right channel audio signal corresponding to the same target channel, and audio upmix signals in target formats corresponding to the multiple target channels can be output based on the first audio signals of the multiple target channels. In this way, the left channel audio signals output from the plurality of audio output channels are different from each other and do not interfere with each other, and the right channel audio signals output from the plurality of audio output channels are different from each other and do not interfere with each other, so that the correlation between the first audio signals of the plurality of target channels can be reduced and the effect of audio upmixing can be improved.On the other hand, since each audio output channel is independent from the others, in the present application, a corresponding first audio signal is generated separately for each target channel, and the first audio signals of multiple target channels do not interfere with each other, resulting in higher localization accuracy and a more favorable spatialization effect of the output audio upmix signal in the target format, thereby further improving the audio upmix effect. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a schematic flowchart of an audio upmix method according to an embodiment of the present application; [Figure 2] 1 is a schematic flow chart illustrating a visualization of feature extraction for a stereo audio signal in accordance with an embodiment of the present application; [Figure 3] 1 is a schematic flowchart of extracting multiple left channel audio signals and multiple right channel audio signals from a stereo audio signature based on multiple audio output channels in an embodiment of the present application; [Figure 4] 1 is a visualized schematic flowchart of generating audio signals for multiple target channels based on first direction sound source signal features and second direction sound source signal features in one embodiment of the present application; [Figure 5] 1 is a visualized schematic flowchart of generating a 5.1 channel audio signal based on an original stereo signal in one embodiment of the present application; [Figure 6] 10 is a schematic flow chart visualized for generating a 5.1 channel audio signal based on an original stereo signal in another embodiment of the present application; [Figure 7] 1 is a block diagram of an audio upmixing device according to an embodiment of the present application; [Figure 8] 1 is a diagram illustrating the internal configuration of an audio device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0018] In order to make the objectives, technical means and advantages of the present application clearer, the present application will be described in more detail below with reference to the drawings and examples. It will be understood that the specific examples described herein are for illustrative purposes only and are not intended to limit the present application.

[0019] In order to improve the spatialization effect of playback, stereo audio signals are usually converted into audio up-mix signals, such as 5.1-channel audio signals, but currently, stereo audio signals are usually directly decoded into 5.1-channel audio signals using Dolby Pro Logic decoding. In this method, the localization accuracy of audio signals of different channels is questionable, which may reduce the spatialization effect of the audio up-mix signal, thereby reducing the effect of the audio up-mix. Meanwhile, there may be high correlation between audio signals of different channels, which may further affect the effect of the audio up-mix.

[0020] It should be noted that the audio upmix signal generated in the present application is not limited to a 5.1-channel audio signal, but may be a 7.1-channel audio signal or a 7.1.2-channel audio signal, etc., and users can select one according to their actual needs. The audio upmix method of the present application can be applied to audio devices, such as earphones, speakers, home theater audio devices, hearing aid devices, etc., and is not limited thereto.

[0021] In one embodiment, as shown in FIG. 1, an audio upmixing method is provided, which includes the following steps 202 to 208.

[0022] In step 202, a stereo audio signal is obtained, and feature extraction is performed on the stereo audio signal to obtain stereo audio features.

[0023] The stereo audio signal is an audio signal with a spatial effect, typically consisting of a left channel audio signal and a right channel audio signal. The stereo audio feature is a high-dimensional feature representing the signal feature of the stereo audio signal, such as a high-dimensional feature matrix. The audio device may request an external device to acquire the stereo audio signal, and then the external device transmits the stereo audio signal to the audio device. The audio device may also directly acquire the stereo audio signal from locally stored audio data.

[0024] As an example, the stereo audio signal may be an original collected stereo signal consisting of a voice signal and a non-voice signal, or the stereo audio signal may be an original stereo signal from which the voice signal has been separated, i.e., the non-voice signal within the original stereo signal.

[0025] For example, step 202 acquires a stereo audio signal, transforms the stereo audio signal from the time domain to the frequency domain to acquire a stereo frequency domain signal, divides the stereo frequency domain signal to acquire a plurality of divided signals, performs feature extraction on each of the divided signals to acquire a plurality of divided signal features, and combines the plurality of divided signal features based on the division bandwidth to acquire a stereo audio feature. In this way, because the feature extraction is performed after dividing the stereo frequency domain signal, the granularity of the feature extraction is finer and the accuracy of the feature extraction is higher, and the accuracy of the finally acquired stereo audio feature is also higher.

[0026] For example, step 202 transforms the stereo audio signal from the time domain to the frequency domain to obtain a stereo frequency domain signal, and performs feature extraction on the stereo frequency domain signal to obtain a stereo audio feature.

[0027] As an example, in this embodiment, a series of GRU modules can be used as feature extraction modules to achieve the feature extraction function. Specifically, refer to FIG. 2. FIG. 2 is a visualized schematic flowchart of feature extraction for a stereo audio signal to obtain stereo audio features in this embodiment. In this embodiment, the stereo audio signal is converted from the time domain to the frequency domain to obtain a stereo frequency domain signal. The stereo frequency domain signal is then input to a frequency division module to obtain frequency-divided signals in multiple frequency division bands. The frequency-divided signals in the multiple frequency division bands are then input to corresponding complex time-frequency encoders for feature encoding to obtain frequency-divided signal-encoded features in the multiple frequency division bands. The frequency-divided signal-encoded features in the multiple frequency division bands are then passed through a series of GRU modules (feature extraction modules) to obtain frequency-divided features extracted in the multiple frequency division bands. The frequency-divided features in the multiple frequency division bands are then passed through a Merge (feature combination module) to combine the features and obtain combined features. Finally, the combined features are passed through the GRU module to map the combined features to a preset feature dimension, thereby obtaining final stereo audio features.

[0028] In step 204, a plurality of left channel audio signals and a plurality of right channel audio signals are extracted from the stereo audio characteristics based on a plurality of audio output channels each for outputting a corresponding left channel audio signal or a right channel audio signal independently of each other.

[0029] In this embodiment, a plurality of audio output channels are set, and these audio output channels are independent of each other and can output left channel audio signals and right channel audio signals with different weights, and one audio output channel can output a corresponding left channel audio signal or right channel audio signal. In this way, since these audio output channels are independent of each other, the left channel audio signals and right channel audio signals output from each audio output channel are different from each other and have low correlation.

[0030] As an example, step 204 includes inputting the stereo audio features into a plurality of audio output channels respectively to output a plurality of left channel audio signals and a plurality of right channel audio signals, where each audio output channel is used to map the stereo audio features to a corresponding left channel audio signal or right channel audio signal, and since the mapping parameters of each audio output channel are different from each other, the plurality of left channel audio signals and the plurality of right channel audio signals thus output have low correlation.

[0031] In step 206, the left channel audio signal and the right channel audio signal corresponding to the same target channel are respectively combined to obtain a first audio signal of multiple target channels, where each target channel corresponds to two audio output channels. The number and types of the multiple target channels are determined according to the target format of the audio up-mix signal to be finally generated. For example, if the audio up-mix signal is a 5.1-channel audio signal, the multiple target channels are a front left channel, a front right channel, a rear left channel, a rear right channel, and a center channel, respectively, and first audio signals of the multiple target channels are a first front left channel signal, a first front right channel signal, a first rear left channel signal, a first rear right channel signal, and a first center channel signal, respectively.

[0032] In this embodiment, the multiple target channels correspond to the multiple audio output channels, and two audio output channels correspond to one target channel. The two audio output channels corresponding to each target channel output a left channel audio signal and a right channel audio signal, respectively. Therefore, each target channel corresponds to one left channel audio signal and one right channel audio signal.

[0033] For example, in step 206, for each target channel, the left channel audio signal and the right channel audio signal output from the two audio output channels corresponding to the target channel are obtained, the obtained left channel audio signal and the right channel audio signal are respectively weighted, and the weighted left channel audio signal and the right channel audio signal are combined to obtain a first audio signal for the target channel. In this way, in this embodiment, by setting a correspondence between the target channel and the audio output channel, the left channel audio signal and the right channel audio signal output from different audio output channels can be combined to obtain a first audio signal for the target channel. Because the audio output channels are independent of each other and do not interfere with each other, the correlation between the combined first audio signals for each target channel is low, resulting in more accurate localization and improving the audio upmix effect.

[0034] In step 208, based on the first audio signals of the plurality of target channels, an audio upmix signal in a target format corresponding to the plurality of target channels is output.

[0035] As an example, step 208 includes combining the first audio signals of the multiple target channels to obtain an audio upmix signal in the target format.

[0036] For example, the audio upmix signal in the target format is a 5.1-channel audio signal, and step 208 includes: performing low-pass filtering on an original stereo signal corresponding to the stereo audio signal to obtain a bass channel signal; and combining the bass channel signal, a first front left channel signal, a first front right channel signal, a first rear left channel signal, a first rear right channel signal, and a first center channel signal to obtain a 5.1-channel audio signal.

[0037] In the audio upmixing method, a stereo audio signal is first obtained, feature extraction is performed on the stereo audio signal to obtain a stereo audio feature, and then a plurality of left channel audio signals and a plurality of right channel audio signals are extracted from the stereo audio feature based on a plurality of audio output channels. Each audio output channel is used to output a corresponding left channel audio signal or a right channel audio signal. Because the audio output channels are independent from each other, the left channel audio signals output from the plurality of audio output channels are different from each other and do not interfere with each other. Similarly, the right channel audio signals output from the plurality of audio output channels are different from each other and do not interfere with each other. This contributes to reducing the correlation between the left channel audio signals output from the plurality of audio output channels and the right channel audio signals output from the plurality of audio output channels. Furthermore, because each target channel corresponds to two audio output channels, first audio signals of a plurality of target channels with lower correlation can be obtained by respectively combining the left channel audio signal and the right channel audio signal corresponding to the same target channel. Audio upmix signals in target formats corresponding to the plurality of target channels can be output based on the first audio signals of the plurality of target channels. In this way, the left channel audio signals output from the plurality of audio output channels are different from each other and do not interfere with each other, and the right channel audio signals output from the plurality of audio output channels are different from each other and do not interfere with each other, so that the correlation between the first audio signals of the plurality of target channels can be reduced and the effect of audio upmixing can be improved.On the other hand, since each audio output channel is independent from the others, in the present application, a corresponding first audio signal is generated separately for each target channel, and the first audio signals of multiple target channels do not interfere with each other, resulting in higher localization accuracy and a more favorable spatialization effect of the output audio upmix signal in the target format, thereby further improving the audio upmix effect.

[0038] In one embodiment, as shown in FIG. 3 , for example, extracting a left channel audio signal output from a plurality of audio output channels and a right channel audio signal output from a plurality of audio output channels from a stereo audio characteristic further includes step 302 and step 304.

[0039] In step 302, directional audio features representing source signal features of different directional directions of the stereo audio signal are extracted from the stereo audio features; The number of azimuth audio features corresponding to the stereo audio features may be one or more, and is not limited here. For example, the azimuth audio feature may be a first azimuth sound source signal feature representing a front azimuth audio feature, and a second azimuth sound source signal feature representing a rear azimuth audio feature.

[0040] As an example, step 302 includes extracting corresponding azimuth audio features from the stereo audio features based on at least one predetermined azimuth feature extraction module for extracting azimuth audio features representing sound source signal features in a predetermined azimuth from the stereo audio features.

[0041] As an example, the predetermined orientation feature extraction module may be a GRU (Gated Recurrent Unit) module.

[0042] In step 304, a plurality of left channel audio signals and a plurality of right channel audio signals are extracted from the azimuth audio feature based on the plurality of audio output channels.

[0043] As an example, step 304 maps, for each orientation audio feature, the orientation audio feature to a corresponding left channel audio signal and a corresponding right channel audio signal using the audio output channel corresponding to the orientation audio feature as an input.

[0044] As an example, if the finally generated audio upmix signal is a 5.1-channel audio signal, the different orientations may be a front-left orientation (corresponding to the front-left channel of the 5.1-channel), a front-right orientation (corresponding to the front-right channel of the 5.1-channel), a rear-left orientation (corresponding to the rear-left channel of the 5.1-channel), a rear-right orientation (corresponding to the rear-right channel of the 5.1-channel), and a center orientation (corresponding to the center channel of the 5.1-channel).

[0045] In this manner, in this embodiment, the left channel audio signal and the right channel audio signal are generated by dividing the stereo audio feature into at least one azimuth audio feature, which prevents interference between the sound source signal features of each azimuth in the stereo audio feature and contributes to improving the accuracy of the finally generated left channel audio signal and right channel audio signal.

[0046] In one embodiment, the orientation audio features include first orientation sound source signal features and second orientation sound source signal features of the stereo audio signal, and each predetermined orientation feature extraction module includes a first orientation feature extraction module and a second orientation feature extraction module. Extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the orientation audio features based on the plurality of audio output channels includes: The method includes extracting a first direction sound source signal feature from the stereo audio feature by a first direction feature extraction module, and extracting a second direction sound source signal feature from the stereo audio feature by a second direction feature extraction module.

[0047] As an example, the first azimuth sound source signal feature may be an azimuth signal feature representing a front sound source signal feature of the stereo audio signal, and the second azimuth sound source signal feature may be an azimuth signal feature representing a rear sound source signal feature of the stereo audio signal.

[0048] As an example, the first azimuth sound source signal feature may be an azimuth signal feature representing a left front sound source signal feature of the stereo audio signal, and the second azimuth sound source signal feature may be an azimuth signal feature representing a right rear sound source signal feature of the stereo audio signal.

[0049] The first direction sound source signal feature and the second direction sound source signal feature may represent sound source signal features of a stereo audio signal in any direction, but the first direction sound source signal feature and the second direction sound source signal feature may represent sound source signal features of a stereo audio signal in any direction.

[0050] In one embodiment, the azimuth audio features include first azimuth sound source signal features and second azimuth sound source signal features of the stereo audio signal, and the plurality of audio output channels include a plurality of first fully-connected networks and a plurality of second fully-connected networks, respectively, and extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the azimuth audio features based on the plurality of audio output channels includes: The method includes: receiving the first directional sound source signal features as input, passing the signals through a plurality of first fully-connected networks for outputting left channel audio signals or right channel audio signals, thereby outputting corresponding left channel audio signals and right channel audio signals; and receiving the second directional sound source signal features as input, passing the signals through a plurality of second fully-connected networks for outputting left channel audio signals or right channel audio signals, thereby outputting corresponding left channel audio signals and right channel audio signals.

[0051] Each audio output channel may include one first fully-connected network or one second fully-connected network, and all of the first fully-connected networks and all of the second fully-connected networks are different from each other.

[0052] Specifically, for each first fully-connected network, the first direction sound source signal features are input to the first fully-connected network, the first fully-connected network fully connects the first direction sound source signal features, and after the full connection, a left channel audio signal or a right channel audio signal is obtained as an output from the first fully-connected network using a preset activation function; for each second fully-connected network, the second direction sound source signal features are input to the second fully-connected network, the second fully-connected network fully connects the second direction sound source signal features, and after the full connection, a left channel audio signal or a right channel audio signal is obtained as an output from the second fully-connected network using a preset activation function; In this embodiment, first, a first direction sound source signal feature and a second direction sound source signal feature of the stereo audio signal are separated from the stereo audio feature. Next, left channel audio signals and right channel audio signals output from the plurality of audio output channels are extracted for the first direction sound source signal feature, and left channel audio signals and right channel audio signals are extracted for the second direction sound source signal feature. After separating the sound source signal features of different directions from the stereo audio feature, the left channel audio signals and right channel audio signals can be generated. In the process of generating the left channel audio signals and the right channel audio signals, the sound source signal features of different directions from the stereo audio feature do not interfere with each other, thereby improving the accuracy of the generated left channel audio signals and right channel audio signals.

[0053] Note that the greater the number of orientation audio features divided from the stereo audio features, the higher the accuracy of the finally generated left channel audio signals and right channel audio signals, but correspondingly, the greater the number of orientation audio features divided from the stereo audio features, the lower the efficiency of the generated left channel audio signals and right channel audio signals. In this embodiment, the stereo audio features are divided into first orientation sound source signal features and second orientation sound source signal features that respectively output two orientations, and the left channel audio signals and right channel audio signals that are output from multiple audio output channels are generated using the first orientation sound source signal features and the second orientation sound source signal features. This makes it possible to achieve both efficiency and accuracy in generating the left channel audio signals and the right channel audio signals.

[0054] In one embodiment, for example, when a 5.1-channel audio signal is finally generated, the first audio signals of the plurality of target channels include a first front left channel signal of the front left channel, a first front right channel signal of the front right channel, a first rear left channel signal of the rear left channel, a first rear right channel signal of the rear right channel, and a first center channel signal of the center channel. The left channel audio signal and the right channel audio signal corresponding to the same target channel are combined to obtain the first audio signals of the plurality of target channels, which is: The method includes combining the left channel audio signal and the right channel audio signal output from the first fully-coupled network corresponding to the front left channel to obtain a first front left channel signal, combining the left channel audio signal and the right channel audio signal output from the first fully-coupled network corresponding to the front right channel to obtain a first front right channel signal, combining the left channel audio signal and the right channel audio signal output from the first fully-coupled network corresponding to the center channel to obtain a first center channel signal, combining the left channel audio signal and the right channel audio signal output from the second fully-coupled network corresponding to the rear left channel to obtain a first rear left channel signal, and combining the left channel audio signal and the right channel audio signal output from the second fully-coupled network corresponding to the rear right channel to obtain a first rear right channel signal.

[0055] Specifically, based on the correspondence between the target channel and the audio output channel, a left channel audio signal and a right channel audio signal corresponding to a front left channel, a left channel audio signal and a right channel audio signal corresponding to a front right channel, and a left channel audio signal and a right channel audio signal corresponding to a center channel are determined from the left channel audio signals and the right channel audio signals output from the different first fully connected networks, and the left channel audio signal and the right channel audio signal corresponding to the front left channel are weighted and combined to obtain a first front left channel signal, and the left channel audio signal and the right channel audio signal corresponding to the front right channel are weighted and combined to obtain a first front right channel signal, and the center a first center channel signal is obtained by weighting and combining the left channel audio signal and the right channel audio signal corresponding to the target channel; a left channel audio signal and a right channel audio signal corresponding to the rear left channel and the rear right channel are determined from the left channel audio signals and right channel audio signals output from different second fully connected networks based on the correspondence between the target channel and the audio output channel; the left channel audio signal and the right channel audio signal corresponding to the rear left channel are weighted and combined to obtain a first rear left channel signal; and the left channel audio signal and the right channel audio signal corresponding to the rear right channel are weighted and combined to obtain a first rear right channel signal.

[0056] In this embodiment, different first fully-connected networks and different second fully-connected networks are configured to generate left and right channel audio signals that are different from each other and do not interfere with each other. Thus, a first front left channel signal, a first front right channel signal, and a first center channel signal are generated by weighting and combining the left and right channel audio signals output from the different first fully-connected networks, respectively. A first rear left channel signal and a first rear right channel signal are generated by weighting and combining the left and right channel audio signals output from the different second fully-connected networks, respectively. By combining the different and non-interfering left and right channel audio signals, the combined first front left channel signal, first front right channel signal, and first center channel signal, as well as the left and right channel audio signals, are not identical to each other and do not interfere with each other. This improves the spatial localization accuracy of 5.1 channel audio signals and reduces correlation between audio signals of different channels, thereby improving the audio upmixing effect.

[0057] In one embodiment, as shown in Figure 4, Figure 4 is a visualized schematic flowchart of generating audio signals for each target channel of a 5.1 channel audio signal based on first direction sound source signal features and second direction sound source signal features in one embodiment. In the figure, GRU_F is a first direction feature extraction module that extracts first direction sound source signal features of the stereo audio signal from the stereo audio features, GRU_S is a second direction feature extraction module that extracts second direction sound source signal features of the stereo audio signal from the stereo audio features, the FC modules connected to GRU_F are each a first fully connected network followed by a Sigmoid activation function, and the FC modules connected to GRU_S are each a second fully connected network followed by a Sigmoid activation function. Sigmoids are connected, L represents the output left channel audio signal, R represents the output right channel audio signal, FL represents the first front left channel signal, FR represents the first front right channel signal, C represents the first center channel signal, SL represents the first rear left channel signal, and SR represents the first rear right channel signal. A circle plus sign represents weighted combination. For example, the weighted combination may be performed by multiplying the output left channel signal and right channel signal by corresponding masks and then adding them.

[0058] In one embodiment, obtaining a stereo audio signal comprises: The method includes obtaining an original stereo signal, performing voice separation on the original stereo signal, obtaining a non-voice signal and a voice signal, and converting the non-voice signal into a stereo audio signal.

[0059] In order to preserve the integrity of the voice, in this embodiment, before performing audio downmixing, the voice signal may be separated from the original stereo signal first, so that a stereo audio signal without the voice signal can be obtained.

[0060] As an example, in this embodiment, the original stereo signal may be directly converted into a stereo audio signal, in which case a voice signal will be present in the stereo audio signal.

[0061] In one embodiment, outputting an audio upmix signal in a target format based on first audio signals of a plurality of target channels includes: The method includes combining the human voice signal with a front left channel signal, a front right channel signal, and a center channel signal among the first audio signals of the plurality of target channels to obtain second audio signals of the plurality of target channels, and outputting an audio up-mix signal in the target format based on the second audio signals of the plurality of target channels.

[0062] In an audio upmix signal, a voice signal is usually present in the front left channel, the front right channel, and the center channel, thereby providing a better auditory experience to the user. Taking a 5.1-channel audio signal as an example, a voice signal is usually present in the front left channel, the front right channel, and the center channel, but not in the rear left channel and the rear right channel, resulting in a better auditory experience for the 5.1-channel audio signal. Therefore, after obtaining first audio signals for multiple target channels and before finally generating an audio upmix signal in a target format, a voice signal is combined with the first audio signals for multiple target channels. In this way, although a complete voice signal is present in the finally generated audio upmix signal in a target format, distortion of the voice signal during audio upmixing can be avoided because the voice signal has not been subjected to a series of audio upmixing operations on a stereo audio signal, and the true integrity of the voice signal is ensured, thereby further improving the audio upmix effect.

[0063] Specifically, the voice signal is weighted and combined with the front left channel signal, the front right channel signal, and the center channel signal among the first audio signals of the plurality of target channels to obtain second audio signals of the plurality of target channels, and the second audio signals of the plurality of target channels are combined to form an audio up-mix signal in the target format.

[0064] In this embodiment, after separating a voice signal from an original stereo signal and generating first audio signals for a plurality of target channels based on the stereo audio signal obtained by separating the voice signal, the voice signal is combined with only the front left channel signal, the front right channel signal, and the center channel signal of the first audio signals for the plurality of target channels, thereby realizing selective audio rendering for the first audio signals for the plurality of target channels. The generated second audio signals for the plurality of target channels are more in line with the user's hearing habits, resulting in a better hearing experience, and thus improving the effect of audio upmixing.

[0065] In one embodiment, taking the case where the audio upmix signal of the target format is a 5.1-channel audio signal as an example, the first audio signals of the multiple target channels include a first front left channel signal, a first front right channel signal, a first center channel signal, a first rear left channel signal, and a first rear right channel signal. Since the voice signal is also a stereo signal, the voice signal includes a left-channel voice signal and a right-channel voice signal. The second audio signals of the multiple target channels include a second front left channel signal, a second front right channel signal, a second center channel signal, a second rear left channel signal, and a second rear right channel signal. Combining the voice signal with the front channel signal and the center channel signal of the first audio signals of the multiple target channels to obtain the second audio signals of the multiple target channels can be performed as follows: The method includes weighting and combining the first front left channel signal and the left channel voice signal to obtain a second front left channel signal, weighting and combining the first front right channel signal and the right channel voice signal to obtain a second front right channel signal, weighting the first rear left channel signal to obtain a second rear left channel signal, weighting the first rear right channel signal to obtain a second rear right channel signal, and weighting and combining the left channel voice signal, the right channel voice signal, and the first center channel signal to obtain a second center channel signal.

[0066] Specifically, weight the first front left channel signal with a first preset weight, weight the left channel voice signal with a second preset weight, add the weighted first front left channel signal and the weighted left channel voice signal to obtain a second front left channel signal, weight the first front right channel signal with a first preset weight, weight the right channel voice signal with a second preset weight, add the weighted first front right channel signal and the weighted right channel voice signal to obtain a second front right channel signal, and add the weighted first front right channel signal and the weighted right channel voice signal to obtain a third preset weight. The first rear left channel signal and the first rear right channel signal are weighted respectively to obtain a second rear left channel signal corresponding to the first rear left channel signal and a second rear right channel signal corresponding to the first rear right channel signal; the first center channel signal is weighted by a fourth preset weight; the left channel voice signal and the right channel voice signal are weighted by a fifth preset weight; and the weighted first center channel signal, the weighted left channel voice signal, and the weighted right channel voice signal are added to obtain a second center channel signal.

[0067] Furthermore, low-pass filtering may be performed on the original stereo signal, and the low-pass filtered original stereo signal may be weighted with a sixth preset weight to obtain a bass channel signal, so that a second front left channel signal, a second front right channel signal, a second center channel signal, a second rear left channel signal, a second rear right channel signal, and a bass channel signal can all be obtained as a 5.1 channel audio signal.

[0068] As an example, the first preset weight represents the importance of non-human voices in the front surround channels (including the front left channel and the front right channel) to the 5.1 channel audio signal, and the more important the non-human voices in the front surround channels are, the higher the first preset weight is. The second preset weight represents the importance of human voices in the front surround channels to the 5.1 channel audio signal, and the more important the human voices in the front surround channels are, the higher the second preset weight is. The third preset weight represents the importance of non-human voices in the rear surround channels (including the rear left channel and the rear right channel) to the 5.1 channel audio signal. The third preset weight represents the importance of non-human voices in the rear surround channels to the 5.1 channel audio signal, and the fourth preset weight represents the importance of non-human voices in the center channel to the 5.1 channel audio signal, and the fourth preset weight represents the importance of non-human voices in the center channel to the 5.1 channel audio signal, and the fourth preset weight represents the importance of non-human voices in the center channel to the 5.1 channel audio signal, and the fifth preset weight represents the importance of human voices in the center channel to the 5.1 channel audio signal, and the fifth preset weight represents the importance of human voices in the center channel to the 5.1 channel audio signal, and the sixth preset weight represents the importance of bass channel signals ....

[0069] In one embodiment, a case where the finally generated audio upmix signal in the target format is a 5.1-channel audio signal will be described as an example with reference to Fig. 5. Fig. 5 is a visualized schematic flowchart of generating a 5.1-channel audio signal based on an original stereo signal in one embodiment, in which stereo L and R are the original stereo signals, voice VL is a left-channel voice signal, voice VR is a right-channel voice signal, and non-voice is a stereo audio signal, and the AI ​​upmix module generates first audio signals of multiple target channels based on the non-voice signals, the first audio signals of the multiple target channels including a first front left channel signal O_FL, a first front right channel signal O_FR, a first rear left channel signal O_SL, a first rear right channel signal O_RL, and a first center channel signal O_C, and the LPF where F_Gain is the first preset weighting, V_Gain is the second preset weighting, S_Gain is the third preset weighting, C_Gain1 is the fourth preset weighting, C_Gain2 is the fifth preset weighting, Bass_Gain is the sixth preset weighting, the circle plus sign represents addition, FL is the second front left channel signal, FR is the second front right channel signal, SL is the second rear left channel signal, RL is the second rear right channel signal, C is the second center channel signal, and Bass is the bass channel signal; thus, FL, FR, SL, RL, C and Bass together constitute a 5.1 channel audio signal.

[0070] In one embodiment, the case where the finally generated audio upmix signal in the target format is a 5.1-channel audio signal will be described as an example with reference to Fig. 6. Fig. 6 is a visualized schematic flowchart of generating a 5.1-channel audio signal based on an original stereo signal in another embodiment, in which the voice signal may not be separated from the original stereo signal and the original stereo signal may be directly used as a stereo audio signal, where stereo L and R are stereo audio signals, the AI ​​upmix module generates first audio signals of multiple target channels based on the stereo, the first audio signals of the multiple target channels include a first front left channel signal O_FL, a first front right channel signal O_FR, a first rear left channel signal O_SL, a first rear right channel signal O_RL, and a first center channel signal O_C, the LPF is a low-pass filter that directly uses a first preset weight F_Gain to filter the first front the front left channel signal O_FL is weighted to a second front left channel signal FL, the first front right channel signal O_FR is weighted to a second front right channel signal FR, the first rear left channel signal O_SL is weighted directly by a third preset weighting S_Gain to a second rear left channel signal SL, the first rear right channel signal O_RL is weighted directly by a fourth preset weighting C_Gain1 to a second center channel signal C, and the low-pass filtered stereo L, R are weighted to a bass channel signal Bass by a sixth preset weighting Bass_Gain; thus, FL, FR, SL, RL, C, and Bass together constitute a 5.1 channel audio signal.

[0071] In one embodiment, the audio upmix method is performed by an audio upmix model, said audio upmix method comprising: The method further includes obtaining each 5.1-channel sound source signal, selecting a target sound source signal from each 5.1-channel sound source signal, extracting a 5-channel target audio signal from the target sound source signal, downmixing the 5-channel target audio signal to obtain a stereo training audio signal, extracting a 5-channel left channel audio signal and a 5-channel right channel audio signal from the stereo training audio signal using an audio upmix model, combining the 5-channel left channel audio signal and the 5-channel right channel audio signal into a 5-channel output audio signal, and optimizing the audio upmix model based on a difference between the 5-channel target audio signal and the 5-channel output audio signal.

[0072] The audio upmix process of steps 202 to 208 may be performed by an audio upmix model. However, when training the audio upmix model, in order to ensure the training effect of the audio upmix model, it is usually necessary to select a target sound source signal with a high training effect from a large number of 5.1 channel sound source signals to train the audio upmix model.

[0073] Specifically, each 5.1-channel sound source signal is obtained, a target sound source signal is selected from each 5.1-channel sound source signal, a 5-channel target audio signal is extracted from the target sound source signal, downmixing is performed on the 5-channel target audio signal using a preset downmix matrix, the stereo signal obtained by the downmixing is used as a stereo training audio signal, and an audio upmixing process is performed based on an audio upmix model. The audio upmix process transforms the stereo training audio signals from the time domain to the frequency domain to obtain stereo frequency-domain training signals, divides the stereo frequency-domain training signals to obtain a plurality of divided training signals, performs feature extraction on the plurality of divided training signals respectively to obtain a plurality of divided training signal features, combines the plurality of divided training signal features based on the division bandwidth to obtain stereo training audio features, inputs the stereo training audio features into 10 audio output channels respectively, outputs 5-channel left channel audio signals and 5-channel right channel audio signals, weights and combines the 5-channel left channel audio signals and the 5-channel right channel audio signals to obtain a 5-channel output audio signal, and iteratively updates and optimizes the audio upmix model based on a model loss calculated based on the difference between the 5-channel target audio signals and the 5-channel output audio signals.

[0074] For the audio upmix process performed by the audio upmix model, please refer to the audio upmix process in the embodiment of the audio upmix method, and a detailed description will be omitted here.

[0075] As an example, extracting a five-channel target audio signal from a target sound source signal is The method includes removing the bass channel signal from the target sound source signal to obtain a five-channel target audio signal.

[0076] As an example, extracting a five-channel target audio signal from a target sound source signal is This includes removing the bass channel signal and the voice signal from the target sound source signal to obtain a 5-channel target audio signal.

[0077] The 5-channel target audio signals may be a front left channel signal, a front right channel signal, a rear left channel signal, a rear right channel signal, and a center channel signal of a 5.1 audio signal.

[0078] In one embodiment, optimizing the audio up-mix model based on a difference between the five-channel target audio signal and the five-channel output audio signal comprises: The method includes generating a first model loss based on an overall signal difference between the five-channel target audio signal and the five-channel output audio signal; generating a second model loss based on a volume difference between different channel audio signals of the five-channel output audio signal; and optimizing an audio upmix model based on the first model loss and the second model loss.

[0079] Specifically, a first model loss is calculated based on an overall signal difference between the 5-channel target audio signal and the 5-channel output audio signal, a second model loss is calculated based on a first volume difference between different channel audio signals of the 5-channel target audio signal and a second volume difference between different channel audio signals of the 5-channel output audio signal, the first model loss and the second model loss are added together to obtain a total model loss, and the audio upmix model is iteratively updated and optimized based on gradient information calculated from the total model loss.

[0080] As an example, the formula for calculating the first model loss based on the overall signal difference between the 5-channel target audio signal and the 5-channel output audio signal is as follows:

[0081] L1=20log(target / estimate-target) where L1 is the first model loss, target is the 5-channel target audio signal, and estimate is the 5-channel output audio signal.

[0082] For example, calculating the second model loss based on a first volume difference between different channel audio signals of the five-channel target audio signal and a second volume difference between different channel audio signals of the five-channel output audio signal may include: The method includes calculating first volume difference values ​​between different channel audio signals of the five-channel target audio signals, calculating second volume difference values ​​between different channel audio signals of the five-channel output audio signals, calculating a difference accumulation result between the first volume difference value and the second volume difference value of each group corresponding to the same channel audio signal, and setting the accumulation result as a second model loss.

[0083] As an example, the method for selecting the target sound source signal from each 5.1 channel sound source signal includes at least one of the following methods.

[0084] The first method is to select a target sound source signal from each 5.1 channel sound source signal based on the correlation between the front left channel signal and the rear left channel signal and the correlation between the front right channel signal and the rear right channel signal of each 5.1 channel sound source signal.

[0085] For a 5.1 channel sound source signal, if the correlation between the front left channel signal and the rear left channel signal is too high, or if the correlation between the front right channel signal and the rear right channel signal is too high, the auditory effect of the 5.1 channel sound source signal will be poor.

[0086] Specifically, for each 5.1-channel sound source signal, a first correlation value between the front left channel signal and the rear left channel signal of the 5.1-channel sound source signal is calculated, and a second correlation value between the front right channel signal and the rear right channel signal of the 5.1-channel sound source signal is calculated. If both the first correlation value and the second correlation value are smaller than a preset correlation threshold, the 5.1-channel sound source signal is determined as a target sound source signal. If both the first correlation value and the second correlation value are equal to or greater than the preset correlation threshold, the 5.1-channel sound source signal is not determined as a target sound source signal.

[0087] The second method is to select a target sound source signal from each 5.1 channel sound source signal based on the signal energy of the rear surround channel signal among the 5.1 channel sound source signals.

[0088] For a 5.1 channel sound source signal, if the signal energy of the rear surround channel signal is relatively small, the spatial hearing sensation of the rear surround of the 5.1 channel sound source signal will be lost, resulting in a poor hearing effect.

[0089] Specifically, for each 5.1-channel sound source signal, the signal energy of the rear surround channel signal of the 5.1-channel sound source signal is determined. If the signal energy is greater than a preset signal energy threshold, the 5.1-channel sound source signal is set as the target sound source signal. If the signal energy is equal to or less than the preset signal energy threshold, the 5.1-channel sound source signal is not set as the target sound source signal.

[0090] In this embodiment, an audio upmix model can be accurately trained and obtained for performing the audio upmix process. The model loss can be reasonably set based on the overall signal difference between the 5-channel target audio signal and the 5-channel output audio signal, the first volume difference between the different channel audio signals of the 5-channel target audio signal, and the second volume difference between the different channel audio signals of the 5-channel output audio signal, so that the audio upmix model can learn a more accurate 5-channel audio signal and ultimately generate a more accurate 5.1-channel audio signal. Meanwhile, the 5.1-channel source signal can be accurately selected, and the 5.1-channel source signal selected has a low correlation between the front left channel signal and the rear left channel signal and a low correlation between the front right channel signal and the rear right channel signal, or has a large signal energy of the rear surround channel signal, so that the quality of the training samples for the audio upmix model can be improved, the trained audio upmix model can be made more accurate, and a foundation for improving the audio upmix effect can be laid.

[0091] Although the steps in the flowcharts according to the above embodiments are indicated in order by arrows, it should be understood that these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated otherwise in this specification, the execution of these steps is not limited to a strict order, and these steps may be executed in other orders. Furthermore, at least some of the steps in the flowcharts according to the above embodiments may include multiple steps or multiple stages, and these steps or stages do not necessarily have to be executed at the same time but may be executed at different times. The order in which these steps or stages are executed is also not necessarily sequential, and they may be executed in order or alternately with other steps or at least some of the steps or stages within other steps.

[0092] Based on the same inventive idea, embodiments of the present application further provide an audio upmixing device for implementing the above audio upmixing method. Since the solution provided by the device is similar to the implementation described in the above method, specific limitations of one or more embodiments of the audio upmixing device provided below can refer to the limitations of the above audio upmixing method, and further description will be omitted here.

[0093] In one embodiment, as shown in FIG. 7 , the audio upmixing device includes a stereo feature extraction module 702, a signal extraction module 704, a combination module 706, and a signal generation module 708; a stereo feature extraction module 702 for obtaining a stereo audio signal and performing feature extraction on the stereo audio signal to obtain stereo audio features; a signal extraction module 704 extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio characteristics based on a plurality of audio output channels, each for outputting a corresponding left channel audio signal or a right channel audio signal independently of each other; the combining module 706 respectively combines the left channel audio signal and the right channel audio signal corresponding to the same target channel to obtain a first audio signal of multiple target channels, each of the target channels corresponding to two of the audio output channels; The signal generation module 708 outputs an audio upmix signal in a target format corresponding to the plurality of target channels based on the first audio signals of the plurality of target channels.

[0094] In one embodiment, the signal extraction module further comprises: An orientation audio feature representing sound source signal features of different orientations of the stereo audio signal is extracted from the stereo audio feature, and a plurality of left channel audio signals and a plurality of right channel audio signals are extracted from the orientation audio feature based on a plurality of the audio output channels.

[0095] In one embodiment, the directional audio features include a first directional sound source signal feature and a second directional sound source signal feature of the stereo audio signal, and the plurality of audio output channels respectively include a plurality of first fully-connected networks and a plurality of second fully-connected networks, and the signal extraction module further comprises: The first direction sound source signal features are input, and the first fully-connected networks are configured to output left channel audio signals or right channel audio signals, thereby outputting corresponding left channel audio signals and right channel audio signals, and the second direction sound source signal features are input, and the second fully-connected networks are configured to output left channel audio signals or right channel audio signals, thereby outputting corresponding left channel audio signals and right channel audio signals.

[0096] In one embodiment, the first audio signals of the plurality of target channels include a first front left channel signal of a front left channel, a first front right channel signal of a front right channel, a first rear left channel signal of a rear left channel, a first rear right channel signal of a rear right channel, and a first center channel signal of a center channel, and the signal extraction module further comprises: The left channel audio signal and the right channel audio signal output from the first fully-connected network corresponding to the front left channel are combined to obtain the first front left channel signal, the left channel audio signal and the right channel audio signal output from the first fully-connected network corresponding to the front right channel are combined to obtain the first front right channel signal, the left channel audio signal and the right channel audio signal output from the first fully-connected network corresponding to the center channel are combined to obtain the first center channel signal, the left channel audio signal and the right channel audio signal output from the second fully-connected network corresponding to the rear left channel are combined to obtain the first rear left channel signal, and the left channel audio signal and the right channel audio signal output from the second fully-connected network corresponding to the rear right channel are combined to obtain the first rear right channel signal.

[0097] In one embodiment, the stereo feature extraction module further comprises: The method includes obtaining an original stereo signal, performing voice separation on the original stereo signal to obtain a non-voice signal and a voice signal, and using the non-voice signal as the stereo audio signal.

[0098] In one embodiment, the signal generation module further comprises: The audio signal is combined with a front left channel signal, a front right channel signal, and a center channel signal among the first audio signals of the plurality of target channels to obtain second audio signals of the plurality of target channels, and an audio up-mix signal in a target format is output based on the second audio signals of the plurality of target channels.

[0099] In one embodiment, the first audio signals of the plurality of target channels include a first front left channel signal, a first front right channel signal, a first rear left channel signal, a first rear right channel signal, and a first center channel signal; the voice signal includes a left channel voice signal and a right channel voice signal; the second audio signals of the plurality of target channels include a second front left channel signal, a second front right channel signal, a second rear left channel signal, a second rear right channel signal, and a second center channel signal; and the signal generating module further comprises: The first front left channel signal and the left channel voice signal are weighted and combined to obtain the second front left channel signal, the first front right channel signal and the right channel voice signal are weighted and combined to obtain the second front right channel signal, the first rear left channel signal is weighted to obtain the second rear left channel signal, the first rear right channel signal is weighted to obtain the second rear right channel signal, and the left channel voice signal, the right channel voice signal and the first center channel signal are weighted and combined to obtain the second center channel signal.

[0100] In one embodiment, the audio upmix process is performed by an audio upmix model, and the audio upmix device further comprises a training module; The training module obtains each 5.1-channel sound source signal, selects a target sound source signal from each of the 5.1-channel sound source signals, extracts 5-channel target audio signals from the target sound source signals, and downmixes the 5-channel target audio signals to obtain stereo training audio signals. The training module extracts 5-channel left channel audio signals and 5-channel right channel audio signals from the stereo training audio signals using the audio up-mix model, combines the 5-channel left channel audio signals and the 5-channel right channel audio signals into a 5-channel output audio signal, and optimizes the audio up-mix model based on a difference between the 5-channel target audio signals and the 5-channel output audio signals.

[0101] In one embodiment, the training module further comprises: A first model loss is generated based on an overall signal difference between the five-channel target audio signal and the five-channel output audio signal, a second model loss is generated based on a first volume difference between different channel audio signals of the five-channel target audio signal and a second volume difference between different channel audio signals of the five-channel output audio signal, and the audio up-mix model is optimized based on the first model loss and the second model loss.

[0102] The modules in the audio upmixing device may be implemented in whole or in part by software, hardware, or a combination thereof. The modules may be built into a processor in an audio device in a hardware form, may be independent, or may be stored in a memory in an audio device in a software form, so that the processor can easily perform the operations corresponding to the modules.

[0103] In one embodiment, an audio device is provided, which may be a terminal, and its internal configuration may be as shown in FIG. 7. The audio device includes a processor, a memory, a communication interface, a display screen, and an input device, all connected via a system bus. The processor of the audio device provides calculation and control functions. The memory of the audio device includes a non-volatile storage medium and an internal memory. An operating system and a computer program are stored in the non-volatile storage medium. The internal memory provides an environment for the execution of the operating system and the computer program stored in the non-volatile storage medium. The communication interface of the audio device communicates with an external terminal via a wired or wireless method, and the wireless method can be realized by Wi-Fi, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, an audio upmixing method is realized.

[0104] Those skilled in the art will understand that the structure shown in FIG. 7 is merely a block diagram of some structures related to the solution of the present application, and does not limit the audio equipment to which the solution of the present application is applied, and that a specific audio equipment may include more or fewer components than those shown, may combine some components, or may have a different component arrangement.

[0105] In one embodiment, an audio device is further provided, including a memory in which a computer program is stored, and a processor, which, when executed by the processor, performs the steps of each of the method embodiments described above.

[0106] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of each of the method embodiments described above.

[0107] In one embodiment, a computer program product is provided that includes a computer program that, when executed by a processor, implements the steps of each of the method embodiments above.

[0108] Those skilled in the art will understand that all or part of the flow of the methods in the above embodiments can be realized by a computer program instructing associated hardware, and that the computer program can be stored in a non-volatile computer-readable storage medium and, when executed, can include the flow of each of the above method embodiments. The memory, database, or other medium mentioned in each embodiment of the present application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM®), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, the RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database in the embodiments may include at least one of a relational database and a non-relational database. The non-relational database may include, but is not limited to, a distributed database based on blockchain. The processor in the embodiments provided herein may be, but is not limited to, a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, programmable logic, data processing logic based on quantum computing, etc.

[0109] The technical features of the above embodiments can be combined in any manner, and for the sake of convenience, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to fall within the scope described in this specification.

[0110] The above examples merely describe some embodiments of the present application, and although the descriptions are specific and detailed, they should not be understood as limiting the scope of the claims of the present application. It should be noted that a person skilled in the art may make some modifications and improvements without departing from the concept of the present application, which also fall within the scope of protection of the present application. Therefore, the scope of protection of the present application should be determined based on the scope of the claims.

Claims

1. obtaining a stereo audio signal and performing feature extraction on the stereo audio signal to obtain stereo audio features; extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio characteristics based on a plurality of audio output channels, each for outputting a corresponding left channel audio signal or a right channel audio signal independently of each other; respectively combining a left channel audio signal and a right channel audio signal corresponding to the same target channel to obtain a first audio signal of a plurality of target channels, each of the target channels corresponding to two of the audio output channels; outputting an audio upmix signal in a target format corresponding to the plurality of target channels based on first audio signals of the plurality of target channels; 1. An audio upmixing method comprising:

2. extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the stereo audio characteristic based on a plurality of the audio output channels, extracting azimuth audio features from the stereo audio signal, the azimuth audio features representing sound source signal features of different azimuths of the stereo audio signal; extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the orientation audio feature based on the plurality of audio output channels; 2. The method of claim 1, comprising:

3. the azimuth audio features include first azimuth sound source signal features and second azimuth sound source signal features of the stereo audio signal, and the plurality of audio output channels respectively include a plurality of first fully-connected networks and a plurality of second fully-connected networks that are different from each other; extracting a plurality of left channel audio signals and a plurality of right channel audio signals from the orientation audio feature based on a plurality of the audio output channels, receiving the first directional sound source signal features as input, and passing the first directional sound source signal features through a plurality of first fully connected networks for outputting left channel audio signals or right channel audio signals, and outputting corresponding left channel audio signals and right channel audio signals; receiving the second directional sound source signal features as input, and passing the second directional sound source signal features through a plurality of second fully connected networks for outputting left channel audio signals or right channel audio signals, and outputting corresponding left channel audio signals and right channel audio signals; 3. The method of claim 2, comprising:

4. the first audio signals of the plurality of target channels include a first front left channel signal for a front left channel, a first front right channel signal for a front right channel, a first rear left channel signal for a rear left channel, a first rear right channel signal for a rear right channel, and a first center channel signal for a center channel; Combining the left channel audio signal and the right channel audio signal corresponding to the same target channel to obtain a first audio signal of a plurality of target channels, Combining the left channel audio signal and the right channel audio signal output from a first fully connected network corresponding to the front left channel to obtain the first front left channel signal; Combining the left channel audio signal and the right channel audio signal output from a first fully connected network corresponding to the front right channel to obtain the first front right channel signal; Combining the left channel audio signal and the right channel audio signal output from a first fully connected network corresponding to the center channel to obtain the first center channel signal; Combining the left channel audio signal and the right channel audio signal output from a second fully connected network corresponding to the rear left channel to obtain the first rear left channel signal; and combining the left channel audio signal and the right channel audio signal output from a second fully connected network corresponding to the rear right channel to obtain the first rear right channel signal.

5. obtaining the stereo audio signal Obtaining an original stereo signal, and performing voice separation on the original stereo signal to obtain a non-voice signal and a voice signal; the non-human voice signal is the stereo audio signal; The method of claim 1 , comprising:

6. outputting an audio up-mix signal in a target format based on the first audio signals of the plurality of target channels, combining the human voice signal with a front left channel signal, a front right channel signal, and a center channel signal among the first audio signals for the plurality of target channels to obtain second audio signals for the plurality of target channels; outputting an audio upmix signal in a target format based on the second audio signals of the plurality of target channels; 6. The method of claim 5, comprising:

7. the first audio signals of the plurality of target channels include a first front left channel signal, a first front right channel signal, a first rear left channel signal, a first rear right channel signal, and a first center channel signal; the voice signal includes a left channel voice signal and a right channel voice signal; the second audio signals of the plurality of target channels include a second front left channel signal, a second front right channel signal, a second rear left channel signal, a second rear right channel signal, and a second center channel signal; Combining the human voice signal with the front channel signal and the center channel signal of the first audio signals for the plurality of target channels to obtain second audio signals for the plurality of target channels includes: weighting and combining the first front left channel signal and the left channel voice signal to obtain the second front left channel signal; weighting and combining the first front right channel signal and the right channel voice signal to obtain the second front right channel signal; weighting the first rear left channel signal to form the second rear left channel signal and weighting the first rear right channel signal to form the second rear right channel signal; weighting and combining the left channel voice signal, the right channel voice signal, and the first center channel signal to obtain the second center channel signal; 7. The method of claim 6, comprising:

8. This is performed by the audio upmix model, and obtaining each 5.1 channel sound source signal, selecting a target sound source signal from each of the 5.1 channel sound source signals, and extracting a 5 channel target audio signal from the target sound source signal; downmixing the five-channel target audio signal to obtain a stereo training audio signal; extracting five left channel audio signals and five right channel audio signals from the stereo training audio signal using the audio up-mix model, and combining the five left channel audio signals and the five right channel audio signals into a five-channel output audio signal; optimizing the audio up-mix model based on a difference between the five-channel target audio signal and the five-channel output audio signal; 2. The method of claim 1, comprising:

9. optimizing the audio up-mix model based on a difference between the five-channel target audio signal and the five-channel output audio signal, generating a first model loss based on an overall signal difference between the five-channel target audio signal and the five-channel output audio signal; generating a second model loss based on a first volume difference between different channel audio signals of the five-channel target audio signal and a second volume difference between different channel audio signals of the five-channel output audio signal; optimizing the audio upmix model based on the first model loss and the second model loss; 9. The method of claim 8, comprising:

10. 10. An audio device comprising a memory in which a computer program is stored and a processor, the audio device implementing the steps of the method according to any one of claims 1 to 9 when the processor executes the computer program.

Citation Information

Patent Citations

  • Signal processing device, and signal processing method

    WO2023162508A1

  • Generation of multichannel audio signal

    WO2024056438A1