Audio processing method, apparatus and audio output system

By extracting the location information of coherent sound and ambient sound from two-channel audio and processing the sky channel in combination with psychoacoustic characteristics, the shortcomings of traditional stereo playback methods in terms of sound source separation and spatial structure are solved, achieving clearer sound source separation and a three-dimensional immersive experience.

CN120751331BActive Publication Date: 2025-12-26ZHEJIANG LEAPMOTOR TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511228865.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-26
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Traditional stereo playback methods are insufficient to meet users' needs for an immersive music experience. Related upmixing algorithms are inadequate in terms of sound field separation and sound source separation. In particular, when processing coherent sound sources, they lack in-depth exploration of the details of energy and phase distribution in the left and right channels, resulting in poor sound source separation and missing spatial structure.

Method used

By extracting coherent sound, the location information of coherent sound, and ambient sound from the first frequency domain signal of dual-channel audio, separating them using the least squares method, extracting sub-coherent sound from multiple target channels based on the location information, and performing mid-to-high frequency processing on the sky channel in combination with psychoacoustic characteristics, the sound source layout in 3D space is reconstructed to generate a multi-channel time-domain audio signal.

Benefits of technology

It significantly improves sound source separation and spatial hierarchy, enhances three-dimensional immersion, solves the shortcomings of traditional stereo playback methods in sound field construction, and achieves clearer sound source separation and better vertical separation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751331B_ABST
    Figure CN120751331B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio processing method and device and an audio output system, and relate to the technical field of audio processing. The method comprises: extracting a first coherent sound, azimuth information of the first coherent sound and ambient sound of a double-channel audio from a first frequency domain signal of the double-channel audio; extracting, based on the azimuth information, sub-coherent sounds corresponding to at least part of sound channels of N target sound channels from the first coherent sound; wherein N is an integer greater than or equal to 3; determining second frequency domain signals of the N target sound channels based on the sub-coherent sounds and the ambient sound of the double-channel audio; and generating time domain audio signals corresponding to the N target sound channels based on the second frequency domain signals. Thus, in the process of expanding the double-channel audio into multi-channel audio, the sound source layout in the three-dimensional space can be effectively reconstructed, and the sound source separation degree can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of audio processing, and in particular to an audio processing method, an audio processing device and an audio output system. BACKGROUND

[0002] At present, most music resources still exist in the format of stereo (two-channel), such as Bluetooth music, mainstream sound sources of APP and contents of radio stations. In order to enhance the auditory experience, an audio output system (such as a vehicle audio system) usually carries multiple loudspeakers. For example, a vehicle audio system can carry multiple loudspeakers, which can be distributed in the instrument panel, front and rear doors, seat headrests, roof and tail area of a vehicle.

[0003] With the popularity of multi-channel and multi-loudspeaker layout in the audio output system, the traditional stereo playback mode has been difficult to meet the immersive music experience needs of users. Although the related upmix algorithm can expand the two-channel audio into multi-channel audio playback, there are still deficiencies in the separation degree of sound field construction and the like. SUMMARY

[0004] Embodiments of the present application provide an audio processing method, an audio processing device and an audio output system, which can effectively reconstruct the sound source layout in a three-dimensional space and significantly improve the sound source separation degree.

[0005] In a first aspect, an embodiment of the present application provides an audio processing method, comprising:

[0006] extracting a first coherent sound, azimuth information of the first coherent sound and ambient sound of two channels from a first frequency domain signal of the two-channel audio;

[0007] extracting, based on the azimuth information, sub-coherent sounds corresponding to at least part of sound channels in N target sound channels from the first coherent sound; wherein N is an integer greater than or equal to 3;

[0008] determining second frequency domain signals of the N target sound channels based on the sub-coherent sounds and the ambient sound of the two channels;

[0009] generating time domain audio signals corresponding to the N target sound channels based on the second frequency domain signals.

[0010] In an implementation manner, the extracting, based on the azimuth information, of the sub-coherent sounds corresponding to at least part of sound channels in the N target sound channels from the first coherent sound comprises:

[0011] for each target sound channel in the at least part of sound channels, determining a first gain value corresponding to the target sound channel in a case of extracting the sub-coherent sound corresponding to the target sound channel based on the azimuth information by using a first gain function, and determining the sub-coherent sound corresponding to the target sound channel based on the first coherent sound and the first gain value.

[0012] In an embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction;

[0013] The target sound channel corresponding sub-coherent sound is determined based on the first coherent sound and the first gain value, including:

[0014] The second gain value related to the frequency of the first coherent sound is determined by using a second gain function;

[0015] The first sound channel corresponding sub-coherent sound is determined based on the first coherent sound, the first gain value corresponding to the first sound channel, and the second gain value.

[0016] In an embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction;

[0017] The second frequency domain signal of the N target sound channels is determined based on each sub-coherent sound and the ambient sound of the dual sound channels, including:

[0018] The weight value corresponding to each sound channel in the dual sound channels is determined based on the layout position of the first sound channel;

[0019] The target ambient sound is generated by weighted sum of the ambient sound of the dual sound channels based on the weight value;

[0020] The second frequency domain signal of the first sound channel is generated by superimposing the target ambient sound and the first sound channel corresponding sub-coherent sound.

[0021] In an embodiment, the at least part of the sound channels includes a second sound channel, and the second sound channel is a sound channel in a second direction;

[0022] The second frequency domain signal of the N target sound channels is determined based on each sub-coherent sound and the ambient sound of the dual sound channels, including:

[0023] The second frequency domain signal of the second sound channel is generated by superimposing the ambient sound of the sound channel related to the second sound channel in the dual sound channels and the second sound channel corresponding sub-coherent sound.

[0024] In an embodiment, the at least part of the sound channels includes a third sound channel, and the third sound channel is a sound channel in a second direction;

[0025] The second frequency domain signal of the N target sound channels is determined based on each sub-coherent sound and the ambient sound of the dual sound channels, including:

[0026] The second frequency domain signal of the third sound channel is generated by superimposing the average value of the ambient sound of the dual sound channels and the third sound channel corresponding sub-coherent sound.

[0027] In an embodiment, the N target sound channels further include a fourth sound channel in addition to the at least part of the sound channels, and the fourth sound channel is a sound channel in a second direction;

[0028] determining second frequency domain signals of the N target sound channels based on each sub-coherent sound and the ambient sound of the two sound channels, comprising:

[0029] determining the ambient sound of the sound channel related to the fourth sound channel direction in the two sound channels as the second frequency domain signal of the fourth sound channel.

[0030] In an implementation, the first coherent sound, the direction information of the first coherent sound and the ambient sound of the two sound channels are extracted from the first frequency domain signal of the two sound channel audio, comprising:

[0031] The first coherent sound, the direction information and the ambient sound of the two sound channels are extracted from the first frequency domain signal by using the least square method.

[0032] In a second aspect, the embodiments of the present application provide an audio processing device, comprising:

[0033] The first extraction unit is configured to extract the first coherent sound, the direction information of the first coherent sound and the ambient sound of the two sound channels from the first frequency domain signal of the two sound channel audio;

[0034] The second extraction unit is configured to extract the sub-coherent sound corresponding to at least part of the sound channels in the N target sound channels from the first coherent sound based on the direction information; wherein N is an integer greater than or equal to 3;

[0035] The determination unit is configured to determine the second frequency domain signals of the N target sound channels based on each sub-coherent sound and the ambient sound of the two sound channels;

[0036] The conversion unit is configured to generate the time domain audio signals corresponding to the N target sound channels based on each second frequency domain signal.

[0037] In a third aspect, the embodiments of the present application provide an audio output system, comprising a plurality of loudspeakers corresponding to the N target sound channels and the audio processing device as described in the second aspect.

[0038] The scheme provided by the embodiments of the present application can extract the first coherent sound, the direction information of the first coherent sound and the ambient sound of the two sound channels from the first frequency domain signal of the two sound channel audio, then extract the sub-coherent sound corresponding to at least part of the sound channels in the N target sound channels from the first coherent sound based on the direction information, then determine the second frequency domain signals of the N target sound channels based on each sub-coherent sound and the ambient sound of the two sound channels, and then generate the time domain audio signals corresponding to the N target sound channels based on each second frequency domain signal. Therefore, in the process of expanding the two sound channel audio into the multi-channel audio, the scheme can effectively reconstruct the sound source layout in the 3D space through the sound source reconstruction (sub-coherent sound separation), the ambient sound and the sub-coherent sound recombination and other technologies, and significantly improve the sound source separation degree. In addition, the scheme also significantly improves the spatial level. BRIEF DESCRIPTION OF DRAWINGS

[0039] The technical solutions and other beneficial effects of the present application will become apparent from the following detailed description of specific embodiments of the present application, taken in conjunction with the accompanying drawings.

[0040] Figure 1 is a flow chart of an audio processing method in embodiments of the present application;

[0041] Figure 2 is a schematic diagram of an audio processing process in embodiments of the present application;

[0042] Figure 3 is a schematic diagram of determining a gain function of a sub-coherent sound based on orientation information in embodiments of the present application;

[0043] Figure 4 is a flow chart of a method of determining a sub-coherent sound corresponding to a sky channel in embodiments of the present application;

[0044] Figure 5 is a schematic diagram of vertical direction listening experience before and after mid-high frequency processing when a main channel and a sky channel exist simultaneously in embodiments of the present application;

[0045] Figure 6 is a flow chart of a method of determining a second frequency domain signal of a sky channel in embodiments of the present application;

[0046] Figure 7 is a schematic diagram of in-vehicle listening experience comparison effect;

[0047] Figure 8 is a structural schematic diagram of an audio processing apparatus in embodiments of the present application.

[0048] The reference signs: 800-audio processing apparatus, 801-first extraction unit, 802-second extraction unit, 803-determination unit, 804-conversion unit. DETAILED DESCRIPTION

[0049] The technical solutions and other beneficial effects of the present application will become apparent from the following detailed description of specific embodiments of the present application, taken in conjunction with the accompanying drawings.

[0050] In the description of the present application, it should be noted that, unless otherwise explicitly specified and limited, the term "and / or" in this paper is only to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / " in this paper generally represents an "or" relationship between the front and rear associated objects without special explanation.

[0051] As described above, with the popularity of multi-channel multi-speaker layout in the audio output system, the traditional stereo playback mode has been difficult to meet the immersive music experience needs of users. Although the related upmix algorithm can expand the two-channel audio into multi-channel audio playback, there are still deficiencies in the separation degree of sound field construction and the like. For example, the related upmix algorithm is insufficient in utilizing the coherent sound, and lacks the ability to remodel the structure of the sound source. Specifically, the related upmix algorithm often only extracts the "coherent sound source" (such as the main singer, various musical instruments, etc.) with high correlation in the stereo signal (such as the PAE method for extracting coherent sound), but does not further mine the details of the energy and phase distribution of the coherent sound in the left and right channels. Therefore, the information of the coherent sound is not fully utilized to remodel the layout of the multi-channel sound field. The related upmix algorithm does not reconstruct the actual positioning of the sound source in space according to the micro left-right distribution structure of the coherent sound, resulting in problems such as poor sound source separation, unstable sound image focusing, and missing spatial structure.

[0052] The embodiments of the present application provide an audio processing method, device and audio output system, which can effectively reconstruct the sound source layout in a three-dimensional space and significantly improve the sound source separation degree.

[0053] In some embodiments, the present application is applied to a car audio system, in which case the two-channel audio can be audio to be played by the car audio system.

[0054] In some embodiments, the present application is applied to an aircraft, a cinema, a conference room or a classroom, etc.

[0055] Figure 1 is a flowchart of an audio processing method in the embodiments of the present application. The method comprises the following steps:

[0056] S101: extracting a first coherent sound, azimuth information of the first coherent sound and an ambient sound of the two-channel audio from a first frequency domain signal of the two-channel audio;

[0057] S103: extracting sub-coherent sounds corresponding to at least part of the N target sound channels from the first coherent sound based on the azimuth information; wherein N is an integer greater than or equal to 3;

[0058] S105: determining second frequency domain signals of the N target sound channels based on the sub-coherent sounds and the ambient sound of the two-channel audio;

[0059] S107: generating time-domain audio signals corresponding to the N target sound channels based on the second frequency-domain signals.

[0060] Figure 1 The scheme provided by the corresponding embodiments can extract the first coherent sound, the orientation information of the first coherent sound, and the ambient sound of the dual sound channels from the first frequency-domain signal of the dual sound channels, then extract the sub-coherent sound corresponding to at least part of the N target sound channels from the first coherent sound based on the orientation information, then determine the second frequency-domain signals of the N target sound channels based on each sub-coherent sound and the ambient sound of the dual sound channels, and then generate time-domain audio signals corresponding to the N target sound channels based on each second frequency-domain signal. Thus, in the process of expanding the dual sound channels into multi-channel audio, the scheme can effectively reconstruct the sound source layout in the 3D space through sound source reconstruction (sub-coherent sound separation), ambient sound and sub-coherent sound recombination, and significantly improve the sound source separation degree. In addition, the scheme also significantly improves the spatial hierarchy.

[0061] Next, steps S101 to S107 are described.

[0062] In step S101, the first coherent sound, the orientation information of the first coherent sound, and the ambient sound of the dual sound channels are extracted from the first frequency-domain signal of the dual sound channels. Here, the dual sound channels usually refer to the left sound channel and the right sound channel. The orientation information can be the energy difference of the first coherent sound in the left and right sound channels.

[0063] As an example, the dual sound channels can be audio to be played by an audio output system, and the audio output system includes a plurality of speakers corresponding to the N target sound channels; wherein N is an integer greater than or equal to 3. The number of the plurality of speakers is greater than or equal to 3. In practice, the plurality of speakers are distributed in the environment (such as a vehicle, an aircraft, a cinema, a conference room, or a classroom, etc.) where the audio output system is located, and at least part of the plurality of speakers are distributed at different positions in the environment.

[0064] The N target channels can include multiple channels of the following: front channels, surround channels, sky channels, etc. The front channels, surround channels, and sky channels are divided based on a 3-dimensional sound field layout. The front channels and surround channels are horizontal-direction channels, and the sky channels are vertical-direction channels. The front channels are the "main sound field" in the 3-dimensional sound field, usually located in front of the user, and are responsible for presenting the main sound source (such as movie dialogue, music lead vocals, etc.). The front channels are usually subdivided into left front, center, and right front channels. The surround channels are located on the sides or rear of the user, and are used to simulate environmental sounds, background sound effects, and create an immersive feeling of being "surrounded by sound". The surround channels can be subdivided into left surround and right surround channels. Further, the surround channels can be subdivided into left surround, right surround, left rear surround, and right rear surround channels. The sky channels are the "vertical dimension" newly added to the 3-dimensional sound field, located above the user (such as the roof, etc.), and are used to simulate sounds coming from above. The sky channels can be subdivided into multiple of the following: left front sky channel, right front sky channel, left rear sky channel, and right rear sky channel, etc.

[0065] In practice, the left front channel, center channel, and right front channel are all main channels. Among them, the main channel usually refers to the channel that carries the main audio signal, and is the core part of the sound field construction. Since the division of channel types is a known technology in the art, it will not be described in more detail here.

[0066] The two-channel audio is a time-domain signal. The first frequency-domain signal is a frequency-domain representation of the two-channel audio, obtained by performing a Short-Time Fourier Transform (STFT) on the two-channel audio. Among them, the STFT is a mathematical tool for analyzing the time-frequency characteristics of a signal, which combines the "local" of the time-domain signal with the frequency-domain analysis, solving the problem that the traditional Fourier transform cannot process non-stationary signals (signals whose frequency changes over time).

[0067] The two-channel audio includes left-channel audio and right-channel audio. Let x represent the two-channel audio, L represent the left channel, and R represent the right channel, represents the audio signal at time t in the left-channel audio, represents the audio signal at time t in the right-channel audio, for example, as shown in Figure 2 , x can be processed into the first frequency-domain signal X by the STFT algorithm: , , It should be understood that is the frequency-domain representation of , is the frequency-domain representation of . Among them, Figure 2 ​Fig. 1 is a schematic diagram of an audio processing process in the embodiments of the present application. In the time domain, t represents time. In the frequency domain, t represents time frame index, and k represents frequency point index.

[0068] Then, the first coherent sound, the direction information of the first coherent sound and the ambient sound of the binaural sound can be extracted from the first frequency domain signal by using a PAE (Primary Ambient Extraction) algorithm. In practice, the PAE algorithm can separate two important components in the sound scene, i.e., the coherent sound component and the ambient sound component, and processing them respectively can improve the auditory experience when reconstructing the sound scene. Among them, the LS (Least-Squares) is a commonly used algorithm for the PAE algorithm. The LS algorithm estimates the input signal and then extracts the coherent sound and ambient sound components, and is the optimal estimation when the estimation error is irrelevant to the input signal, and the estimation weight is obtained by combining the model assumption to complete the extraction.

[0069] Next, the principle of the PAE-LS algorithm is introduced.

[0070] In the time domain, the audio signal in the binaural audio x can be split into coherent sound and ambient sound, which can be expressed as: ,

[0071] ;

[0072] ;

[0073] Among them, , represent the coherent sound and ambient sound in , , represent the coherent sound and ambient sound in . It should be pointed out that the coherent sound of the left and right channels (i.e., , ) is actually the same signal, only the energy size is different, so it can be written as:

[0074] ;

[0075] Among them, represents the size ratio relationship between and .

[0076] The audio signal in the time domain , can be expressed as:

[0077] ;

[0078] ​ ;

[0079] wherein, represents the ambient sound in the time domain, specifically the time domain representation of the above . represents the ambient sound in the time domain, specifically the time domain representation of the above . is and the coherent sound in the time domain, specifically the time domain representation of the above , . represents the direction information of the coherent sound.

[0080] For a determined time frame index t and frequency point index k, the ambient sound (such as , ) in the left and right channels in the above formula has the same short-time energy, which can be denoted as , and the short-time energy of the coherent sound (such as ) can be denoted as . It can be obtained that:

[0081] ;

[0082] ;

[0083] wherein, can represent the short-time energy of the left channel frequency domain signal (such as ), can represent the short-time energy of the right channel frequency domain signal (such as ), and B is the direction information of the coherent sound. The subscript S of

[0084] Taking the estimation of the coherent sound as an example, the calculation formula of the estimation value of the coherent sound can be represented as:

[0085] ;

[0086] wherein, and represent the estimation weights to be solved. The estimation error of the coherent sound can be represented as:

[0087] ;

[0088] wherein the subscript represents the coherent sound. In the LS algorithm, when the estimation error is completely irrelevant to the frequency domain signal , , the obtained weight is the optimal estimation, that is:

[0089] ;

[0090] .

[0091] Through the calculation process as described above, the first coherent sound , the direction information B of the first coherent sound and the binaural ambient sound , of the binaural audio x can be extracted from the first frequency domain signal X of the binaural audio x. Specifically, as shown in FIG. 1, the direction information B of the coherent sound Figure 2 , , the ambient sound of the left channel and the ambient sound of the right channel can be extracted from X: , , .

[0092] It should be noted that all the subscripts L in the foregoing represent the left channel, and all the subscripts R represent the right channel.

[0093] Through the LS algorithm, the , contains the position information of the coherent sound. According to the diagram of Figure 3 , it can be learned that for a determined time frame index t and frequency point index k, |B| = 1 represents that the sizes of the left and right channel coherent sounds are consistent, at this time, the sound source direction is in the middle; |B| < 1 represents that the left channel coherent sound is larger, at this time, the sound source direction is on the left, and the smaller |B| is, the more left the sound source is. |B| > 1 represents that the right channel coherent sound is larger, at this time, the sound source direction is on the right, and the larger |B| is, the more right the sound source is. Wherein, Figure 3 is a schematic diagram of a gain function for determining a sub-coherent sound based on the direction information in the embodiments of the present application. The gain function can be referred to as a gain function.

[0094] In Figure 3 , mode represents mode, the subscripts FL, FR and C represent the left front channel, the right front channel and the center channel respectively, represents the sub-coherent sound corresponding to the left front channel, represents the sub-coherent sound corresponding to the center channel, represents the sub-coherent sound corresponding to the right front channel. The curve pointed by the label 301 represents , may represent the gain value corresponding to the left front channel in the case of extracting the sub-coherent sound corresponding to the left front channel based on the orientation information B. The curve pointed by the label 302 represents , , may represent the gain value corresponding to the middle channel in the case of extracting the sub-coherent sound corresponding to the middle channel based on the orientation information B. The curve pointed by the label 303 represents , , may represent the gain value corresponding to the right front channel in the case of extracting the sub-coherent sound corresponding to the right front channel based on the orientation information B. The curve pointed by the label 304 represents . It can be seen from . It can be seen from Figure 3

[0095] ;

[0096] wherein the sum of the formulae is equal to 1. Therefore, the overall spatial energy distribution is balanced and high in restoration.

[0097] Based on the information in and , the embodiments of the present application design a set of sub-coherent sound extraction algorithms, such as the 3D upmixing algorithm shown in Figure 2 .

[0098] In step S103, the 3D upmixing algorithm is used to extract the sub-coherent sound corresponding to at least part of the N target channels from the first coherent sound S based on the orientation information B. The at least part of the channels can include at least one of the following: the first channel, the second channel, and the third channel. The first channel is a channel in the first direction, and the second channel and the third channel are channels in the second direction. The first direction is the vertical direction, and the second direction is the horizontal direction. Further, the first channel is the sky channel, the second channel includes the left front channel and / or the right front channel, and the third channel is the middle channel. Figure 2 The extraction of the sub-coherent sound corresponding to the middle channel , the sub-coherent sound corresponding to the left front channel , the sub-coherent sound corresponding to the right front channel , and the sub-coherent sound corresponding to the sky channel is schematically shown based on . Among them, the subscript S in represents the sky channel. It should be noted that in the present application, in addition to the foregoing and ​The subscript S other than the subscript S of the sky channel represents a sky channel; in addition, all the subscripts FL in the present application represent a left front channel, all the subscripts FR represent a right front channel, and all the subscripts C represent a center channel.

[0099] Specifically, for each target channel in the at least part of the channels, a first gain value corresponding to the target channel can be determined by using a first gain function, in a case that a sub-coherent sound corresponding to the target channel is extracted based on the bearing information B, and the sub-coherent sound corresponding to the target channel is determined based on the first coherent sound S and the first gain value.

[0100] Further, in an embodiment, in a case that the at least part of the channels includes a second channel, and the second channel includes a left front channel and a right front channel, for a determined time frame index t and a frequency point index k, the first gain value corresponding to the left front channel can be represented by , and the first gain value corresponding to the right front channel can be represented by , the sub-coherent sound corresponding to the left front channel , and the sub-coherent sound corresponding to the right front channel The algorithmic formulae can be:

[0101]

[0102]

[0103] In an embodiment, in a case that the at least part of the channels includes a third channel, and the third channel is a center channel, for a determined time frame index t and a frequency point index k, the first gain value corresponding to the center channel can be represented by , and the sub-coherent sound corresponding to the center channel The algorithmic formulae can be:

[0104]

[0105] In an embodiment, the at least part of the channels includes a first channel, and the first channel includes a sky channel. The sky channel is a channel in a vertical direction.

[0106] It should be noted that the related upmix algorithm lacks processing of the vertical direction of the sound field, and the spatial experience is flat and not "3D". For example, the related upmix algorithm mainly focuses on stretching and expanding the sound image in the horizontal direction, and lacks thinking and utilization of the uniqueness of the sky channel. Even, some algorithms directly copy the signals or ambient sounds of the main channels to the upper sky speakers, lacking spatial structure design of the sound field. The spatial restoration still stays in the horizontal dimension, lacking the upper and lower three-dimensional spatial levels, and the listening experience is flat and the spatial texture is insufficient.

[0107] ​​​The scheme provided by the embodiments of the present application can introduce psychoacoustic characteristics of pitch perception, intelligently map part of high-frequency high-pitch parts to the sky channel loudspeaker, simulate the auditory effect of "high pitch and high position", and thus enhance the level and immersion of the vertical space.

[0108] Under the influence of psychoacoustics, when playing sounds of different frequencies, listeners generally believe that high-frequency sounds come from a higher position and low-frequency sounds come from a lower position. Taking a vehicle as an example, the in-vehicle loudspeaker layout with a sky channel is roughly divided into two layers (the sky channel and other channels), and this physical layout provides the feasibility of improving the quality of the vertical directional sound field, thereby truly realizing 3D upmixing.

[0109] Based on the above analysis, the scheme provided by the embodiments of the present application performs high-frequency processing on the sky channel based on psychoacoustics to improve the vertical listening experience. Specifically, the flow shown in Figure 4 may be executed to determine the sub-coherent sound corresponding to the sky channel. In this regard, Figure 4 is a flowchart of a method for determining the sub-coherent sound corresponding to the sky channel in the embodiments of the present application. As shown in Figure 4 , the method includes the following steps:

[0110] S401: Using a first gain function, determine the first gain value of the sky channel corresponding to the sub-coherent sound extracted based on the directional information;

[0111] S403: Using a second gain function, determine the second gain value related to the frequency of the first coherent sound;

[0112] S405: Based on the first coherent sound, the first gain value of the sky channel corresponding to the sky channel, and the second gain value, determine the sub-coherent sound corresponding to the sky channel.

[0113] In this regard, the first gain value of the sky channel corresponding to the sky channel can be represented by , the second gain value related to the frequency of the first coherent sound S can be represented by , and the sub-coherent sound corresponding to the sky channel may be represented by the algorithm formula:

[0114] ;

[0115] In this regard, by using , the sky channel processing based on psychoacoustics is realized. It should be understood that Different from other modes, for the sky channel, the gain value curve of the frequency point indicated by the frequency point index k will reduce the spatial attenuation compared with other modes when the frequency of the frequency point is higher (finally the gain value is constant when the frequency is larger), so as to retain as much high-frequency sound signal as possible; since the high frequency does not play a dominant role in sound source positioning, this method will not interfere with the improved separation and spatial sense of the overall scheme. In addition, The additional gain given to different frequency bands can be specifically represented. According to psychoacoustics, the scheme provided in the embodiments of the present application gives a certain enhancement to the signal components with higher frequencies.

[0116] It should be noted that, The initial sub-coherent sound corresponding to the sky channel can be represented, and by applying , the initial sub-coherent sound corresponding to the sky channel can be reprocessed in the middle and high frequencies based on psychoacoustics.

[0117] In general, as Figure 5 shows, the direct vertical position difference of high and low frequencies is increased, so that the listener can feel better vertical separation and atmosphere. Among them, Figure 5 is a vertical direction listening feeling schematic diagram before and after the middle and high frequency processing when the main channel and the sky channel exist at the same time in the embodiments of the present application.

[0118] Continuing to refer to Figure 1 , in step S105, based on each sub-coherent sound and the ambient sound of the binaural channel, the second frequency domain signal of the N target channels is determined. Wherein, the binaural channel is the left channel and the right channel, Figure 2 schematically shows that based on the sub-coherent sound corresponding to the center channel , the sub-coherent sound corresponding to the left front channel , the sub-coherent sound corresponding to the right front channel , the sub-coherent sound corresponding to the sky channel , the ambient sound of the left channel and the ambient sound of the right channel , the ambient sound coherent sound is recombined to generate the second frequency domain signal corresponding to the center channel , the second frequency domain signal corresponding to the left front channel , the second frequency domain signal corresponding to the right front channel , and the second frequency domain signal corresponding to the sky channel .

[0119] Specifically, in one embodiment, the at least part of the channels includes a first channel. Taking the first channel as the sky channel as an example, the second frequency domain signal of the sky channel can be determined by executing the flow as shown in Figure 6 . Wherein, Figure 6This is a flowchart illustrating the method for determining the second frequency domain signal of the sky channel in an embodiment of this application. For example... Figure 6 As shown, the determination method includes the following steps:

[0120] S601: Determine the weight value of each channel in the two channels based on the layout position of the sky channel;

[0121] S603: Generates target ambient sound by weighted summation of the two-channel ambient sound based on the weight values;

[0122] S605: Superimpose the sub-coherent sound corresponding to the sky channel and the target ambient sound to generate the second frequency domain signal of the sky channel.

[0123] Specifically, the target ambient sound can be represented as:

[0124] ;

[0125] in, , Ambient sound for the left channel The weight value, Ambient sound for the right channel The weight values. Additionally... It is also used to reflect the left-right layout of the overhead sound channels; the further to the left the layout, the better. The closer to 1, the more rightward the layout. The closer it is to 0.

[0126] The second frequency domain signal of the sky channel The algorithm formula can be:

[0127] .

[0128] By adopting Figure 6 The described method determines the second frequency domain signal corresponding to the sky channel, and can combine the layout position of the sky channel to superimpose the ambient sound and the subcoherent sound after mid-to-high frequency processing in a certain proportion to enhance the sense of spatial atmosphere.

[0129] In one embodiment, at least some of the aforementioned channels include a second channel. The sub-coherent sound corresponding to the second channel and the ambient sound of the channels in the stereo setup that are directionally related to the second channel can be superimposed to generate a second frequency domain signal for the second channel. Taking a second channel comprising a left front channel and a right front channel as an example, the second frequency domain signal of the left front channel... and the second frequency domain signal of the right front channel The algorithm formula can be:

[0130] ;

[0131] .

[0132] In an embodiment, the at least part of the sound channels includes a third sound channel, and the average of the ambient sound of the two sound channels and the sub-coherent sound corresponding to the third sound channel can be superimposed to generate a second frequency domain signal of the third sound channel. Taking the third sound channel as a center sound channel, the algorithm formula of the second frequency domain signal corresponding to the center sound channel can be:

[0133] .

[0134] In an embodiment, the N target sound channels further include a fourth sound channel in addition to the at least part of the sound channels, and the fourth sound channel is a sound channel in a second direction. The ambient sound of the sound channel in the two sound channels related to the direction of the fourth sound channel can be determined as the second frequency domain signal of the fourth sound channel. Taking the fourth sound channel as a surround sound channel, which includes a left surround sound channel and a right surround sound channel, as an example, the ambient sound of the left sound channel can be determined as the second frequency domain signal of the left surround sound channel, and the ambient sound of the right sound channel can be determined as the second frequency domain signal of the right surround sound channel.

[0135] It should be noted that the related up-mixing algorithm does not make a suitable combination and distribution strategy for coherent and incoherent signals, and often simply distributes the extracted ambient sound to some secondary sound channels (such as surround sound channels, sky sound channels, etc.). According to the explanation of step S105 in the foregoing, the scheme provided in the embodiments of the present application can effectively recombine and distribute the ambient sound and the sub-coherent sound in combination with the spatial layout, so as to improve the spatial surround feeling and the atmosphere.

[0136] Continuing to refer to Figure 1 , in step S107, based on each second frequency domain signal, a time domain audio signal corresponding to the N target sound channels is generated. Specifically, each second frequency domain signal is subjected to ISTFT (Inverse Short-Time Fourier Transform, inverse short-time Fourier transform) respectively, to generate a time domain audio signal corresponding to the N target sound channels. Among them, Figure 2 schematically shows that the second frequency domain signal , , , is subjected to ISTFT respectively, to generate a time domain audio signal corresponding to the center sound channel, a time domain audio signal corresponding to the left front sound channel, a time domain audio signal corresponding to the right front sound channel, and a time domain audio signal corresponding to the sky sound channel.

[0137] When Figure 1The described audio processing method is applied to the case of a vehicle audio system. After generating time-domain audio signals corresponding to the N target sound channels, the vehicle audio system can play the time-domain audio signals corresponding to the target sound channels using the speakers mounted. The real vehicle effect comparison display is shown in Figure 7 . Wherein, Figure 7 is a schematic diagram of the in-vehicle listening experience comparison effect.

[0138] Figure 8 is a structural schematic diagram of an audio processing device in an embodiment of the present application. As shown in Figure 8 , the audio processing device 800 includes:

[0139] The first extraction unit 801 is configured to extract a first coherent sound, azimuth information of the first coherent sound, and an ambient sound of the binaural audio from a first frequency domain signal of the binaural audio;

[0140] The second extraction unit 802 is configured to extract sub-coherent sounds corresponding to at least part of the N target sound channels from the first coherent sound based on the azimuth information; wherein N is an integer greater than or equal to 3;

[0141] The determination unit 803 is configured to determine second frequency domain signals of the N target sound channels based on each sub-coherent sound and the ambient sound of the binaural audio;

[0142] The conversion unit 804 is configured to generate time-domain audio signals corresponding to the N target sound channels based on each second frequency domain signal.

[0143] In an embodiment, the second extraction unit 802 is configured to extract sub-coherent sounds corresponding to at least part of the N target sound channels from the first coherent sound based on the azimuth information, comprising:

[0144] The second extraction unit 802 is configured to, for each target sound channel in the at least part of the sound channels, determine a first gain value corresponding to the target sound channel in the case of extracting the sub-coherent sound corresponding to the target sound channel based on the azimuth information using a first gain function, and determine the sub-coherent sound corresponding to the target sound channel based on the first coherent sound and the first gain value.

[0145] In an embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction;

[0146] The second extraction unit 802 is configured to determine the sub-coherent sound corresponding to the target sound channel based on the first coherent sound and the first gain value, comprising:

[0147] determining a second gain value related to the frequency of the first coherent sound using a second gain function;

[0148] The sub-coherent sound corresponding to the first sound channel is determined based on the first coherent sound, the first gain value corresponding to the first sound channel, and the second gain value.

[0149] In an embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction.

[0150] The determination unit 803 is configured to determine the second frequency domain signals of the N target sound channels based on the sub-coherent sounds and the ambient sound of the binaural sound, including:

[0151] The weight value corresponding to each sound channel in the binaural sound is determined based on the layout position of the first sound channel.

[0152] The ambient sound of the binaural sound is weighted and summed based on the weight value to generate a target ambient sound.

[0153] The second frequency domain signal of the first sound channel is generated by superimposing the sub-coherent sound corresponding to the first sound channel and the target ambient sound.

[0154] In an embodiment, the at least part of the sound channels includes a second sound channel, and the second sound channel is a sound channel in a second direction.

[0155] The determination unit 803 is configured to determine the second frequency domain signals of the N target sound channels based on the sub-coherent sounds and the ambient sound of the binaural sound, including:

[0156] The determination unit 803 is configured to superimpose the sub-coherent sound corresponding to the second sound channel and the ambient sound of the sound channel in the binaural sound related to the direction of the second sound channel to generate the second frequency domain signal of the second sound channel.

[0157] In an embodiment, the at least part of the sound channels includes a third sound channel, and the third sound channel is a sound channel in a second direction.

[0158] The determination unit 803 is configured to determine the second frequency domain signals of the N target sound channels based on the sub-coherent sounds and the ambient sound of the binaural sound, including:

[0159] The determination unit 803 is configured to superimpose the average value of the ambient sound of the binaural sound and the sub-coherent sound corresponding to the third sound channel to generate the second frequency domain signal of the third sound channel.

[0160] In an embodiment, the N target sound channels further include a fourth sound channel in addition to the at least part of the sound channels, and the fourth sound channel is a sound channel in a second direction.

[0161] The determination unit 803 is configured to determine the second frequency domain signals of the N target sound channels based on the sub-coherent sounds and the ambient sound of the binaural sound, including:

[0162] The determining unit 803 is configured to determine the ambient sound of the channel in the two channels that is oriented to the fourth channel as the second frequency domain signal of the fourth channel.

[0163] In one embodiment, the first extraction unit 801 is configured to extract a first coherent sound, the location information of the first coherent sound, and the ambient sound of the two channels from a first frequency domain signal of the two-channel audio, including:

[0164] The first extraction unit 801 is configured to extract first coherent sound, directional information and two-channel ambient sound from the first frequency domain signal using the least squares method.

[0165] It should be noted that other aspects and implementation details of the audio processing device 800 are the same as or similar to the audio processing method described above, and will not be repeated here.

[0166] This application embodiment also provides an audio output system, which includes multiple speakers corresponding to N target channels, and as follows: Figure 8 The audio processing device 800 shown is described. This audio output system can be located in a vehicle, aircraft, cinema, conference room, or classroom, etc., without specific limitations. When the audio output system is located in a vehicle, it can be referred to as a vehicle audio system.

[0167] This application embodiment also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements, for example, Figure 1 The described audio processing method.

[0168] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the following... Figure 1 The audio processing method described.

[0169] This application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implements the following: Figure 1 The described audio processing method.

[0170] Based on the foregoing description, the solution provided in this application has the following advantages:

[0171] Advantage 1: Enhanced 3D immersion

[0172] The existing mixing algorithms only consider the horizontal direction and lack spatial representation in the vertical dimension, resulting in a flat listening experience. However, the embodiments of this application introduce psychoacoustic features of pitch effect and spatial response design values ​​in the vertical direction, which can effectively expand the vertical listening experience and thus form a 3D sound effect.

[0173] Advantage two: clearer sound source separation

[0174] The related upmix algorithm is not perfect in processing coherent sound, lacks reconstruction of its spatial physical characteristics, and leads to blurred sound image and insufficient separation; the embodiment of the present application fully utilizes the internal spatial orientation information of coherent sound, obtains multiple sub-coherent sounds in different spatial orientations and performs spatial re-projection to adapt to the multi-channel and multi-speaker system in the vehicle, which can significantly increase the separation and level.

[0175] Advantage three: more reasonable sound source combination

[0176] The related upmix algorithm does not make appropriate combination and distribution strategy for coherent and incoherent signals, and often simply distributes the extracted environmental sound to some secondary channels (surround channels, sky channels, etc.); the embodiment of the present application combines spatial layout to effectively recombine and distribute environmental sound and coherent sound (sub-coherent sound), which can improve the spatial surround feeling and atmosphere.

[0177] The above is only part of the embodiments of the present application, and does not limit the present application in any form. The protection scope of the embodiments of the present application is not limited thereto, and any person skilled in the art can easily think of simple modifications, equivalent changes and modifications within the technical range disclosed in the embodiments of the present application, which should be covered within the protection scope of the embodiments of the present application.

Claims

1. An audio processing method, characterized by, The method comprises: extracting a first coherent sound, azimuth information of the first coherent sound, and ambient sound of a two-channel audio from a first frequency domain signal of the two-channel audio; for each of at least part of N target channels, determining a first gain value corresponding to the target channel based on extracting a sub-coherent sound corresponding to the target channel based on the azimuth information by using a first gain function, and determining the sub-coherent sound corresponding to the target channel based on the first coherent sound and the first gain value; wherein N is an integer greater than or equal to 3; the at least part of the channels includes a sky channel in a vertical direction, and a psychoacoustic characteristic of pitch perception is introduced to perform mid-high frequency processing on the sky channel when extracting the sub-coherent sound corresponding to the sky channel; determining a second frequency domain signal of the N target channels based on the spatial layout of the N target channels, the sub-coherent sound, and the ambient sound of the two-channel audio; generating time-domain audio signals corresponding to the N target channels based on the second frequency domain signal.

2. The audio processing method of claim 1, wherein, The method for determining the sub-coherent sound corresponding to the target channel based on the first coherent sound and the first gain value comprises: determining a second gain value related to the frequency of the first coherent sound by using a second gain function; determining the sub-coherent sound corresponding to the sky channel based on the first coherent sound, the first gain value corresponding to the sky channel, and the second gain value.

3. The audio processing method of claim 1, wherein, The method for determining a second frequency domain signal of the N target channels based on the spatial layout of the N target channels, the sub-coherent sound, and the ambient sound of the two-channel audio comprises: determining a weight value corresponding to each channel of the two-channel audio based on the layout position of the sky channel; performing weighted summation on the ambient sound of the two-channel audio based on the weight value to generate a target ambient sound; superimposing the sub-coherent sound corresponding to the sky channel and the target ambient sound to generate the second frequency domain signal of the sky channel.

4. The audio processing method of claim 1, wherein, The at least part of the channels includes a second channel, and the second channel is a channel in a second direction; The method for determining a second frequency domain signal of the N target channels based on the spatial layout of the N target channels, the sub-coherent sound, and the ambient sound of the two-channel audio comprises: superimposing the sub-coherent sound corresponding to the second channel and the ambient sound of the channel of the two-channel audio related to the direction of the second channel to generate the second frequency domain signal of the second channel.

5. The audio processing method of claim 1, wherein, The at least part of the channels includes a third channel, and the third channel is a channel in a second direction; The method for determining a second frequency domain signal of the N target channels based on the spatial layout of the N target channels, the sub-coherent sound, and the ambient sound of the two-channel audio comprises: superimposing an average value of the ambient sound of the two-channel audio and the sub-coherent sound corresponding to the third channel to generate the second frequency domain signal of the third channel.

6. The audio processing method of claim 1, wherein, The N target channels further include a fourth channel in addition to the at least part of the channels, and the fourth channel is a channel in a second direction; determining, based on the spatial layout of the N target sound channels, the sub-coherent sounds of each of the N target sound channels, and the ambient sound of the dual sound channels, second frequency domain signals of the N target sound channels, comprising: determining, as the second frequency domain signal of the fourth sound channel, the ambient sound of the sound channel of the dual sound channels related to the fourth sound channel direction.

7. The audio processing method of any of claims 1-6, wherein, the extracting of the first coherent sound, the direction information of the first coherent sound, and the ambient sound of the dual sound channels from the first frequency domain signal of the dual sound channel audio, comprising: extracting the first coherent sound, the direction information, and the ambient sound of the dual sound channels from the first frequency domain signal by using a least square method.

8. An audio processing apparatus, characterized by comprising: comprising: a first extraction unit configured to extract the first coherent sound, the direction information of the first coherent sound, and the ambient sound of the dual sound channels from the first frequency domain signal of the dual sound channel audio; a second extraction unit configured to, for each of at least part of the N target sound channels, determine, by using a first gain function, a first gain value of the target sound channel in a case where the sub-coherent sound corresponding to the target sound channel is extracted based on the direction information, and determine the sub-coherent sound corresponding to the target sound channel based on the first coherent sound and the first gain value; wherein N is an integer greater than or equal to 3; the at least part of the sound channels includes a sky sound channel in a vertical direction, and a psychoacoustic characteristic of pitch perception is introduced to perform mid-high frequency processing on the sky sound channel when the sub-coherent sound corresponding to the sky sound channel is extracted; a determination unit configured to determine, based on the spatial layout of the N target sound channels, the sub-coherent sounds of each of the N target sound channels, and the ambient sound of the dual sound channels, second frequency domain signals of the N target sound channels; a conversion unit configured to generate time domain audio signals corresponding to the N target sound channels based on the second frequency domain signals.

9. An audio output system, characterized by comprising a plurality of loudspeakers corresponding to N target sound channels, and the audio processing device according to claim 8.

Citation Information

Patent Citations

  • Audio signal separation method and device, electronic equipment and storage medium

    CN119785819A