Audio processing method and device and audio output system

By extracting the directional information of coherent sound and ambient sound from the two-channel audio, and using the 3D upmixing algorithm and speaker layout to reconstruct the three-dimensional sound source layout, the shortcomings of traditional stereo playback methods in sound field construction are solved, achieving clearer sound source separation and better three-dimensional immersion.

CN120751331AActive Publication Date: 2025-10-03ZHEJIANG LEAPMOTOR TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511228865.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-03
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Traditional stereo playback methods are unable to meet users' needs for an immersive music experience. The correlated upmixing algorithm has shortcomings in the separation and spatial performance of sound field construction. In particular, when processing coherent sound sources, it lacks in-depth exploration of the energy and phase distribution details in the left and right channels, resulting in poor sound source separation and lack of spatial structure.

Method used

By extracting coherent sound, its azimuth information, and ambient sound from the frequency domain signal of two-channel audio, and using a 3D upmixing algorithm to extract sub-coherent sounds of multiple target channels from the coherent sound, and combining the speaker layout position and psychoacoustic characteristics, the sound source layout in 3D space is reconstructed to generate a multi-channel audio signal.

Benefits of technology

It significantly improves the sound source separation and spatial layering, enhances the three-dimensional immersion, solves the shortcomings of traditional stereo playback methods in sound field construction, and achieves clearer sound source separation and a better listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751331A_ABST
    Figure CN120751331A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio processing method and device and an audio output system, and relates to the technical field of audio processing. The method comprises the following steps: extracting first coherent sound, azimuth information of the first coherent sound and environment sound of two channels from a first frequency domain signal of two-channel audio; based on the azimuth information, extracting sub coherent sounds corresponding to at least part of sound channels in the N target sound channels from the first coherent sound; wherein N is an integer greater than or equal to 3; determining second frequency domain signals of N target sound channels based on the sub coherent sound and the environment sound of the double sound channels; and generating time domain audio signals corresponding to the N target sound channels based on the second frequency domain signals. Therefore, in the process of expanding the dual-channel audio into the multi-channel audio, the sound source layout in the three-dimensional space can be effectively reconstructed, and the sound source separation degree is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of audio processing technology, and in particular to an audio processing method, device, and audio output system. Background Art

[0002] Currently, most music sources, such as Bluetooth music, mainstream app audio sources, and radio content, are still available in stereo (two-channel) format. Audio output systems (such as in-car audio systems) typically feature multiple speakers to enhance the listening experience. For example, in-car audio systems can be equipped with multiple speakers, which can be located in the vehicle's dashboard, front and rear doors, seat headrests, roof, and rear areas.

[0003] With the widespread adoption of multi-channel, multi-speaker layouts in audio output systems, traditional stereo playback methods are no longer able to meet users' demands for an immersive music experience. While correlated upmixing algorithms can expand two-channel audio into multi-channel playback, they still lack sufficient separation in sound field construction. Summary of the Invention

[0004] The embodiments of the present application provide an audio processing method, device, and audio output system that can effectively reconstruct the sound source layout in a three-dimensional space and significantly improve the sound source separation.

[0005] In a first aspect, an embodiment of the present application provides an audio processing method, comprising: Extracting a first coherent sound, position information of the first coherent sound, and a binaural ambient sound from a first frequency domain signal of the binaural audio; Extracting, based on the direction information, sub-coherent sounds corresponding to at least some of the N target sound channels from the first coherent sound; wherein N is an integer greater than or equal to 3; Determine second frequency domain signals of N target channels based on the sub-coherent sounds and the binaural ambient sound; Based on each second frequency domain signal, time domain audio signals corresponding to N target channels are generated.

[0006] In one embodiment, extracting sub-coherent sounds corresponding to at least some of the N target channels from the first coherent sound based on the direction information includes: For each target channel in the at least some of the channels, a first gain value corresponding to the target channel is determined using a first gain function when extracting a sub-coherent sound corresponding to the target channel based on the azimuth information. The sub-coherent sound corresponding to the target channel is then determined based on the first coherent sound and the first gain value.

[0007] In one embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction; Determining a sub-coherent sound corresponding to a target channel based on the first coherent sound and the first gain value includes: Determining a second gain value related to the frequency of the first coherent sound using a second gain function; A sub-coherent sound corresponding to the first channel is determined based on the first coherent sound, the first gain value corresponding to the first channel, and the second gain value.

[0008] In one embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction; Determining second frequency domain signals of N target channels based on the sub-coherent sounds and the two-channel ambient sound includes: Determine the weight value corresponding to each channel in the dual channels based on the layout position of the first channel; Perform weighted summation of the binaural ambient sounds based on the weight values ​​to generate the target ambient sound; The sub-coherent sound corresponding to the first sound channel and the target ambient sound are superimposed to generate a second frequency domain signal of the first sound channel.

[0009] In one embodiment, the at least part of the sound channels includes a second sound channel, and the second sound channel is a sound channel in a second direction; Determining second frequency domain signals of N target channels based on the sub-coherent sounds and the two-channel ambient sound includes: The sub-coherent sound corresponding to the second channel and the ambient sound of the channel in the dual channels that is azimuthally related to the second channel are superimposed to generate a second frequency domain signal of the second channel.

[0010] In one embodiment, the at least part of the sound channels includes a third sound channel, and the third sound channel is a sound channel in the second direction; Determining second frequency domain signals of N target channels based on the sub-coherent sounds and the two-channel ambient sound includes: The average value of the dual-channel ambient sound and the sub-coherent sound corresponding to the third channel are superimposed to generate a second frequency domain signal of the third channel.

[0011] In one embodiment, the N target sound channels further include a fourth sound channel in addition to at least some of the above-mentioned sound channels, and the fourth sound channel is a sound channel in the second direction; Determining second frequency domain signals of N target channels based on the sub-coherent sounds and the two-channel ambient sound includes: The ambient sound of the channel in the dual channels that is azimuthally related to the fourth channel is determined as the second frequency domain signal of the fourth channel.

[0012] In one embodiment, extracting a first coherent sound, position information of the first coherent sound, and a binaural ambient sound from a first frequency domain signal of a binaural audio includes: The first coherent sound, the direction information and the binaural ambient sound are extracted from the first frequency domain signal by using the least square method.

[0013] In a second aspect, an embodiment of the present application provides an audio processing device, comprising: A first extraction unit is configured to extract a first coherent sound, position information of the first coherent sound, and a binaural ambient sound from a first frequency domain signal of the binaural audio; A second extraction unit is configured to extract, from the first coherent sound, sub-coherent sounds corresponding to at least some of the N target sound channels based on the direction information; wherein N is an integer greater than or equal to 3; a determining unit configured to determine second frequency domain signals of N target channels based on the sub-coherent sounds and the binaural ambient sound; The conversion unit is configured to generate time domain audio signals corresponding to N target channels based on each second frequency domain signal.

[0014] In a third aspect, an embodiment of the present application provides an audio output system, comprising a plurality of speakers corresponding to N target channels, and the audio processing device as described in the second aspect.

[0015] The solution provided by the embodiment of the present application can extract the first coherent sound, the orientation information of the first coherent sound, and the binaural ambient sound from the first frequency domain signal of the binaural audio, and then extract the sub-coherent sounds corresponding to at least some of the N target channels from the first coherent sound based on the orientation information. Then, based on each sub-coherent sound and the binaural ambient sound, the second frequency domain signals of the N target channels are determined, and then based on each second frequency domain signal, the time domain audio signals corresponding to the N target channels are generated. Therefore, in the process of expanding binaural audio to multi-channel audio, this solution can effectively reconstruct the sound source layout in 3D space through technologies such as sound source reconstruction (sub-coherent sound separation) and recombining ambient sound and sub-coherent sound, thereby significantly improving the sound source separation. In addition, this solution also significantly improves the sense of spatial hierarchy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The following detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings will make the technical solutions and other beneficial effects of the present application apparent.

[0017] Figure 1 is a flow chart of the audio processing method in an embodiment of the present application; Figure 2 is a schematic diagram of the audio processing process in an embodiment of the present application; Figure 3 is a schematic diagram of determining a gain function of a sub-coherent sound based on azimuth information in an embodiment of the present application; Figure 4is a flowchart of a method for determining sub-coherent sounds corresponding to a sky channel in an embodiment of the present application; Figure 5 2. This is a schematic diagram of the vertical listening experience before and after mid-high frequency processing when the main channel and the sky channel exist simultaneously in an embodiment of the present application; Figure 6 is a flowchart of a method for determining a second frequency domain signal of a sky channel in an embodiment of the present application; Figure 7 This is a schematic diagram of the contrasting effects of in-car listening; Figure 8 It is a structural diagram of the audio processing device in an embodiment of the present application.

[0018] Reference numerals: 800 - audio processing device, 801 - first extraction unit, 802 - second extraction unit, 803 - determination unit, 804 - conversion unit. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0020] In the description of this application, it should be noted that, unless otherwise specified or limited, the term "and / or" herein is merely a description of an association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the character " / " herein, unless otherwise specified, generally indicates that the associated objects are in an "or" relationship.

[0021] As mentioned earlier, with the prevalence of multi-channel, multi-speaker configurations in audio output systems, traditional stereo playback methods are no longer able to meet users' demands for an immersive music experience. While correlation upmixing algorithms can expand two-channel audio into multi-channel playback, they still have shortcomings in areas such as sound field separation. For example, correlation upmixing algorithms underutilize coherent sound and lack the ability to reconstruct the sound source structure. Specifically, correlation upmixing algorithms often perform simple extraction (using methods such as PAE) of highly correlated "coherent sound sources" in stereo signals (such as lead vocals and various instruments), but fail to delve into the detailed energy and phase distribution of these coherent sound sources in the left and right channels. Consequently, they fail to fully utilize the coherent sound information to reshape the multi-channel sound field layout. Correlation upmixing algorithms fail to reconstruct the actual spatial positioning of sound sources based on the microscopic left-right distribution structure of the coherent sound, resulting in poor sound source separation, unstable sound image focus, and a loss of spatial structure.

[0022] The embodiments of the present application provide an audio processing method, device, and audio output system that can effectively reconstruct the sound source layout in a three-dimensional space and significantly improve the sound source separation.

[0023] In some embodiments, the present application is applied to a car audio system. In this case, the two-channel audio may be the audio to be played by the car audio system.

[0024] In some embodiments, the present application is applied to aircraft, theaters, conference rooms or classrooms, etc.

[0025] Figure 1 This is a flow chart of the audio processing method in the embodiment of the present application. The method includes the following steps: S101: extracting a first coherent sound, position information of the first coherent sound, and a binaural ambient sound from a first frequency domain signal of a binaural audio; S103: extracting sub-coherent sounds corresponding to at least some of the N target channels from the first coherent sound based on the direction information, where N is an integer greater than or equal to 3; S105: Determine second frequency domain signals of N target channels based on the sub-coherent sounds and the dual-channel ambient sound; S107: Generate time domain audio signals corresponding to N target channels based on the second frequency domain signals.

[0026] Figure 1The corresponding embodiment provides a solution that can extract the first coherent sound, the orientation information of the first coherent sound, and the binaural ambient sound from the first frequency domain signal of the binaural audio. Based on the orientation information, the solution then extracts the sub-coherent sounds corresponding to at least some of the N target channels from the first coherent sound. Subsequently, based on each sub-coherent sound and the binaural ambient sound, the second frequency domain signals of the N target channels are determined. Then, based on each second frequency domain signal, the time domain audio signals corresponding to the N target channels are generated. Thus, in the process of expanding binaural audio to multi-channel audio, this solution can effectively reconstruct the sound source layout in three-dimensional space through techniques such as sound source reconstruction (sub-coherent sound separation) and recombining ambient sound and sub-coherent sound, significantly improving sound source separation. Furthermore, this solution significantly enhances the sense of spatial layering.

[0027] Next, steps S101 to S107 will be described.

[0028] In step S101, a first coherent sound, positional information of the first coherent sound, and binaural ambient sound are extracted from a first frequency domain signal of a two-channel audio signal. Binary audio typically refers to a left channel and a right channel. The positional information may be the energy difference between the first coherent sound in the left and right channels.

[0029] As an example, two-channel audio may be audio to be played by an audio output system that includes multiple speakers corresponding to N target channels, where N is an integer greater than or equal to 3. The number of these multiple speakers is greater than or equal to 3. In practice, these multiple speakers are distributed throughout an environment (such as a vehicle, aircraft, theater, conference room, or classroom) in which the audio output system is located, and at least some of these multiple speakers are located at different locations within the environment.

[0030] The N target channels can include various channels, including front channels, surround channels, and overhead channels. These channels are based on the 3D sound field layout. Front and surround channels are horizontal channels, while overhead channels are vertical channels. Front channels are the "main sound field" in the 3D sound field, typically located in front of the user and responsible for presenting the primary sound source (such as movie dialogue or lead vocals). Front channels are typically divided into left front, center, and right front channels. Surround channels are located to the sides or behind the user and simulate ambient and background sounds, creating an immersive feeling of being "surrounded by sound." Surround channels can be further divided into left surround and right surround channels. Furthermore, surround channels can be further divided into left surround, right surround, left back surround, and right back surround channels. The overhead channel is a newly added "vertical dimension" to the 3D sound field, located above the user (such as on the roof of a car) and simulating sounds from above. The sky channel can be further divided into the following categories: left front sky channel, right front sky channel, left rear sky channel, and right rear sky channel.

[0031] In practice, the left front channel, center channel, and right front channel are all considered main channels. A main channel generally refers to the channel that carries the primary audio signal and is the core component of the sound field. Since the classification of channels is well known in the art, we will not elaborate on it here.

[0032] Two-channel audio is a time-domain signal. The first frequency-domain signal is the frequency-domain representation of the two-channel audio, obtained by performing a short-time Fourier transform (STFT) on the two-channel audio. The STFT is a mathematical tool used to analyze the time-frequency characteristics of a signal. Its core approach is to combine the "local" nature of the time-domain signal with frequency-domain analysis, addressing the problem that traditional Fourier transforms cannot handle non-stationary signals (signals whose frequency varies with time).

[0033] Dual-channel audio includes left-channel audio and right-channel audio. x represents dual-channel audio, L represents left channel, and R represents right channel. Represents the audio signal at time t in the left channel audio, Take the audio signal at time t in the right channel audio as an example, Figure 2 As shown, the STFT algorithm can be used to convert x: 、 Processed into the first frequency domain signal X: 、 It should be understood that for The frequency domain representation of for Frequency domain representation of . Among them, Figure 2 FIG1 is a schematic diagram of the audio processing process in an embodiment of the present application. In the time domain, t represents time. In the frequency domain, t represents the time frame index, and k represents the frequency point index.

[0034] Next, the PAE (Primary Ambient Extraction) algorithm can be used to extract the first coherent sound, the azimuth information of the first coherent sound, and the binaural ambient sound from the first frequency domain signal. In practice, the PAE algorithm can separate the two important components of the sound scene: the coherent sound component and the ambient sound component. Processing them separately can enhance the auditory experience when reconstructing the sound scene. Among them, the least squares (LS) method is a commonly used algorithm in the PAE algorithm. The LS algorithm estimates the input signal and then extracts the coherent sound and ambient sound components. When the estimation error is uncorrelated with the input signal, it is the optimal estimate. Combined with the model assumptions, the estimation weight is calculated to complete the extraction.

[0035] Next, the principle of the PAE-LS algorithm is introduced.

[0036] In the time domain, the audio signal in the two-channel audio x 、 It can be divided into coherent sound and ambient sound, as shown below: ; ; in, 、 In turn, The coherent sound and ambient sound in 、 In turn, It should be noted that the coherent sound of the left and right channels (i.e. 、 ) are actually the same signal, only the energy size is different, so it can be written as: ; in, express and The size ratio relationship.

[0037] Audio signal in the time domain 、 After STFT, it can be expressed as: ; ; in, express The ambient sound in the Frequency domain representation of . express The ambient sound in the Frequency domain representation of . for and The coherent sound in 、 Frequency domain representation of . express direction information.

[0038] For a certain time frame index t and frequency index k, the ambient sound of the left and right channels (such as 、 ) have the same short-term energy and can be expressed as , coherent sound (such as ) can be expressed as . We can get: ; ; in, Can represent the left channel frequency domain signal (like ) of short-term energy, Can represent the right channel frequency domain signal (like ) is the short-time energy of the coherent sound, and B is the directional information of the coherent sound. The subscript S indicates coherent sound.

[0039] To estimate coherent sound For example, coherent sound Estimated value of The calculation formula can be expressed as: ; in, and Represents the estimated weight to be sought. The estimated error It can be expressed as: ; in, Subscript In the LS algorithm, when the estimation error With frequency domain signal 、 When there is no correlation, the weight obtained is the optimal estimate, that is: ; .

[0040] Through the calculation process as described above, the first coherent sound can be extracted from the first frequency domain signal X of the two-channel audio x. , first coherent sound Position information B and dual-channel ambient sound 、 Specifically, Figure 2 As shown, you can get it from X: 、 Extracting coherent sound 、 Location information , Ambient sound of the left channel and ambient sound on the right channel .

[0041] It should be noted that all subscripts L in the above text represent the left channel, and all subscripts R represent the right channel.

[0042] Through the LS algorithm, we can get , It contains the location information of the coherent sound. Figure 3 From the diagram, we can see that for a certain time frame index t and frequency index k, |B|=1 means that the coherent sound of the left and right channels is the same, and the sound source is centered; |B|<1 means that the coherent sound of the left channel is louder, and the sound source is on the left, and the smaller |B| is, the closer the sound source is to the left. |B|>1 means that the coherent sound of the right channel is louder, and the sound source is on the right, and the larger |B| is, the closer the sound source is to the right. Figure 3 FIG. 1 is a schematic diagram of a gain function for determining sub-coherent sound based on azimuth information in an embodiment of the present application. The gain function may be referred to as a gain function.

[0043] exist Figure 3 In the format, mode indicates the mode, and the subscripts FL, FR, and C indicate the left front channel, right front channel, and center channel, respectively. Indicates the sub-coherent sound corresponding to the left front channel, Indicates the sub-coherent sound corresponding to the center channel, The curve pointed to by number 301 represents the sub-coherent sound corresponding to the right front channel. , It can be expressed that the sub-coherent sound corresponding to the left front channel is extracted based on the azimuth information B. The curve pointed to by reference numeral 302 represents the gain value of the left front channel. , It can be expressed that the sub-coherent sound corresponding to the center channel is extracted based on the azimuth information B. The curve pointed to by number 303 represents the gain value of the center channel. , It can be expressed that the sub-coherent sound corresponding to the right front channel is extracted based on the azimuth information B. In the case of , the gain value corresponding to the right front channel. The line pointed by the number 304 represents .from Figure 3 It can be seen that: ; The sum of these equations is always equal to 1. Therefore, the overall spatial energy distribution is balanced and highly reducible.

[0044] based on and Based on the information in the embodiment of the present application, a sub-coherent sound extraction algorithm is designed, such as Figure 2 The 3D upmixing algorithm shown in .

[0045] In step S103, a 3D upmixing algorithm is used to extract sub-coherent sounds corresponding to at least some of the N target channels from the first coherent sound S based on the azimuth information B. The at least some of the channels may include at least one of the following: a first channel, a second channel, and a third channel. The first channel is a channel in a first direction, and the second and third channels are channels in a second direction. The first direction is a vertical direction, and the second direction is a horizontal direction. Furthermore, the first channel is a sky channel, the second channel includes a left front channel and / or a right front channel, and the third channel is a center channel. Figure 2 Schematically shows the ,from Extract the sub-coherent sound corresponding to the center channel , the sub-coherent sound corresponding to the left front channel , the sub-coherent sound corresponding to the right front channel , and the sub-coherent sound corresponding to the sky channel .in, The subscript S in represents the sky channel. and All subscripts S except the subscript S in represent the sky channel; in addition, all subscripts FL in this application represent the left front channel, all subscripts FR represent the right front channel, and all subscripts C represent the center channel.

[0046] Specifically, for each target channel in the at least some of the channels mentioned above, a first gain function can be used to determine a first gain value corresponding to the target channel when extracting the sub-coherent sound corresponding to the target channel based on the orientation information B, and the sub-coherent sound corresponding to the target channel can be determined based on the first coherent sound S and the first gain value.

[0047] Furthermore, in one embodiment, when the at least part of the channels includes a second channel, and the second channel includes a left front channel and a right front channel, for a determined time frame index t and a frequency index k, the first gain value corresponding to the left front channel can be expressed as Indicates that the first gain value corresponding to the right front channel can be used Indicates that the sub-coherent sound corresponding to the left front channel Sub-coherent sound corresponding to the right front channel The algorithm formula can be: ; .

[0048] In one embodiment, when the at least part of the channels include a third channel, and the third channel is a center channel, for a determined time frame index t and frequency index k, the first gain value corresponding to the center channel can be obtained by Indicates that the sub-coherent sound corresponding to the center channel The algorithm formula can be: .

[0049] In one embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel includes a sky channel. The sky channel is a vertical sound channel.

[0050] It's important to note that some upmixing algorithms lack vertical processing for the sound field, resulting in a flat, un-3D spatial experience. For example, these algorithms primarily focus on horizontally stretching the sound image, failing to consider and utilize the unique characteristics of the overhead channel. Some algorithms even copy the main channel signal or ambient sound directly to the overhead speakers, lacking a clear understanding of the sound field's spatial structure. Spatial restoration remains horizontal, lacking the three-dimensional spatial hierarchy above and below, resulting in a flat, lacking spatial texture.

[0051] The solution provided in the embodiment of the present application can introduce the psychoacoustic characteristics of pitch perception, and insightfully map part of the high-frequency and treble parts to the sky channel speakers, simulating the auditory effect of "high sound and high position", thereby enhancing the sense of layering and immersion in the vertical space.

[0052] Due to the influence of psychoacoustics, when playing sounds of different frequencies, listeners generally perceive high-frequency sounds as coming from higher positions and low-frequency sounds as coming from lower positions. For example, in a vehicle with an overhead sound channel, the speaker layout is generally divided into two layers within the vehicle (the overhead sound channel and the other channels). This physical layout makes it possible to enhance the vertical sound field quality, thus achieving true 3D upmixing.

[0053] Based on the above analysis, the solution provided in the embodiment of the present application performs mid-high frequency processing on the sky channel based on psychoacoustics to improve the vertical listening experience. Figure 4 The process shown in the figure determines the sub-coherent sound corresponding to the sky channel. Figure 4 FIG. 1 is a flow chart of a method for determining sub-coherent sounds corresponding to the sky channel in an embodiment of the present application. Figure 4 As shown, the determination method includes the following steps: S401: Determine, using a first gain function, a first gain value corresponding to the sky channel when extracting sub-coherent sounds corresponding to the sky channel based on azimuth information; S403: Determine a second gain value related to the frequency of the first coherent sound using a second gain function; S405: Determine a sub-coherent sound corresponding to the sky channel based on the first coherent sound, the first gain value corresponding to the sky channel, and the second gain value.

[0054] Among them, the first gain value corresponding to the sky channel can be used Indicates that the second gain value related to the frequency of the first coherent sound S can be expressed as Indicates that the sub-coherent sound corresponding to the sky channel The algorithm formula can be: ; Among them, by using , realizes the sky channel processing based on psychoacoustics. It should be understood that Unlike other modes mentioned above, for the sky channel, when the frequency index k indicates a higher frequency, its gain curve will reduce spatial attenuation compared to other modes (ultimately, when the frequency is higher, the gain value is always 1), thereby retaining as much high-frequency sound signal as possible; since high frequencies do not play a dominant role in sound source localization, this method will not interfere with the separation and spatial perception improved by the overall solution. In addition, It can specifically represent the additional gain given to different frequency bands. According to psychoacoustics, the solution provided in the embodiment of the present application gives a certain enhancement to the signal components with higher frequencies.

[0055] It should be pointed out that It can represent the initial sub-coherent sound corresponding to the sky channel, by The application of can realize the mid-high frequency reprocessing of the initial sub-coherent sound corresponding to the sky channel based on psychoacoustics.

[0056] Overall, such as Figure 5 As shown in the figure, the vertical position difference between high and low frequency melodies increases, so that the listener can feel better vertical separation and atmosphere. Figure 5 This is a schematic diagram of the vertical listening experience before and after mid-high frequency processing when the main channel and the sky channel exist at the same time in an embodiment of the present application.

[0057] Continue to read Figure 1 In step S105, based on each sub-coherent sound and the dual-channel ambient sound, the second frequency domain signals of N target channels are determined. The dual channels are the left channel and the right channel. Figure 2 Schematically shows the sub-coherent sound corresponding to the center channel. , the sub-coherent sound corresponding to the left front channel , the sub-coherent sound corresponding to the right front channel , sub-coherent sound corresponding to the sky channel , Ambient sound of the left channel and ambient sound on the right channel , perform coherent sound reconstruction of the ambient sound and generate the second frequency domain signal corresponding to the center channel , the second frequency domain signal corresponding to the left front channel , the second frequency domain signal corresponding to the right front channel , the second frequency domain signal corresponding to the sky channel .

[0058] Specifically, in one embodiment, the at least part of the channels include the first channel. Taking the first channel as the sky channel as an example, the following can be performed: Figure 6 The process shown determines the second frequency domain signal of the sky channel. Figure 6 FIG. 1 is a flow chart of a method for determining a second frequency domain signal of a sky channel in an embodiment of the present application. Figure 6 As shown, the determination method includes the following steps: S601: Determine a weight value corresponding to each channel in the dual-channel system based on the layout position of the sky channel; S603: Perform weighted summation on the binaural ambient sounds based on the weight values ​​to generate a target ambient sound; S605: Superimpose the sub-coherent sound corresponding to the sky channel and the target ambient sound to generate a second frequency domain signal of the sky channel.

[0059] Specifically, the target ambient sound can be expressed as: ; in, , Ambient sound for the left channel The weight value of Ambient sound for the right channel In addition, It is also used to reflect the left and right layout of the sky channel. The closer it is to 1, the further to the right the layout is. The closer to 0.

[0060] The second frequency domain signal of the sky channel The algorithm formula can be: .

[0061] By adopting Figure 6 The described determination method determines the second frequency domain signal corresponding to the sky channel, and can be combined with the layout position of the sky channel to superimpose the ambient sound and the sub-coherent sound after mid- and high-frequency processing in a certain proportion to enhance the sense of spatial atmosphere.

[0062] In one embodiment, the at least part of the channels include a second channel, and the sub-coherent sound corresponding to the second channel and the ambient sound of the channel in the binaural channel that is azimuthally related to the second channel can be superimposed to generate a second frequency domain signal of the second channel. Taking the second channel including the left front channel and the right front channel as an example, the second frequency domain signal of the left front channel is and the second frequency domain signal of the right front channel The algorithm formula can be: ; .

[0063] In one embodiment, the at least part of the channels include a third channel, and the average value of the ambient sound of the two channels and the sub-coherent sound corresponding to the third channel can be superimposed to generate a second frequency domain signal of the third channel. Taking the third channel as the center channel as an example, the second frequency domain signal corresponding to the center channel is The algorithm formula can be: .

[0064] In one embodiment, the N target channels also include a fourth channel in addition to at least some of the aforementioned channels, and the fourth channel is a channel in the second direction. The ambient sound of the dual-channel channel azimuthally related to the fourth channel can be determined as the second frequency domain signal of the fourth channel. For example, taking the fourth channel as a surround channel, where the surround channels include a left surround channel and a right surround channel, the ambient sound of the left channel can be determined as the second frequency domain signal of the left surround channel, and the ambient sound of the right channel can be determined as the second frequency domain signal of the right surround channel.

[0065] It should be noted that the correlation upmixing algorithm does not implement a suitable combination and distribution strategy for coherent and incoherent signals, often simply distributing the extracted ambient sound to a few sub-channels (such as surround channels and sky channels). However, as explained above with respect to step S105, the solution provided in this embodiment of the application can effectively re-combine and distribute the ambient sound and sub-coherent sound in conjunction with the spatial layout, thereby enhancing the sense of spatial surround and atmosphere.

[0066] Continue to read Figure 1 In step S107, based on each second frequency domain signal, a time domain audio signal corresponding to N target channels is generated. Specifically, an Inverse Short-Time Fourier Transform (ISTFT) is performed on each second frequency domain signal to generate a time domain audio signal corresponding to N target channels. Figure 2 Schematically shows the second frequency domain signal 、 、 、 Perform ISTFT separately to generate the time domain audio signal corresponding to the center channel , the time domain audio signal corresponding to the left front channel , the time domain audio signal corresponding to the right front channel Time domain audio signal corresponding to the sky channel .

[0067] when Figure 1When the audio processing method described is applied to a car audio system, after generating time-domain audio signals corresponding to N target channels, the car audio system can use each speaker on board to play the time-domain audio signals corresponding to the target channels. Figure 7 As shown. Among them, Figure 7 This is a schematic diagram showing the comparison of the listening experience inside the car.

[0068] Figure 8 Schematic diagram of the structure of the audio processing device in the embodiment of the present application. Figure 8 As shown, the audio processing device 800 includes: The first extraction unit 801 is configured to extract a first coherent sound, position information of the first coherent sound, and a binaural ambient sound from a first frequency domain signal of the binaural audio; The second extraction unit 802 is configured to extract sub-coherent sounds corresponding to at least some of the N target channels from the first coherent sound based on the direction information, where N is an integer greater than or equal to 3; A determining unit 803 is configured to determine second frequency domain signals of N target channels based on the sub-coherent sounds and the binaural ambient sound; The conversion unit 804 is configured to generate time-domain audio signals corresponding to N target channels based on each second frequency-domain signal.

[0069] In one embodiment, the second extraction unit 802 is configured to extract, from the first coherent sound, sub-coherent sounds corresponding to at least some of the N target channels based on the orientation information, including: The second extraction unit 802 is configured to determine, for each target channel in the at least some of the channels, using a first gain function, a first gain value corresponding to the target channel when extracting the sub-coherent sound corresponding to the target channel based on the orientation information, and determine the sub-coherent sound corresponding to the target channel based on the first coherent sound and the first gain value.

[0070] In one embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction; The second extraction unit 802 is configured to determine the sub-coherent sound corresponding to the target channel based on the first coherent sound and the first gain value, including: Determining a second gain value related to the frequency of the first coherent sound using a second gain function; A sub-coherent sound corresponding to the first channel is determined based on the first coherent sound, the first gain value corresponding to the first channel, and the second gain value.

[0071] In one embodiment, the at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction; The determining unit 803 is configured to determine the second frequency domain signals of N target channels based on the sub-coherent sounds and the binaural ambient sound, including: Determine the weight value corresponding to each channel in the dual channels based on the layout position of the first channel; Perform weighted summation of the binaural ambient sounds based on the weight values ​​to generate the target ambient sound; The sub-coherent sound corresponding to the first sound channel and the target ambient sound are superimposed to generate a second frequency domain signal of the first sound channel.

[0072] In one embodiment, the at least part of the sound channels includes a second sound channel, and the second sound channel is a sound channel in a second direction; The determining unit 803 is configured to determine the second frequency domain signals of N target channels based on the sub-coherent sounds and the binaural ambient sound, including: The determining unit 803 is configured to superimpose the sub-coherent sound corresponding to the second channel and the ambient sound of the channel in the two channels that is directionally related to the second channel to generate a second frequency domain signal of the second channel.

[0073] In one embodiment, the at least part of the sound channels includes a third sound channel, and the third sound channel is a sound channel in the second direction; The determining unit 803 is configured to determine the second frequency domain signals of N target channels based on the sub-coherent sounds and the binaural ambient sound, including: The determining unit 803 is configured to superimpose the average value of the dual-channel ambient sound and the sub-coherent sound corresponding to the third channel to generate a second frequency domain signal of the third channel.

[0074] In one embodiment, the N target sound channels further include a fourth sound channel in addition to at least some of the above-mentioned sound channels, and the fourth sound channel is a sound channel in the second direction; The determining unit 803 is configured to determine the second frequency domain signals of N target channels based on the sub-coherent sounds and the binaural ambient sound, including: The determining unit 803 is configured to determine the ambient sound of the channel in the dual channels that is directionally related to the fourth channel as the second frequency domain signal of the fourth channel.

[0075] In one embodiment, the first extraction unit 801 is configured to extract the first coherent sound, the orientation information of the first coherent sound, and the binaural ambient sound from the first frequency domain signal of the binaural audio, including: The first extraction unit 801 is configured to extract the first coherent sound, the direction information and the binaural ambient sound from the first frequency domain signal by using the least square method.

[0076] It should be noted that other aspects and implementation details of the audio processing device 800 are the same as or similar to the audio processing method described above and will not be repeated here.

[0077] The embodiment of the present application further provides an audio output system, which includes a plurality of speakers corresponding to N target channels, and Figure 8 The audio processing device 800 is shown. The audio output system can be located in a car, an aircraft, a theater, a conference room, or a classroom, etc., which is not specifically limited here. When the audio output system is located in a car, it can be called a car audio system.

[0078] The embodiment of the present application further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following is achieved: Figure 1 Describes the audio processing method.

[0079] The embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the following is achieved: Figure 1 Describes the audio processing method.

[0080] The present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the following Figure 1 Describes the audio processing method.

[0081] According to the above description, the solution provided by the embodiment of the present application has the following advantages: Advantage 1: Enhanced 3D immersion The related upmixing algorithm only considers the horizontal direction and lacks spatial expression in the vertical dimension, resulting in a flat listening experience; however, the embodiment of the present application introduces the psychoacoustic characteristics of the pitch effect and the spatial response design value in the vertical direction, which can effectively expand the vertical listening experience and thus form a 3D sound effect.

[0082] Advantage 2: Clearer sound source separation The related upmixing algorithm does not handle coherent sound perfectly and lacks remodeling of its spatial physical characteristics, resulting in blurred sound and image and insufficient separation. The embodiment of the present application makes full use of the intrinsic spatial orientation information of coherent sound, obtains multiple sub-coherent sounds in different spatial orientations and performs spatial reprojection to adapt to the multi-channel multi-speaker system in the car, which can significantly increase separation and layering.

[0083] Advantage 3: More reasonable sound source combination The related upmixing algorithm does not make a suitable combination and allocation strategy for the processing of coherent and incoherent signals, and often simply allocates the extracted ambient sound to some sub-channels (surround sound channels, sky sound channels, etc.); the embodiment of the present application combines the spatial layout to effectively recombine and allocate the ambient sound and coherent sound (sub-coherent sound), which can enhance the spatial surround and atmosphere.

[0084] The above description is only a partial implementation of the embodiments of the present application and does not constitute any form of limitation to the application. The protection scope of the embodiments of the present application is not limited thereto. Any simple modifications, equivalent changes and modifications that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed in the embodiments of the present application should be covered within the protection scope of the embodiments of the present application.

Claims

1. An audio processing method, characterized in that: include: Extracting a first coherent sound, position information of the first coherent sound, and a binaural ambient sound from a first frequency domain signal of the binaural audio; Extracting, based on the orientation information, sub-coherent sounds corresponding to at least some of the N target channels from the first coherent sound, wherein N is an integer greater than or equal to 3; Determining second frequency domain signals of the N target channels based on each of the sub-coherent sounds and the dual-channel ambient sound; Based on each of the second frequency domain signals, time domain audio signals corresponding to the N target channels are generated.

2. The audio processing method according to claim 1, wherein: Extracting sub-coherent sounds corresponding to at least some of the N target channels from the first coherent sound based on the direction information includes: For each target channel in the at least some of the channels, a first gain function is used to determine a first gain value corresponding to the target channel when the sub-coherent sound corresponding to the target channel is extracted based on the orientation information, and the sub-coherent sound corresponding to the target channel is determined based on the first coherent sound and the first gain value.

3. The audio processing method according to claim 2, characterized in that The at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction; Determining the sub-coherent sound corresponding to the target channel based on the first coherent sound and the first gain value includes: determining, using a second gain function, a second gain value related to the frequency of the first coherent sound; The sub-coherent sound corresponding to the first channel is determined based on the first coherent sound, the first gain value corresponding to the first channel, and the second gain value.

4. The audio processing method according to claim 1, wherein: The at least part of the sound channels includes a first sound channel, and the first sound channel is a sound channel in a first direction; Determining second frequency domain signals of the N target channels based on each of the sub-coherent sounds and the dual-channel ambient sound includes: Determining a weight value corresponding to each of the two channels based on the layout position of the first channel; Performing weighted summation on the binaural ambient sounds based on the weight values ​​to generate a target ambient sound; The sub-coherent sound corresponding to the first sound channel and the target ambient sound are superimposed to generate the second frequency domain signal of the first sound channel.

5. The audio processing method according to claim 1, wherein: The at least part of the sound channels includes a second sound channel, and the second sound channel is a sound channel in a second direction; Determining second frequency domain signals of the N target channels based on each of the sub-coherent sounds and the dual-channel ambient sound includes: The sub-coherent sound corresponding to the second channel and the ambient sound of a channel in the two channels that is azimuthally related to the second channel are superimposed to generate the second frequency domain signal of the second channel.

6. The audio processing method according to claim 1, wherein: The at least part of the sound channels includes a third sound channel, and the third sound channel is a sound channel in the second direction; Determining second frequency domain signals of the N target channels based on each of the sub-coherent sounds and the dual-channel ambient sound includes: An average value of the dual-channel ambient sound and the sub-coherent sound corresponding to the third channel are superimposed to generate the second frequency domain signal of the third channel.

7. The audio processing method according to claim 1, characterized in that: The N target sound channels further include a fourth sound channel other than the at least part of the sound channels, and the fourth sound channel is a sound channel in the second direction; Determining second frequency domain signals of the N target channels based on each of the sub-coherent sounds and the dual-channel ambient sound includes: The ambient sound of a channel in the dual channels that is directionally related to the fourth channel is determined as a second frequency domain signal of the fourth channel.

8. The audio processing method according to any one of claims 1 to 7, characterized in that: The step of extracting the first coherent sound, the position information of the first coherent sound, and the binaural ambient sound from the first frequency domain signal of the binaural audio includes: The first coherent sound, the direction information, and the binaural ambient sound are extracted from the first frequency domain signal using a least squares method.

9. An audio processing device, characterized in that: include: A first extraction unit is configured to extract a first coherent sound, position information of the first coherent sound, and a binaural ambient sound from a first frequency domain signal of the binaural audio; a second extraction unit configured to extract, from the first coherent sound, sub-coherent sounds corresponding to at least some of the N target sound channels based on the orientation information; wherein N is an integer greater than or equal to 3; a determining unit configured to determine second frequency domain signals of the N target channels based on the sub-coherent sounds and the binaural ambient sound; The conversion unit is configured to generate time-domain audio signals corresponding to the N target channels based on each of the second frequency-domain signals.

10. An audio output system, characterized in that: The system comprises a plurality of loudspeakers corresponding to N target sound channels, and the audio processing device according to claim 9.

Citation Information

Patent Citations

  • Stereophonic audio processing method and device

    CN104378728A

  • Audio processing method, device and equipment and storage medium

    CN111615045A

  • Apparatus and method for audio processing

    CN115190414A

  • Audio signal separation method and device, electronic equipment and storage medium

    CN119785819A

  • Headphones for the electroacoustic conversion of a stereo signal

    WO2010003160A1