Audio processing method and device, electronic equipment and storage medium

By using stereo impulse response in audio signal processing to simulate the reflection and attenuation effect of the acoustic space, the problem of unreal single-channel impulse response data generated by the simulator is solved, and the reverberation effect of the audio signal is achieved more realistic sound effects.

CN120282088APending Publication Date: 2025-07-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202410030700.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the single-channel impulse response data generated by the simulator is not true enough, resulting in poor reverberation effect, affecting the nature and authenticity of sound effects, and limiting the application of reverberation processing in audio signal processing.

Method used

By determining the stereo impulse response, the sound reflection and attenuation effect in the acoustic space is simulated, and a second audio signal is generated by convolutional processing, followed by audio mixing, enhancing the spatial and stereoscopic sense of the audio signal.

Benefits of technology

It realizes that the reverberation effect of the audio signal is more realistic, increases the spatial and layering of the sound, makes the sound more full and three-dimensional, and solves the problem of unreal single-channel impulse response data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282088A_ABST
    Figure CN120282088A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio processing method and device, electronic equipment and a storage medium. The method comprises the following steps: determining a target pulse response of a first audio signal, wherein the target pulse response is a stereo pulse response used for simulating sound reflection and attenuation effects of the first audio signal in an acoustic space; performing convolution processing on the first audio signal and the target pulse response to obtain a second audio signal; and performing audio mixing on the first audio signal and the second audio signal to output a third audio signal. According to the scheme provided by the invention, the stereo pulse response data capable of simulating the reflection and attenuation effects of the sound in the space can be synchronously determined when reverberation is carried out on the audio signal, and reverberation is carried out on the audio signal by using the stereo pulse response data, so that the sound is heard to be generated in different space environments; the reverberation effect of the audio signal is more vivid, the sense of space and the sense of layering of the sound are increased, and the sound is fuller and stereoscopic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and in particular, to an audio processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the continuous development of technology, the means of audio signal processing have become increasingly diverse. For example, reverberation technology is used to process audio signals to simulate the sound reflection and attenuation effects in different environments, giving people an immersive experience. However, the single-channel impulse response data generated by the simulator is not realistic enough, resulting in poor reverberation effects, unnatural and unrealistic sounds, thus affecting the quality and fidelity of the reverberation, and further causing the sound effect generation using reverberation processing to not be widely popularized and applied. Summary of the Invention

[0003] The present disclosure provides an audio processing method, apparatus, electronic device, and storage medium to achieve a more realistic reverberation effect for audio signals, increase the sense of space and layering of sounds, and make the sounds more full and three-dimensional.

[0004] In a first aspect, an embodiment of the present disclosure provides an audio processing method, the method including:

[0005] Determine a target impulse response of a first audio signal, where the target impulse response is a stereo impulse response for simulating the sound reflection and attenuation effects of the first audio signal in an acoustic space;

[0006] Obtain a second audio signal by performing convolution processing on the first audio signal and the target impulse response;

[0007] Output a third audio signal by performing audio mixing on the first audio signal and the second audio signal.

[0008] In a second aspect, an embodiment of the present disclosure further provides an audio processing apparatus, the apparatus including:

[0009] A determination module, configured to determine a target impulse response of a first audio signal, where the target impulse response is a stereo impulse response for simulating the sound reflection and attenuation effects of the first audio signal in an acoustic space;

[0010] A reverberation processing module, configured to obtain a second audio signal by performing convolution processing on the first audio signal and the target impulse response;

[0011] An audio output module, configured to output a third audio signal by performing audio mixing on the first audio signal and the second audio signal.

[0012] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device including:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein

[0015] the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the audio processing method according to any one of the above embodiments.

[0016] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable medium storing computer instructions for causing a processor to implement the audio processing method according to any one of the above embodiments when executed.

[0017] In the embodiment of the present disclosure, when a first audio signal is received, a target impulse response for reverberating the first audio signal is determined. The target impulse response is a stereo impulse response for simulating the sound reflection and attenuation effects of the first audio signal in an acoustic space. By convolving the first audio signal with the target impulse response, a second audio signal after reverberation processing is obtained. Furthermore, the first audio signal and the second audio signal can be mixed to output a third audio signal. This solution can solve the problem that the mono-channel impulse response data generated by the simulator is not realistic enough, resulting in poor reverberation effects. When reverberating an audio signal, a stereo impulse response data that can simulate the sound reflection and attenuation effects in space is determined synchronously. Using the stereo impulse response data to reverberate the audio signal makes the sound seem to be generated in different spatial environments, achieving a more realistic reverberation effect for the audio signal, increasing the sense of space and layering of the sound, and making the sound more full and three-dimensional.

[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.

[0020] Figure 1 is a flowchart of an audio processing method provided by an embodiment of the present disclosure;

[0021] Figure 2aIt is a schematic diagram showing the sound after multiple reflections, scattering, and attenuation provided by an embodiment of the present disclosure;

[0022] Figure 2b It is a signal schematic diagram of the impulse response provided by an embodiment of the present disclosure;

[0023] Figure 2c It is a schematic diagram of performing convolutional reverberation on an audio signal provided by an embodiment of the present disclosure;

[0024] Figure 3 It is a flowchart schematic diagram of another audio processing method provided by an embodiment of the present disclosure;

[0025] Figure 4 It is a schematic diagram of the principle of generating a stereo impulse response by a stereo impulse response generator provided by an embodiment of the present disclosure;

[0026] Figure 5 It is a schematic diagram of the structure of an audio processing device provided by an embodiment of the present disclosure;

[0027] Figure 6 It is a schematic diagram of the structure of an electronic device for implementing an audio processing method provided by an embodiment of the present disclosure. Detailed Embodiments

[0028] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0029] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0030] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0031] Note that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0032] Note that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes, and are not used to limit the scope of these messages or information.

[0034] Figure 1 The figure is a schematic flow chart of an audio processing method provided by an embodiment of this disclosure. The embodiments of this disclosure are applicable to the situation of performing reverberation processing on an audio signal. This method can be executed by an audio processing device, which can be implemented in the form of software and / or hardware, and is generally integrated in any electronic device with network communication functions. The electronic device can be a mobile terminal, a PC or a server, etc.

[0035] As Figure 1 shown, the audio processing method of the embodiments of this disclosure may include the following processes:

[0036] S110. Determine the target impulse response of the first audio signal. The target impulse response is a stereo impulse response used to simulate the sound reflection and attenuation effects of the first audio signal in an acoustic space.

[0037] Referring to Figure 2a , for the first audio signal emitted from a sound source, when the first audio signal propagates in an actual or virtual acoustic space, it will interact with the objects in the acoustic space, generating sound reflection and attenuation, and these reflections and attenuations will affect aspects such as the timbre, volume and duration of the sound. For example, when the sound of the first audio signal propagates indoors, it will be reflected by obstacles such as walls, ceilings, and floors, and each reflection will absorb some by the obstacles. In this way, when the sound stops, the sound will be reflected and absorbed multiple times indoors before finally disappearing. This phenomenon is called reverberation, and this period of time is called the reverberation time.

[0038] Reverberation is a description of the auditory perception formed after multiple reflections and attenuations of sound in an acoustic space. The multiple reflections, scattering and attenuation "aftertones" after the direct sound will enhance the sense of space, sound depth and clarity in subjective listening. Referring to Figure 2b, in order to endow the first audio signal with a reverberation effect, a corresponding target impulse response for reverberation can be allocated to the first audio signal. The target impulse response is a stereo impulse response used to simulate the sound reflection and attenuation effects of the first audio signal in an actual or virtual acoustic space. As a mathematical model for simulating these sound reflection and attenuation effects, the target impulse response can be generated according to the characteristics of the actual or virtual acoustic space.

[0039] Among them, the stereo impulse response can describe the reflection and attenuation of the input audio signal. The stereo impulse response contains acoustic characteristics and spatial information. By analyzing the stereo impulse response, the response at different frequencies and the propagation and reflection of sound in space can be obtained. In audio processing, the stereo impulse response is used to simulate the acoustic effects of different spaces, such as the reverberation, reflection, and attenuation of a room. By applying the target impulse response to audio signal processing, different spatial acoustic reverberation effects can be simulated.

[0040] S120. Obtain a second audio signal by performing convolution processing on the first audio signal and the target impulse response.

[0041] Based on the implementation of the modeled reverberation effect, digital filters are often used to achieve the reverberation effect, such as the Schroeder algorithm and the Moorer algorithm. These algorithms are simple to implement and have low computational complexity. However, there has always been a lack in the authenticity and naturalness of this kind of reverberation. Therefore, after generating a suitable stereo impulse response for the first audio signal, the first audio signal can be convolved with the target impulse response to apply the target impulse response to audio signal processing to simulate different spatial reverberation effects. In this process, the target impulse response, as a stereo impulse response data, can simulate the process of sound reflection, scattering, and attenuation in space. By performing appropriate reverberation processing on the first audio signal, the first audio signal can be made more realistic, natural, stereo, mellow, and clear.

[0042] Optionally, in the convolution processing of the first audio signal and the target impulse response, the target impulse response represents the reflection and attenuation of the sound of the first audio signal in the acoustic space. The target impulse response can be considered as a model function that describes the situation of the sound after multiple reflections and attenuations in the acoustic space after being emitted from the sound source. By performing convolution processing on the first audio signal and the target impulse response, the propagation and reflection of sound in the acoustic space can be simulated, thereby obtaining a second audio signal with a reverberation effect.

[0043] As an optional but non-limiting implementation manner, obtaining a second audio signal by performing convolution processing on the first audio signal and the target impulse response includes the following steps A1 - A2:

[0044] Step A1: Convolve the left and right channel audio signals of the first audio signal with the target impulse response respectively to obtain the left channel audio processing result and the right channel audio processing result corresponding to the first audio signal.

[0045] Step A2: Mix the left channel audio processing result and the right channel audio processing result corresponding to the first audio signal respectively for left and right channel mixing, and then output them to the left and right channels, and generate a second audio signal with a stereo effect based on the audio signals output to the left and right channels.

[0046] See Figure 2c , after determining the target impulse response of the first audio signal, there are a left channel audio signal and a right channel audio signal corresponding to the first audio signal. Convolve the left channel audio signal corresponding to the first audio signal with the target impulse response to obtain the left channel audio processing result corresponding to the first audio signal, and convolve the right channel audio signal corresponding to the first audio signal with the target impulse response to obtain the right channel audio processing result corresponding to the first audio signal.

[0047] See Figure 2c , for the left channel audio processing result wet_L and the right channel audio processing result wet_R corresponding to the first audio signal, a part of the right channel audio processing result corresponding to the first audio signal can be mixed in the left channel audio processing result corresponding to the first audio signal, and the mixed result is output to the left channel; at the same time, a part of the left channel audio processing result corresponding to the first audio signal can also be mixed in the right channel audio processing result corresponding to the first audio signal, and the mixed result is output to the right channel. By mixing the left channel audio processing result and the right channel audio processing result respectively for left and right channel mixing, the reverberation intensity of the audio signals of the left and right channels can be dynamically adjusted, and the audio signals output to the left and right channels are further mixed into a second audio signal with a stereo effect.

[0048] As an optional but non-limiting implementation manner, convolving the left and right channel audio signals of the first audio signal with the target impulse response respectively includes the following steps B1 - B3:

[0049] Step B1: Frame and window the left and right channel audio signals of the first audio signal respectively, and obtain the frequency domain results corresponding to the left and right channel audio signals of the first audio signal respectively by performing Fourier transform on each windowed left and right channel audio frame segment.

[0050] Step B2: Window the target impulse response, and obtain the frequency domain result of the target impulse response by performing Fourier transform on the windowed target impulse response.

[0051] Step B3: Fourier inverse transform is performed after the frequency-domain results corresponding to the left and right channel audio signals of the first audio signal are respectively multiplied in the frequency domain with the frequency-domain result of the target impulse response.

[0052] See Figure 2c , the left and right channel audio signals of the input first audio signal are respectively framed. For example, the left channel audio signal of the first audio signal is divided into small segments of audio signals, and each small segment of signal is called a frame, denoted as a left channel audio frame segment. The frame length is generally taken as 20 - 50 milliseconds. Further, a window function is applied to each left and right channel audio frame segment. The function of the window function is to gradually reduce the amplitude of the signal at the start and end of the audio frame to reduce the discontinuity at the frame boundary. Common window functions include rectangular window, Hanning window, Hamming window, etc.

[0053] See Figure 2c , Fourier transform (FFT) is respectively performed on each windowed left and right channel audio frame segment to obtain the frequency-domain results corresponding to the left and right channel audio signals of the first audio signal. At the same time, a window function can be directly applied to the target impulse response, and the frequency-domain result of the target impulse response is obtained by performing Fourier transform on the windowed target impulse response. Further, the frequency-domain results corresponding to the left and right channel audio signals of the first audio signal above are respectively multiplied with the frequency-domain result of the target impulse response, and the multiplied result is converted to the time domain through Fourier inverse transform. Then, the data is spliced using the overlap-and-add method / overlap-save method, and the left channel audio processing result wet_L corresponding to the first audio signal and the right channel audio processing result wet_R corresponding to the first audio signal are output.

[0054] Optionally, converting the multiplication result of the frequency-domain results corresponding to the left and right channel audio signals of the first audio signal with the frequency-domain result of the target impulse response to the time domain means converting the processing performed in the frequency domain or other domains back to the time domain, which can be achieved through inverse Fourier transform (IFFT) or other appropriate time-domain conversion methods.

[0055] Optionally, the overlap-and-add method or overlap-save method to splice data is a commonly used data splicing method in signal processing for combining segmented processed data into a continuous output signal. Overlap-and-add method: In this method, adjacent audio frame segments have a certain overlapping part in time. When splicing, the overlapping part of each audio frame segment is added to the starting part of the next audio frame segment for smooth transition; Overlap-save method: Similar to the overlap-and-add method, but when splicing, the overlapping part of each audio frame segment is retained instead of being added, so that more original audio signal information can be retained.

[0056] As an optional but non-limiting implementation, the left-channel audio processing result and the right-channel audio processing result corresponding to the first audio signal are respectively mixed for the left and right channels and then output to the left and right channels, including the following steps C1 - C2:

[0057] Step C1: Mix the right-channel audio processing result corresponding to the first audio signal into the left-channel audio processing result corresponding to the first audio signal according to a preset left-right channel mixing ratio and output it to the left channel.

[0058] Step C2: Mix the left-channel audio processing result corresponding to the first audio signal into the right-channel audio processing result corresponding to the first audio signal according to a preset left-right channel mixing ratio and output it to the right channel to generate a second audio signal with a stereo effect.

[0059] For the left-channel audio processing result wet_L and the right-channel audio processing result wet_R corresponding to the first audio signal, the stereo effect after mixing is S. The following process can be used for mixing: For the left-channel audio processing result wet_L and the right-channel audio processing result wet_R corresponding to the first audio signal, the left and right channel audio signals after mixing can be calculated respectively according to the mixing weight (such as p) as follows:

[0060] Left-channel input L = (1 - p) * left-channel processing result + p * right-channel processing result

[0061] Right-channel input R = (1 - p) * right-channel processing result + p * left-channel processing result

[0062] The audio signal L of the left channel and the audio signal R of the right channel are respectively input to the left and right channels to generate a stereo effect S. Among them, the value range of the mixing weight p should be from 0 to 1, which is used to control the balance of the left and right channels, and this can be set in response to user operations to determine the value of the mixing weight.

[0063] S130: Output a third audio signal by mixing the first audio signal and the second audio signal.

[0064] Optionally, the first audio signal (dry sound) and the second audio signal (wet sound) are mixed according to a specified ratio to obtain a third audio signal with a richer and more natural audio effect. For example,

[0065] Obtain the first audio signal (dry sound) and the second audio signal (wet sound), determine the dry-wet sound mixing ratio of the first audio signal (dry sound) and the second audio signal (wet sound), and mix the first audio signal (dry sound) and the second audio signal (wet sound) according to the determined dry-wet sound mixing ratio, which can be achieved through mathematical operations such as an adder or a multiplier. The mixed third audio signal will contain the components of the first audio signal and the processed second audio signal, thereby obtaining a richer and more natural audio effect.

[0066] In the technical solution of the embodiment of the present disclosure, when the first audio signal is received, the target impulse response used for reverberation processing of the first audio signal will be determined. The target impulse response is a stereo impulse response used to simulate the sound reflection and attenuation effects of the first audio signal in an acoustic space. The second audio signal after reverberation processing is obtained by convolving the first audio signal with the target impulse response. Furthermore, the first audio signal and the second audio signal can be mixed to output a third audio signal. This solution can solve the problem that the single-channel impulse response data generated by the simulator is not realistic enough, resulting in poor reverberation effects. When reverberating the audio signal, the stereo impulse response data that can simulate the sound reflection and attenuation effects in space will be determined synchronously. Using the stereo impulse response data to reverberate the audio signal makes the sound seem to be generated in different spatial environments, realizing a more realistic reverberation effect of the audio signal, increasing the sense of space and layering of the sound, and making the sound more full and three-dimensional.

[0067] Figure 3 It is a schematic flowchart of another audio processing method provided by the embodiment of the present disclosure. The technical solution of this embodiment further optimizes the process of determining the target impulse response of the first audio signal in the foregoing embodiment on the basis of the technical solution of the foregoing embodiment. This embodiment can be combined with each optional solution in one or more of the foregoing embodiments.

[0068] As Figure 3 shown, the audio processing method of the embodiment of the present disclosure may include the following processes:

[0069] S310. Determine the reference impulse response of the first audio signal. The reference impulse response is a mono impulse response representing the sound reflection and attenuation effects of the first audio signal in an acoustic space, and the reference impulse response is obtained by measuring or synthesizing the sound response in the acoustic space.

[0070] Optionally, the reference impulse response is a mono impulse response obtained by collecting in a real environment, physical modeling, or psychoacoustic modeling.

[0071] When reverberating, most often professional recording equipment is used to collect multi-channel impulse response data in a real environment, and then the input audio signal is processed using the real collected impulse response data to generate this sense of the sound space. Although using professional recording equipment to collect multi-channel impulse response data in a real environment has good effects when reverberating, professional recording equipment is usually expensive and requires recording in a specific environment, which may incur additional costs, making it impossible to widely popularize and apply the use of reverberation processing for sound effect generation. Therefore, a mono impulse response obtained through physical modeling or psychoacoustic modeling can be used.

[0072] S320. Delay and gain the reference impulse response of the first audio signal to obtain the target impulse response of the first audio signal, and the target impulse response can simulate the time difference and intensity difference generated when the sound of the first audio signal reaches the first sound pickup position and the second sound pickup position in space.

[0073] Among them, the target impulse response is a stereo impulse response used to simulate the sound reflection and attenuation effects of the first audio signal in the acoustic space.

[0074] See Figure 4 , the reference impulse response of the first audio signal, as the impulse response for reverberation, can be from a mono impulse response obtained by real environment collection, physical modeling, or psychoacoustic modeling. Determine the time difference between the sound source of the first audio signal reaching the first sound pickup position and the second sound pickup position respectively, and delay the reference impulse response of the first audio signal through the time difference to simulate the time difference effect generated when the sound of the first audio signal reaches the first sound pickup position and the second sound pickup position in space. At the same time, determine the intensity difference between the sound source of the first audio signal reaching the first sound pickup position and the second sound pickup position respectively, and further use the intensity difference to gain the reference impulse response of the first audio signal to simulate the intensity difference effect generated when the sound of the first audio signal reaches the first sound pickup position and the second sound pickup position in space.

[0075] See Figure 4 , the ITD algorithm can be used to calculate the time difference when the sound source of the first audio signal reaches the first sound pickup position and the second sound pickup position respectively, and the ILD can be used to calculate the sound pressure level difference generated by the different intensities when the sound source of the first audio signal reaches the first sound pickup position and the second sound pickup position respectively. Among them, the ITD algorithm parameters and the ILT algorithm parameters both include: the distance T between the first sound pickup position and the second sound pickup position, the distance d from the sound source to the first sound pickup position, the angular direction θ of the sound source relative to the perpendicular bisector of the line connecting the first sound pickup position and the second sound pickup position, and the sound propagation speed c in the air.

[0076] As an optional but non-limiting implementation, delaying and gaining the reference impulse response of the first audio signal to obtain the target impulse response of the first audio signal includes the following steps D1 - D3:

[0077] Step D1: Determine the time difference when the sound source of the first audio signal reaches the first sound pickup location and the second sound pickup location respectively, and use it as the target time difference.

[0078] Refer to Figure 4 , calculate the time difference when the sound source reaches the first sound pickup location: According to the propagation speed of sound in the air, the time difference when the sound source reaches the first sound pickup location can be calculated. The following formula can be used for calculation: Δt = d / c; where Δt represents the time difference when the sound source reaches the first sound pickup location, d represents the distance from the sound source to the first sound pickup location, and c represents the propagation speed of sound in the air. Calculate the time difference between reaching the first sound pickup location and the second sound pickup location: According to the angle direction θ of the sound source relative to the perpendicular bisector of the line connecting the first sound pickup location and the second sound pickup location and the distance between the first sound pickup location and the second sound pickup location, the time difference between reaching the first sound pickup location and the second sound pickup location can be calculated. The following formula can be used for calculation: ITD = T * sin(θ) / c; where ITD represents the time difference between the first sound pickup location and the second sound pickup location, T represents the distance between the first sound pickup location and the second sound pickup location, θ represents the angle direction θ of the sound source relative to the perpendicular bisector of the line connecting the first sound pickup location and the second sound pickup location, and c represents the propagation speed of sound in the air.

[0079] Step D2: Determine the sound pressure level difference generated by the different intensities when the sound source of the first audio signal reaches the first sound pickup location and the second sound pickup location respectively, and determine it as the target intensity difference.

[0080] Refer to Figure 4 , calculate the distance difference between the sound source of the first audio signal reaching the first sound pickup location and the second sound pickup location: According to the distance d from the sound source of the first audio signal to the center of the line connecting the first sound pickup location and the second sound pickup location and the angle direction θ of the sound source relative to the perpendicular bisector of the line connecting the first sound pickup location and the second sound pickup location, calculate the distance difference Δd between the sound source and the first sound pickup location and the second sound pickup location. The following formula can be used for calculation: Δd = T * sin(θ). Calculate the amplitude ratio between the first sound pickup location and the second sound pickup location: According to the distance difference Δd between the sound source of the first audio signal and the first sound pickup location and the second sound pickup location, the amplitude ratio α between the first sound pickup location and the second sound pickup location can be calculated. The following formula can be used for calculation: α = (d + Δd) / (d - Δd).

[0081] Step D3: Delay the reference impulse response of the first audio signal according to the target time difference, and after delaying the reference impulse response of the first audio signal, perform gain on the reference impulse response of the first audio signal according to the target time difference to obtain the target impulse response of the first audio signal.

[0082] See Figure 4 , delay the reference impulse response of the first audio signal by the time difference of ITD to simulate the time difference when the sound of the first audio signal arrives at different positions (between the left and right ears) in the acoustic space, which can be implemented using a delay module in digital signal processing technology. After delaying the reference impulse response of the first audio signal, further gain processing can be performed on the reference impulse response of the first audio signal to simulate the amplitude difference between the first sound pickup part and the second sound pickup part. The gain processing can be achieved by multiplying the reference impulse response signal of the first audio signal by the amplitude ratio α.

[0083] S330: Obtain the second audio signal by performing convolution processing on the first audio signal and the target impulse response.

[0084] S340: Output the third audio signal by performing audio mixing on the first audio signal and the second audio signal.

[0085] The technical solution of the embodiments of the present disclosure, when receiving the first audio signal, will determine the target impulse response used for reverberation processing of the first audio signal. The target impulse response is a stereo impulse response used to simulate the sound reflection and attenuation effects of the first audio signal in the acoustic space. By performing convolution processing on the first audio signal and the target impulse response, the second audio signal after reverberation processing is obtained, and then the first audio signal and the second audio signal can be mixed to output the third audio signal. This solution can solve the problem that the single-channel impulse response data generated by the simulator is not realistic enough to result in poor reverberation effects. When performing reverberation on the audio signal, the stereo impulse response data that can simulate the sound reflection and attenuation effects in the space will be determined synchronously. Using the stereo impulse response data to perform reverberation on the audio signal makes the sound seem to be generated in different spatial environments, realizing a more realistic reverberation effect of the audio signal, increasing the sense of space and hierarchy of the sound, and making the sound more full and three-dimensional.

[0086] Figure 5 It is a schematic structural diagram of an audio processing device provided by the embodiments of the present disclosure. The embodiments of the present disclosure are applicable to the situation of performing reverberation processing on audio signals. The audio processing device can be implemented in the form of software and / or hardware, and is generally integrated on any electronic device with network communication functions. The electronic device can be a mobile terminal, a PC, or a server, etc.

[0087] As shown Figure 5 below, the audio processing device according to the embodiments of the present disclosure may include the following: a determination module 510, a reverberation processing module 520, and an audio output module 530. Among them:

[0088] The determination module 510 is configured to determine a target impulse response of the first audio signal, where the target impulse response is a stereo impulse response for simulating the sound reflection and attenuation effects of the first audio signal in an acoustic space;

[0089] The reverberation processing module 520 is configured to obtain a second audio signal by performing convolution processing on the first audio signal and the target impulse response;

[0090] The audio output module 530 is configured to output a third audio signal by mixing the first audio signal and the second audio signal.

[0091] Based on the technical solutions of the above embodiments, optionally, determining the target impulse response of the first audio signal includes:

[0092] Determining a reference impulse response of the first audio signal, where the reference impulse response is a mono impulse response representing the sound reflection and attenuation effects of the first audio signal in an acoustic space, and the reference impulse response is obtained by measuring or synthesizing the sound response in the acoustic space;

[0093] Performing delay and gain on the reference impulse response of the first audio signal to obtain the target impulse response of the first audio signal, where the target impulse response can simulate the time difference and intensity difference generated when the sound of the first audio signal reaches the first sound pickup part and the second sound pickup part in space.

[0094] Based on the technical solutions of the above embodiments, optionally, the reference impulse response is a mono impulse response obtained by collecting in a real environment, physical modeling, or psychoacoustic modeling.

[0095] Based on the technical solutions of the above embodiments, optionally, performing delay and gain on the reference impulse response of the first audio signal to obtain the target impulse response of the first audio signal includes:

[0096] Determining the time difference when the sound source of the first audio signal reaches the first sound pickup part and the second sound pickup part respectively as the target time difference;

[0097] Determining the sound pressure level difference generated by the different intensities when the sound source of the first audio signal reaches the first sound pickup part and the second sound pickup part respectively and determining it as the target intensity difference;

[0098] Delay the reference impulse response of the first audio signal according to the target time difference, and after delaying the reference impulse response of the first audio signal, obtain the target impulse response of the first audio signal by gain processing the reference impulse response of the first audio signal according to the target time difference.

[0099] Based on the technical solution of the above embodiment, optionally, obtaining the second audio signal by performing convolution processing on the first audio signal and the target impulse response includes:

[0100] Convolve the left and right channel audio signals of the first audio signal with the target impulse response respectively to obtain the left channel audio processing result and the right channel audio processing result corresponding to the first audio signal;

[0101] Mix the left channel audio processing result and the right channel audio processing result corresponding to the first audio signal in the left and right channels respectively and output them to the left and right channels, and generate a second audio signal with a stereo effect based on the audio signals output to the left and right channels.

[0102] Based on the technical solution of the above embodiment, optionally, convolving the left and right channel audio signals of the first audio signal with the target impulse response respectively includes:

[0103] Frame and window the left and right channel audio signals of the first audio signal respectively, and obtain the frequency domain results corresponding to the left and right channel audio signals of the first audio signal respectively by performing Fourier transform on each windowed left and right channel audio frame segment;

[0104] Window the target impulse response, and obtain the frequency domain result of the target impulse response by performing Fourier transform on the windowed target impulse response;

[0105] Perform frequency domain multiplication on the frequency domain results corresponding to the left and right channel audio signals of the first audio signal and the frequency domain result of the target impulse response respectively, and then perform inverse Fourier transform.

[0106] Based on the technical solution of the above embodiment, optionally, mixing the left channel audio processing result and the right channel audio processing result corresponding to the first audio signal in the left and right channels respectively and outputting them to the left and right channels includes:

[0107] Mix the right channel audio processing result corresponding to the first audio signal into the left channel audio processing result corresponding to the first audio signal according to the preset left and right channel mixing ratio and output it to the left channel;

[0108] Mix the left-channel audio processing result corresponding to the first audio signal in the right-channel audio processing result corresponding to the preset left and right channel mixing ratio and output it to the right channel to generate a second audio signal with a stereo effect.

[0109] In the technical solution of the embodiment of the present disclosure, when a first audio signal is received, a target impulse response used for reverberation processing of the first audio signal will be determined. The target impulse response is a stereo impulse response used to simulate the sound reflection and attenuation effects of the first audio signal in an acoustic space. By convolving the first audio signal with the target impulse response, a second audio signal after reverberation processing is obtained. Furthermore, the first audio signal and the second audio signal can be mixed and output as a third audio signal. This solution can solve the problem that the single-channel impulse response data generated by the simulator is not realistic enough, resulting in poor reverberation effects. When reverberating an audio signal, a stereo impulse response data that can simulate the sound reflection and attenuation effects in space will be determined synchronously. Using the stereo impulse response data to reverberate the audio signal makes the sound seem to be generated in different spatial environments, realizing a more realistic reverberation effect of the audio signal, increasing the sense of space and hierarchy of the sound, and making the sound more full and three-dimensional.

[0110] The audio processing device provided by the embodiment of the present disclosure can execute the audio processing method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the audio processing method.

[0111] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiment of the present disclosure.

[0112] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Referring below to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiment of the present disclosure (such as Figure 6 the terminal device or server in). The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiment of the present disclosure.

[0113] AsFigure 6 As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An editing / output (I / O) interface 605 is also connected to the bus 604.

[0114] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0115] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the audio processing method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the audio processing method of the embodiment of the present disclosure are executed.

[0116] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0117] The electronic device provided by the embodiment of the present disclosure and the audio processing method provided by the above embodiment belong to the same inventive concept. The technical details not described in detail in this embodiment may be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0118] An embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the audio processing method provided in the above embodiment.

[0119] It should be noted that the computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0120] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of the communication network include a local area network (“LAN”), a wide area network (“WAN”), the Internet (for example, the Internet), and a peer-to-peer network (for example, an ad hoc peer-to-peer network), as well as any currently known or future-developed network.

[0121] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0122] The above computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine a target impulse response of a first audio signal, where the target impulse response is a stereo impulse response for simulating the sound reflection and attenuation effects of the first audio signal in an acoustic space; obtain a second audio signal by performing convolution processing on the first audio signal and the target impulse response; and output a third audio signal by performing audio mixing on the first audio signal and the second audio signal.

[0123] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0125] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Wherein, the name of the unit does not, in some cases, constitute a limitation on the unit itself. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0126] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, by way of non-limiting illustration, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), Systems on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so forth.

[0127] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be either a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0128] The foregoing description is only a preferred embodiment of the present disclosure and an illustration of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in this disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in this disclosure.

[0129] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0130] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An audio processing method, characterized in that, The method includes: Determining a target impulse response of a first audio signal, where the target impulse response is a stereo impulse response for simulating the sound reflection and attenuation effects of the first audio signal in an acoustic space; Obtaining a second audio signal by performing convolution processing on the first audio signal and the target impulse response; Outputting a third audio signal by performing audio mixing on the first audio signal and the second audio signal.

2. The method according to claim 1, wherein Determining the target impulse response of the first audio signal includes: Determining a reference impulse response of the first audio signal, where the reference impulse response is a mono impulse response representing the sound reflection and attenuation effects of the first audio signal in an acoustic space, and the reference impulse response is obtained by measuring or synthesizing the sound response in the acoustic space; Performing delay and gain on the reference impulse response of the first audio signal to obtain the target impulse response of the first audio signal, where the target impulse response can simulate the time difference and intensity difference generated when the sound of the first audio signal reaches the first sound pickup position and the second sound pickup position in space.

3. The method according to claim 2, wherein The reference impulse response is a mono impulse response obtained by collecting in a real environment, physical modeling, or psychoacoustic modeling.

4. The method according to claim 2, wherein Performing delay and gain on the reference impulse response of the first audio signal to obtain the target impulse response of the first audio signal includes: Determining the time difference when the sound source of the first audio signal reaches the first sound pickup position and the second sound pickup position respectively as the target time difference; Determining the sound pressure level difference generated by the different intensities when the sound source of the first audio signal reaches the first sound pickup position and the second sound pickup position respectively and determining it as the target intensity difference; Delaying the reference impulse response of the first audio signal according to the target time difference, and after delaying the reference impulse response of the first audio signal, performing gain on the reference impulse response of the first audio signal according to the target time difference to obtain the target impulse response of the first audio signal.

5. The method according to claim 1, wherein Obtaining a second audio signal by performing convolution processing on the first audio signal and the target impulse response includes: Convolving the left and right channel audio signals of the first audio signal with the target impulse response respectively to obtain the left channel audio processing result and the right channel audio processing result corresponding to the first audio signal; Mixing the left channel audio processing result and the right channel audio processing result corresponding to the first audio signal in the left and right channels respectively and outputting them to the left and right channels, and generating a second audio signal with a stereo effect based on the audio signals output to the left and right channels.

6. The method according to claim 5, characterized in that Convolving the left and right channel audio signals of the first audio signal with the target impulse response respectively includes: Framing and windowing the left and right channel audio signals of the first audio signal respectively, and obtaining the frequency domain results corresponding to the left and right channel audio signals of the first audio signal respectively by performing Fourier transform on each windowed left and right channel audio frame segment; Windowing the target impulse response, and obtaining the frequency domain result of the target impulse response by performing Fourier transform on the windowed target impulse response; The frequency-domain results corresponding to the left and right channel audio signals of the first audio signal are respectively multiplied in the frequency domain with the frequency-domain result of the target impulse response and then subjected to inverse Fourier transform.

7. The method according to claim 5, characterized in that, Mixing the left-channel audio processing result and the right-channel audio processing result corresponding to the first audio signal respectively for left and right channel mixing and then outputting to the left and right channels, including: Mixing the right-channel audio processing result corresponding to the first audio signal into the left-channel audio processing result corresponding to the first audio signal according to a preset left and right channel mixing ratio and outputting to the left channel; Mixing the left-channel audio processing result corresponding to the first audio signal into the right-channel audio processing result corresponding to the first audio signal according to a preset left and right channel mixing ratio and outputting to the right channel to generate a second audio signal with a stereo effect.

8. An audio processing device, characterized in that, The apparatus includes: A determination module, configured to determine a target impulse response of a first audio signal, where the target impulse response is a stereo impulse response for simulating the sound reflection and attenuation effects of the first audio signal in an acoustic space; A reverberation processing module, configured to obtain a second audio signal by performing convolution processing on the first audio signal and the target impulse response; An audio output module, configured to output a third audio signal by performing audio mixing on the first audio signal and the second audio signal.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the audio processing method according to any one of claims 1-7.

10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the audio processing method according to any one of claims 1-7 when executed by a computer processor.

Citation Information

Cited By

  • Method, device and application for realizing audio delay effect based on pulse response sequence

    CN121260180A

  • Method, device and application for implementing audio delay effect based on impulse response sequence

    CN121260180B

  • Room sound calibration method based on hearing tendency control and related equipment

    CN122054066A

  • Room sound calibration method based on hearing preference control and related device

    CN122054066B