Audio data processing method and electronic equipment
By performing filter calibration on the rendered spatial audio in electronic devices, the problem of degraded sound quality was solved, the timbre of the audio data was improved and power consumption was reduced, adapting to various rendering scenarios and enhancing the intelligence of the device.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2026-04-16
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies show a significant decrease in sound quality after rendering spatial audio, making it difficult to effectively improve the timbre loss of audio data.
By using filters in electronic devices to calibrate the timbre of rendered spatial audio, including calibrating the direct sound waveform, early reflection waveform, and late reverberation waveform separately, and combining different filter processing methods for different audio rendering scenarios, the power consumption is reduced and the sound quality is improved.
It effectively improves the sound quality of rendered audio data, maintains spatial sense while reducing power consumption, adapts to various spatial audio rendering scenarios, and enhances the intelligence of electronic devices.
Smart Images

Figure CN122054067A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to an audio data processing method and an electronic device. Background Technology
[0002] With the iteration of audio technology, spatial audio technology has been widely applied to various electronic devices, such as mobile phones and headphones. Understandably, electronic devices using spatial audio technology can render raw audio data based on the mechanisms of human hearing to reproduce the propagation, reflection, and reverberation characteristics of sound in three-dimensional space, creating an immersive listening experience that closely resembles a real sound field. However, after rendering, the sound quality of the raw audio data will show a significant decrease. Summary of the Invention
[0003] This application provides an audio data processing method and an electronic device for improving the sound quality of rendered audio data.
[0004] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide an audio data processing method applied to an electronic device. The electronic device is one that supports playback functionality. Optionally, the electronic device can be connected to headphones to play audio. Alternatively, the electronic device can play audio through a built-in speaker. In some embodiments, the electronic device can render and play spatial audio with a sense of space.
[0005] It is understandable that electronic devices can simulate different auditory effects of spatial audio played at different distances from the sound source (the distance between the sound source and the listener).
[0006] Take, for example, an electronic device connected to headphones, through which audio is played.
[0007] When the distance between the simulated sound source and the listener in the electronic device is the first distance, the corresponding first spatial audio is generated based on the first original audio (that is, the original data of the audio data to be played) through spatial audio rendering technology.
[0008] After rendering the first spatial audio, the first filter is used to perform timbre calibration on the first direct sound waveform in the first spatial audio, thereby improving the timbre loss of the rendered audio.
[0009] With the electronic device simulating a second distance between the sound source and the listener, a second spatial audio is rendered based on the first original audio using spatial audio rendering technology. The first and second spatial audios differ, for example, they have different auditory effects. After rendering the second spatial audio, a second filter is used to calibrate the timbre of the second direct sound waveform in the second spatial audio. A third filter is used to calibrate the timbre of the first early reflection waveform in the second spatial audio. A fourth filter is used to calibrate the timbre of the first late reverberation waveform in the second spatial audio.
[0010] In the above embodiments, when faced with scenarios that render different spatial audio, different methods can be used to calibrate the sound quality of the spatial audio, thereby improving the sound quality of the spatial audio, reducing power consumption, and enhancing the intelligence of electronic devices.
[0011] In some embodiments, when an electronic device is connected to headphones, calibrated spatial audio (e.g., first spatial audio and second spatial audio) can be played by the connected headphones. Exemplarily, the electronic device can connect to the headphones via wired or wireless means. After calibrating the spatial audio, the electronic device can send it to the headphones, instructing them to play it. In this way, the user wearing the headphones can hear audio with minimal timbre loss and a sense of space.
[0012] In a possible embodiment, if the electronic device is the headphones themselves, then the step of determining whether headphones are connected is skipped. After completing the spatial audio rendering, the timbre of the spatial audio is calibrated according to the method mentioned in the previous embodiment. Then, the calibrated spatial audio is played so that the user wearing the electronic device hears audio with minimal timbre loss and a sense of space.
[0013] In some embodiments, the electronic device may use binaural room impulse response (BRIR) or Ambisonics to process the original audio to obtain the corresponding spatial audio.
[0014] Taking BRIR-based audio rendering as an example, a first BRIR can be obtained, which is capable of rendering a binaural room impulse response at a sound source distance of a first distance. Then, the first BRIR is convolved with the first original audio to obtain the corresponding first spatial audio. Alternatively, a second BRIR can be obtained, which is capable of rendering a binaural room impulse response at a sound source distance of a second distance. Then, the second BRIR is convolved with the first original audio to obtain the corresponding second spatial audio. Thus, by combining the methods mentioned in the previous embodiments, the problem of timbre loss caused by the BRIR rendering process can be solved.
[0015] Taking audio rendering using Ambisonics as an example, multiple spherical points are selected as sound source locations on a sphere with a radius of a first distance. First-order ambisonics (FOA) or high-order ambisonics (HOA) impulse signals are obtained for each spherical point. Then, the first original audio is Ambisonics encoded to obtain the corresponding encoded signal. Convolution is performed between the encoded signal and the FOA (or HOA) impulse signal to output the FOA (or HOA) signal. The FOA (or HOA) signal is then Ambisonics decoded and convolved with the head correlation transfer function to obtain the spatial audio corresponding to that spherical point. Finally, the spatial audio corresponding to multiple spherical points is superimposed to obtain the corresponding first spatial audio. Similarly, a second spatial audio corresponding to a second distance can be rendered based on the first original audio, which will not be elaborated further here. In this way, combined with the method mentioned in the previous embodiments, the problem of timbre loss caused by FOA (or HOA) impulse response convolution, Ambisonics decoding, and the final head correlation transfer function convolution can be solved.
[0016] In the above embodiments, not only can the timbre loss problem in BRIR rendering scenarios be solved, but also the timbre loss problem in Ambisonics rendering scenarios. The electronic device can cope with various spatial audio rendering scenarios, increasing the applicability of the method.
[0017] In some embodiments, if the electronic device switches to enable Head-Related Impulse Response (HRIR) rendering of audio, then the loss of timbre can also be calibrated.
[0018] Continuing with the example of an electronic device connected to headphones, when the distance between the simulated sound source and the listener is the first distance, the electronic device can render a third spatial audio based on the second original audio. For example, it can convolve the HRIR corresponding to the first distance with the second original audio to generate the third spatial audio. Then, it uses a fifth filter to calibrate the timbre of the third spatial audio. The calibrated third spatial audio is then played through the headphones.
[0019] When simulating a second distance between the sound source and the listener, the electronic device can render a fourth spatial audio based on the second original audio. For example, it can convolve the second original audio with the HRIR corresponding to the second distance to generate the fourth spatial audio. Then, a sixth filter is used to calibrate the timbre of the fourth spatial audio. The calibrated fourth spatial audio is then played through headphones.
[0020] In the above embodiments, by combining the shorter and smoother effects of HRIR data (third-space audio and fourth-space audio), the electronic device can calibrate the audio rendered by HRIR by directly calibrating the timbre of the entire spatial audio, thereby expanding the application range.
[0021] In some embodiments, the electronic device may include filters capable of calibrating timbre, such as a fifth filter and a sixth filter. Thus, after the spatial audio is actually rendered, the fifth or sixth filter can be used to improve the timbre loss that occurred during the rendering process.
[0022] Taking the generation of the fifth filter as an example: Multiple first HRIRs are obtained. These first HRIRs are response signals formed when a unit pulse signal reaches the listener directly from multiple first sound source locations, which are different locations at the same distance from the listener. Based on these multiple first HRIRs, a first target HRIR is fused, for example, by superimposing multiple first HRIRs. Compared to the first HRIRs, the positional information carried by the aforementioned first target HRIR is eliminated, but the timbre information is preserved. Then, based on the first target HRIR, the corresponding fifth filter is generated. For example, the spectral information of the first target HRIR can be inverted to obtain the corresponding fifth filter.
[0023] In the above embodiments, the filter generated in the above manner has the ability to improve the sound quality of spatial audio rendered based on HRIR.
[0024] In some embodiments, after rendering the spatial audio, the electronic device can segment the portion of the spatial audio that needs to be timbre calibrated. For example, a first direct sound waveform can be segmented from the first spatial audio. Another example is segmenting a second direct sound waveform, a first early reflection waveform, and a first late reverberation waveform from the second spatial audio.
[0025] Taking the segmentation of the second spatial audio as an example, optionally, the segmented second direct sound waveform includes the first sampling point. It can be understood that spatial audio corresponds to multiple sampling points, each with a corresponding amplitude. The aforementioned first sampling point is the sampling point with the largest amplitude in the second spatial audio. Alternatively, the segmented first late reverberation waveform includes the second sampling point. The second sampling point is the starting point of the target line segment in the energy curve corresponding to the second spatial audio. The target line segment is a line segment that conforms to the characteristics of a ramp-like descent. Further optionally, the first early reflection waveform includes the waveform in the second spatial audio located between the second direct sound waveform and the first late reverberation waveform.
[0026] By using an appropriate filter, the second direct sound waveform, the first early reflection waveform, and the first late reverberation waveform are calibrated in timbre, and then the calibrated waveforms are superimposed to obtain the final spatial audio to be played.
[0027] Taking the division of the first spatial audio as an example, after calibrating the first direct sound waveform, the calibrated part (the first direct sound waveform) and the uncalibrated part (the early reflection waveform and the late reverberation waveform in the first spatial audio) can be superimposed to obtain the final spatial audio to be played.
[0028] In the above embodiments, the effect of calibrating the timbre is improved by calibrating the spatial audio in segments.
[0029] In some embodiments, the electronic device acquires a first BRIR corresponding to a first distance, and divides the first BRIR into a third direct sound waveform, a second early reflection waveform, and a second late reverberation waveform. A second filter is generated based on the third direct sound waveform. A third filter is generated based on the second early reflection waveform. A fourth filter is generated based on the second late reverberation waveform.
[0030] In the above embodiments, by generating filters for different waveforms (direct sound waveform, early reflection waveform, and late reverberation waveform) in advance, the electronic device can use the adapted filters to process different waveforms in the spatial audio during operation, thereby improving the accuracy of timbre calibration.
[0031] In some embodiments, the electronic device generates a third filter based on a second early reflection waveform as follows: First, the first spectral information of the second early reflection waveform is smoothed. Then, the smoothed first spectral information is inverted to obtain the corresponding second spectral information. Finally, the second spectral information is mapped to the time domain to generate the corresponding third filter.
[0032] In the above embodiments, the generated third filter is made more stable. Similarly, the process of generating the fourth filter is similar to that of the third filter, except that the fourth filter is generated based on the late reverberation waveform.
[0033] In some embodiments, the first frequency point in the second spectral information mentioned in the above embodiments has an opposite phase to the second frequency point of the first spectral information, and the corresponding absolute intensity values are the same. The first frequency point and the second frequency point are the same, both belonging to the first frequency band. Thus, the generated third filter can perform full timbre calibration on the frequency points of the first frequency band.
[0034] In some embodiments, the third frequency point in the second spectral information mentioned in the above embodiments is out of phase with the fourth frequency point in the first spectral information, and the absolute intensity value corresponding to the fourth frequency point is greater than the absolute intensity value corresponding to the third frequency point, and the difference is not greater than a preset intensity threshold. The third and fourth frequency points are the same, both belonging to the second frequency band. Thus, the generated third filter can perform limited timbre calibration on the frequency points of the second frequency band.
[0035] In this way, while being able to calibrate the timbre, it avoids timbre distortion caused by over-calibration of certain frequency bands.
[0036] In some embodiments, the frequency point of the first frequency band is lower than that of the second frequency band; for example, the second frequency band is a high-frequency band, and the first frequency band is a low-frequency band. Thus, when using the filters mentioned in the above embodiments to process spatial audio, the stability of low-frequency timbre can be ensured, and the room timbre information in the high-frequency portion can be preserved, making the calibrated spatial audio sound more natural.
[0037] In some embodiments, when the electronic device is not connected to headphones, or after disconnecting from headphones, it can play audio data using its built-in speaker. Taking the playback of third raw audio as an example, the electronic device can render fifth spatial audio based on the third raw audio and a crosstalk cancellation filter, and then play the fifth spatial audio through the built-in speaker of the electronic device.
[0038] In the above embodiments, the electronic device can dynamically adjust the method of calibrating the tone according to whether or not headphones are connected, so as to meet the needs of diverse application scenarios.
[0039] In some embodiments, the process of generating a compensated crosstalk cancellation filter capable of calibrating timbre is as follows: Multiple second HRIRs are acquired, each being a response signal formed by a unit pulse signal reaching the listener directly from multiple second sound source locations. The multiple second sound source locations are different points at the same distance from the listener. A second target HRIR is fused from the multiple second HRIRs. The second target HRIR is inverted to generate a corresponding seventh filter. A compensated HRIR is generated by convolving the seventh filter with a second HRIR. Then, the corresponding compensated crosstalk cancellation filter is calculated based on the compensated HRIR.
[0040] In some embodiments, the electronic device needs to determine that the energy proportion corresponding to the third direct sound waveform is greater than a first proportional threshold before triggering the activation of the first filter to perform timbre calibration on the first direct sound waveform in the first spatial audio. That is, when it is determined that the energy proportion of the direct sound waveform is greater than the first proportional threshold, only the timbre of the direct sound waveform in the spatial audio is calibrated. In this way, the overall timbre calibration effect of the spatial audio is ensured while reducing operating power consumption.
[0041] In some embodiments, before performing timbre calibration on the second direct sound waveform in the second spatial audio using the second filter, a second BRIR corresponding to the second distance is obtained. The electronic device determines that the energy proportion corresponding to the fourth direct sound waveform in the second BRIR is not greater than a first proportion threshold. That is, when it is determined that the energy proportion of the direct sound waveform is not greater than the first proportion threshold, multiple waveform regions of the spatial audio are calibrated to ensure the overall timbre calibration effect of the spatial audio.
[0042] In a second aspect, embodiments of this application provide an electronic device, a memory, and one or more processors; the memory is coupled to one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and one or more processors call the computer instructions to cause the electronic device to perform the methods described in the first aspect and any of its implementations.
[0043] Thirdly, embodiments of this application provide a computer-readable storage medium including instructions. When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method described in the first aspect and any implementation thereof.
[0044] Fourthly, embodiments of this application provide a computer program product, including a computer program or instructions that, when executed on an electronic device, cause the electronic device to perform the methods described in the first aspect and any of its implementations.
[0045] Fifthly, embodiments of this application provide a chip system applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the methods described in the first aspect and any of its implementations.
[0046] It should be understood that the second to fifth aspects of the embodiments of this application correspond to the technical solutions of the first aspect of the embodiments of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description
[0047] Figure 1 This is a schematic diagram of audio data propagation provided in an embodiment of this application; Figure 2 A schematic diagram of HRIR provided for embodiments of this application; Figure 3 This is an example diagram illustrating the propagation of audio data within a room model provided in an embodiment of this application. Figure 4 One of the example diagrams of the BRIR provided in the embodiments of this application; Figure 5 Example diagram of BRIR provided for embodiments of this application; Figure 6 This is one of the flowcharts illustrating the steps of generating a compensation filter provided in an embodiment of this application; Figure 7 At least one set of example test scenarios for candidate HRIRs provided in the embodiments of this application; Figure 8This is the second flowchart illustrating the steps for generating a compensation filter as provided in the embodiments of this application. Figure 9 Example diagram provided for generating a combined window for dividing a direct acoustic waveform in an embodiment of this application; Figure 10 Example diagram provided for generating a combined window for dividing late reverberation waveforms in embodiments of this application; Figure 11 Example diagrams of energy curves corresponding to different target distances provided in embodiments of this application; Figure 12 Example diagram provided for generating a combined window for dividing early reflection waveforms in embodiments of this application; Figure 13 This is an example diagram of the generation of inverse filter 2 provided in an embodiment of this application; Figure 14 This is the third flowchart illustrating the steps for generating a compensation filter in an embodiment of this application. Figure 15 An example diagram illustrating the uniform selection of multiple points on a sphere as provided in this application embodiment; Figure 16 A flowchart rendered by Ambisonics provided for an embodiment of this application; Figure 17 One of the example diagrams of the configuration interface for the spatial audio function provided in the embodiments of this application; Figure 18 Example diagram of the audio data processing method provided in the embodiments of this application; Figure 19 The second example diagram shows the configuration interface for the spatial audio function provided in the embodiments of this application; Figure 20 An example diagram illustrating the processing of audio data in a scenario where the electronic device provided in this application is not connected to headphones; Figure 21 This is an example diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0048] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.
[0049] The user operations mentioned in the embodiments of this application can also be replaced with other operations. For example, the operation may include one or more of the following: swipe operation, multi-finger swipe operation, single screen operation, multi-tap screen operation, multi-finger single screen operation, multi-finger multi-tap screen operation, long press operation. An operation may also consist of multiple sub-operations, which are not specifically limited in this embodiment of the application.
[0050] To make the following embodiments clear and concise, a brief introduction to the relevant concepts or technologies is given first: Spatial audio rendering technology: a technology that simulates the propagation of audio in three-dimensional space. Specifically, it is based on acoustic principles and signal processing technology, and achieves sound localization and spatialization by precisely controlling the time difference and intensity difference of each channel.
[0051] When rendering spatial audio, it is necessary to consider one or more of the following information: the location of the simulated sound source, the room model in which it is located, the distance between the sound source and the listener, and the orientation of the sound source. Among these, the room module can describe the environmental conditions for audio signal transmission, such as the shape and size of the room, the location and material of reflective surfaces within the room, and the positions of the sound source and the listener within the room model.
[0052] For example, by rendering an audio signal, a corresponding direct sound signal and / or reflected sound signal is generated.
[0053] Direct sound signal: An audio signal emitted from a sound source that reaches the listener directly through a straight path without any reflection or scattering.
[0054] Reflected sound signal: The audio signal emitted from the sound source, after being reflected and scattered once or multiple times by the reflective surfaces of the room model (such as walls, floors, ceilings, objects, etc.), reaches the listener.
[0055] Understandably, reflected sound signals exhibit hysteresis, are room-model related, and thus possess spatial characteristics. For example, based on the number of reflections and the delay between the reflected sound signal and the direct sound signal, reflected sound signals can be categorized into early reflection signals and late reverberation signals.
[0056] Early reflected signal: The audio signal emitted from the sound source, after passing through the reflective surface in the room model and undergoing 1 to 2 reflections and scatterings, arrives at the listener 5 to 80 ms after the direct sound signal reaches the listener.
[0057] Late reverberation signal: The audio signal emitted from the sound source, after passing through the reflective surfaces in the room model and undergoing multiple dense reflections and scatterings, arrives at the listener 80 ms after the direct sound signal reaches the listener.
[0058] In some embodiments, the methods for rendering audio may include: head-related impulse response (HRIR) rendering, binaural room impulse response (BRIR) rendering, and Ambisonics rendering.
[0059] HRIR rendering: Based on HRIR and audio signal convolution processing, the corresponding binaural audio signal is simulated, also known as the audio rendering signal. This binaural audio signal includes the direct sound signal from the sound source to the listener, excluding reflected sound signals; that is, it is unaffected by the room model.
[0060] Understandably, this HRIR can be used to describe the characteristics of a unit pulse signal from a sound source being transmitted to the left and right ears. This HRIR can restore the position of the sound source relative to the listener, but it does not have spatial characteristics.
[0061] like Figure 1 As shown, the straight-line distance 'a' of the unit pulse signal emitted from sound source A to the listener's left ear 101 is different from the straight-line distance 'b' to the listener's right ear 102. Thus, there is a time difference (i.e., binaural time difference) between the time the unit pulse signal from sound source A reaches the left ear 101 and the time it reaches the right ear 102, and a sound level difference (i.e., binaural sound level difference) between the signal after reaching the left ear 101 and the signal after reaching the right ear 102. Furthermore, after reflection from the auricle and head, the HRIR waveform curve received by the unit pulse signal in the left ear 101 differs from the HRIR waveform curve received by the right ear 102.
[0062] Figure 2 This illustrates the impulse response signal formed when a unit pulse signal from sound source A reaches the listener's ear. For example... Figure 2 As shown, curve 201 is the waveform of the HRIR received by the left ear 101 in the time domain. Curve 202 is the waveform of the HRIR received by the right ear 102 in the time domain.
[0063] In some embodiments, the HRIR varies depending on the location between the sound source and the listener. Additionally, different timbre information related to the human head results in different HRIRs. Here, head-related timbre information refers to the frequency domain characteristics formed by the scattering, reflection, diffraction, and filtering effects of sound waves of different frequencies produced by the physiological structures of the human head, auricle, and torso.
[0064] For example, HRIR = f(Position, Tone). Here, Position is the target position information used to describe the positional relationship between the sound source and the listener. Tone refers to the timbre information related to the human head. Taking the listener as the user's ear as an example, Position can include the horizontal azimuth and vertical elevation angles of the sound source relative to the listener's head, centered on the listener's head.
[0065] BRIR rendering: Based on BRIR and audio signal convolution processing, it simulates the corresponding binaural room audio signal, also known as the audio rendering signal. The binaural room audio signal includes: the direct sound signal reaching the listener, and the reflected sound signal formed by the audio signal after reflection and scattering in the room model.
[0066] Understandably, this BRIR can be used to describe the characteristics of a unit pulse signal from a sound source, transmitted to the left and right ears within a room model.
[0067] like Figure 3 As shown, within room model 301, sound sources B and C are located at different positions, and their positional relationships with the listener's head are also different. For example, sound source B has a horizontal azimuth angle of 0° and a vertical elevation angle of 0° relative to the listener's head, and a distance of 1m. Sound source C has a horizontal azimuth angle of 0° and a vertical elevation angle of 0° relative to the listener's head, and a distance of 3m.
[0068] For example, the unit pulse signal emitted by sound source B can reach the listener not only directly from sound source B, but also after reflection and scattering within room model 301.
[0069] Figure 4 The diagram shows the impulse response signal formed when a unit pulse signal from sound source B reaches the listener's ear.
[0070] like Figure 4 As shown, curve 401 is the waveform curve of the BRIR received by the left ear in the time domain, also known as the time-domain waveform. Specifically, the portion of curve 401 within time region 402 represents the waveform curve of the pulsating response signal (i.e., the direct sound signal) of the unit pulse signal from sound source B directly reaching the listener's left ear. The portions of curve 401 within time regions 403 and 404 represent the waveform curve of the pulsating response signal (i.e., the reflected sound signal) of the unit pulse signal from sound source B after reflection and scattering within room model 301, reaching the listener's left ear.
[0071] Continue as Figure 4As shown, curve 405 is the simulated waveform of the BRIR received by the right ear in the time domain. Specifically, the portion of curve 405 within time region 406 represents the waveform of the pulsating response signal (i.e., the direct sound signal) from the unit pulse signal from sound source B directly reaching the right ear. The portions of curve 405 within time regions 407 and 408 represent the waveform of the pulsating response signal (i.e., the reflected sound signal) from the unit pulse signal from sound source B after reflection and scattering within room model 301, reaching the right ear.
[0072] For example, the unit pulse signal emitted by the sound source C can reach the listener not only directly from the sound source C, but also after reflection and scattering within the room model 301.
[0073] Figure 5 The diagram shows the impulse response signal formed when a unit pulse signal from sound source C reaches the listener's ear.
[0074] like Figure 5 As shown, curve 501 is the waveform curve of the BRIR received by the left ear in the time domain. Specifically, the portion of curve 501 within time region 502 represents the waveform curve of the pulsating response signal (i.e., the direct sound signal) of the unit pulse signal from sound source C directly reaching the listener's left ear. The portions of curve 501 within time regions 503 and 504 represent the waveform curve of the pulsating response signal (i.e., the reflected sound signal) of the unit pulse signal from sound source C after reflection and scattering within room model 301, reaching the listener's left ear.
[0075] Continue as Figure 5 As shown, curve 505 is the simulated BRIR waveform received by the right ear in the time domain. Specifically, the portion of curve 505 within time region 506 represents the waveform of the pulsating response signal (i.e., the direct sound signal) from the unit pulse signal from sound source C directly reaching the right ear. The portions of curve 505 within time regions 507 and 508 represent the waveform of the pulsating response signal (i.e., the reflected sound signal) from the unit pulse signal from sound source C after reflection and scattering within room model 301, reaching the right ear.
[0076] like Figure 4 and Figure 5 As shown, within the same room model, the BRIR values differ depending on the positional relationship between the sound source and the listener.
[0077] Ambisonics rendering: a spatial audio method capable of displaying the entire sound field. The three-dimensional spatial audio signal obtained through Ambisonics rendering, also known as the rendered signal, consists of multiple channels, and the number of channels is related to the order of the three-dimensional spatial audio signal.
[0078] Whether using HRIR rendering, BRIR rendering, or Ambisonics rendering, the generated rendering signal will suffer from reduced audio quality.
[0079] To address the aforementioned problems, this application provides an audio data processing method applied to an electronic device. If HRIR rendering is used to process audio data to be played, a target HRIR is obtained before rendering. This target HRIR does not contain target location information but retains timbre information. A corresponding rendering signal is generated using the target HRIR and the audio data to be processed. Then, a pre-configured compensation filter is used to calibrate the timbre of the rendering signal.
[0080] If BRIR or Ambisonics is used to render the audio data to be processed, after rendering is complete, the rendered signal is split into a direct sound signal part and a reflected sound signal part. Different compensation filters are used to calibrate the timbre of the direct sound signal part and the reflected sound signal part respectively.
[0081] For example, the aforementioned electronic device can be a device with its own playback capability, such as a playback device. Examples include headphones, speakers, tablets, laptops, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, etc. In this example, the electronic device can perform the above method to calibrate the rendered signal and play the calibrated rendered signal.
[0082] As another example, the aforementioned electronic device can also be a device connected to a playback device, such as a mobile phone connected to headphones. In this example, the electronic device can perform the above method to calibrate the rendered signal, and the playback device can play the calibrated rendered signal.
[0083] In subsequent embodiments, taking a mobile phone with connected headphones as an example, the implementation details of this method will be described: In some embodiments, multiple compensation filters need to be pre-configured in the electronic device. For example, compensation filter 1, compensation filter 2, and compensation filter 3. Compensation filter 1 is used to calibrate the rendering signal generated by the HRIR rendering method. Compensation filter 2 is used to calibrate the rendering signal generated by the BRIR rendering method. Compensation filter 3 is used to calibrate the rendering signal generated by the Ambisonics rendering method.
[0084] In some embodiments, an electronic device may be configured with multiple compensation filters 1. Different compensation filters 1 correspond to different target distances. The target distance refers to the simulated distance between the sound source and the listener.
[0085] As one implementation method, such as Figure 6 As shown, the method for generating compensation filter 1 is as follows: S101, Obtain at least one set of candidate HRIRs.
[0086] As described in the previous embodiments, HRIR = f(target position information, timbre information), that is, different candidate HRIRs correspond to different target position information and / or different timbre information. A set of candidate HRIRs may include multiple candidate HRIRs (e.g., the first HRIR).
[0087] In some embodiments, the target location information corresponding to the same group of candidate HRIRs is different, while the corresponding timbre information is the same. For example, the target location information of the same group of candidate HRIRs may have the same target distance but different horizontal azimuth and vertical elevation angles. For instance, multiple first HRIRs may be HRIRs simulating a direct path from the first sound source location to the listener.
[0088] like Figure 7 As shown, in an anechoic chamber, the impulse response signal formed when a unit pulse signal from sound source 701 reaches both ears of the listener can be tested, which is the first candidate HRIR. The impulse response signal formed when a unit pulse signal from sound source 702 reaches both ears of the listener can be tested, which is the second candidate HRIR. The impulse response signal formed when a unit pulse signal from sound source 703 reaches both ears of the listener can be tested, which is the third candidate HRIR. The impulse response signal formed when a unit pulse signal from sound source 704 reaches both ears of the listener can be tested, which is the fourth candidate HRIR. The target distance between sound sources 701 to 704 and the listener's head is distance 'a', and the corresponding first to fourth candidate HRIRs form a group of candidate HRIRs.
[0089] Understandably, in an anechoic chamber, it's also possible to test the candidate HRIR for more or fewer sound sources. Figure 7 The number of sound sources tested is only an example and is not intended as a specific limitation.
[0090] In possible embodiments, the method described in the foregoing embodiments can also be used to obtain multiple sets of candidate HRIRs. The target distances corresponding to the target location information of different sets of candidate HRIRs are different.
[0091] Continue as Figure 7 As shown, in an anechoic chamber, the impulse response signals formed when a unit pulse signal from sound source 705 reaches both ears of the listener can be tested, which is the 5th candidate HRIR. The impulse response signals formed when a unit pulse signal from sound source 706 reaches both ears of the listener can be tested, which is the 6th candidate HRIR. The impulse response signals formed when a unit pulse signal from sound source 707 reaches both ears of the listener can be tested, which is the 7th candidate HRIR. The impulse response signals formed when a unit pulse signal from sound source 708 reaches both ears of the listener can be tested, which is the 8th candidate HRIR. The target distance between sound sources 705 to 708 and the listener's head is always distance b. Distances a and b differ, and the corresponding 5th to 8th candidate HRIRs form another set of candidate HRIRs.
[0092] S102, merges the target HRIR corresponding to each group of candidate HRIRs.
[0093] The target HRIR does not contain target location information but retains timbre information. In some embodiments, candidate HRIRs from the same group can be superimposed in the time domain to obtain a target HRIR that retains only timbre information (e.g., the first target HRIR).
[0094] Understandably, candidate HRIRs for the same sound source can include candidate HRIRs for the left ear and candidate HRIRs for the right ear. S102 above can also involve superimposing candidate HRIRs for the left ear within the same group to obtain the corresponding target HRIR for the left ear. Similarly, superimposing candidate HRIRs for the right ear within the same group yields the corresponding target HRIR for the right ear. The aforementioned target HRIRs for the left and right ears are collectively referred to as the target HRIRs corresponding to this group of candidate HRIRs.
[0095] In other embodiments, after superimposing the candidate HRIRs of the same group, an averaging process can be performed. For example, after superimposing the left ear candidate HRIRs from the first to the fourth candidate HRIRs, dividing by 4 yields the corresponding left ear target HRIR. Similarly, after superimposing the right ear candidate HRIRs from the first to the fourth candidate HRIRs, dividing by 4 yields the corresponding right ear target HRIR.
[0096] In addition, different groups of candidate HRIRs correspond to different target HRIRs, that is, different target distances correspond to different target HRIRs.
[0097] S103, transform the target HRIR from the time domain to the frequency domain to obtain the corresponding spectrum information to be processed 1.
[0098] In some embodiments, the implementation details of transforming the target HRIR from the time domain to the frequency domain can be found in relevant technologies, and will not be elaborated here.
[0099] Understandably, the target HRIR includes the left ear target HRIR and the right ear HRIR. Accordingly, S104 above can be: converting the left ear target HRIR and the right ear HRIR from the time domain to the frequency domain respectively to obtain the spectrum information to be processed for the left ear and the spectrum information to be processed for the right ear.
[0100] S104, based on the spectrum information to be processed 1, fit the corresponding compensation filter 1.
[0101] In some embodiments, S104 may be: generating an inverse filter a corresponding to the spectrum information 1 to be processed, converting the inverse filter a to the time domain, and obtaining the corresponding compensation filter 1.
[0102] Understandably, the spectrum information to be processed 1 includes the spectrum information to be processed corresponding to the left ear and the spectrum information to be processed corresponding to the right ear. Accordingly, compensation filter 1 for the left ear and compensation filter 1 for the right ear can be generated according to the method described above.
[0103] In other embodiments, after generating target HRIRs for different target distances, compensation filters 1 corresponding to different target distances can be obtained based on the target HRIRs corresponding to different target distances.
[0104] In some embodiments, the electronic device is configured with a default target distance among multiple target distances, such as the default distance. When using HRIR rendering, the electronic device can process the audio signal to be played based on the target HRIR (excluding target position information) corresponding to the default distance to obtain the corresponding rendered signal. Then, the compensation filter 1 corresponding to the default distance is used to calibrate the timbre of the rendered signal.
[0105] In other possible embodiments, when using HRIR rendering, the electronic device can process the audio signal to be played based on the HRIR corresponding to the default distance (which includes target position information, such as a horizontal azimuth of 0 and a vertical elevation of 0) to obtain the corresponding rendered signal. Then, the timbre of the rendered signal is calibrated using the compensation filter 1 corresponding to the default distance.
[0106] In other embodiments, a user can operate an electronic device to select a target distance from multiple target distances. When using HRIR rendering, in response to detecting that the user has selected a target distance (e.g., the selected distance), the audio signal to be played can be processed according to the target HRIR (excluding target position information) corresponding to the selected distance to obtain the corresponding rendered signal. Then, the timbre of the rendered signal is calibrated using the compensation filter 1 corresponding to the selected distance. Specific implementation details can be found in subsequent embodiments and will not be elaborated here.
[0107] In other possible embodiments, when using HRIR rendering, in response to detecting that the user has selected a target distance (e.g., the selected distance), the audio signal to be played can be processed according to the HRIR corresponding to the selected distance (containing target position information, such as a horizontal azimuth of 0 and a vertical elevation of 0) to obtain the corresponding rendered signal. Then, the timbre of the rendered signal is calibrated using the compensation filter 1 corresponding to the selected distance.
[0108] In some embodiments, the electronic device may be configured with multiple sets of compensation filters 2. Different sets of compensation filters 2 correspond to different target distances. The target distance refers to the simulated distance between the sound source and the listener.
[0109] For example, a set of compensation filters 2 includes a compensation filter a for calibrating the direct sound signal and a compensation filter b for calibrating the reflected sound signal. The compensation filter b for calibrating the reflected sound signal includes a compensation filter c for calibrating the early reflection signal and a compensation filter d for calibrating the late reverberation signal. Compensation filters a, c, and d are different.
[0110] As one implementation method, such as Figure 8 As shown, the method for generating compensation filter 2 is as follows: S201, obtain BRIR.
[0111] In some embodiments, a BRIR corresponding to at least one room model is obtained (e.g., a first BRIR corresponding to a first distance, or a second BRIR corresponding to a second distance). The process of generating the BRIR can be referred to the foregoing embodiments and related technologies, and will not be elaborated here. Furthermore, different room models correspond to different target distances; that is, different target distances correspond to different BRIRs.
[0112] Additionally, a BRIR corresponding to a room model may include a left ear BRIR and a right ear BRIR.
[0113] S202, perform signal partitioning on the BRIR.
[0114] In some embodiments, the left ear BRIR and the right ear BRIR can be segmented separately. Taking the signal segmentation of the left ear BRIR as an example, after signal segmentation, the direct sound waveform, early reflection waveform, and late reverberation waveform corresponding to the left ear BRIR are obtained.
[0115] The direct sound waveform is the portion of the left ear BRIR time-domain waveform that indicates the direct sound signal. The early reflection waveform is the portion of the left ear BRIR time-domain waveform that indicates the early reflection signal. The late reverberation waveform is the portion of the left ear BRIR time-domain waveform that indicates the late reverberation signal.
[0116] As one implementation method, the process of dividing the direct sound waveform from the time-domain waveform of the left ear BRIR is as follows: (1) Obtain the maxpoint in the time domain waveform of the left ear BRIR.
[0117] Specifically, the amplitude of the maxpoint (first sampling point) in the time-domain waveform of the left ear BRIR is greater than the amplitudes of other sampling points in the time-domain waveform of the left ear BRIR. For example, the maxpoint could be... Figure 9 Sampling point 901 in the sample.
[0118] (2) Based on maxpoint, extract the direct sound waveform from the time domain waveform of the left ear BRIR.
[0119] In some embodiments, 2N sampling points are obtained based on maxpoint. These 2N sampling points include maxpoint. Among the 2N sampling points, maxpoint is neither the earliest nor the latest sampling point. Based on the obtained 2N sampling points, direct sound waveforms (third direct sound waveform and fourth direct sound waveform) are extracted from the time-domain waveform of the left ear BRIR.
[0120] For example, based on the acquired 2N sampling points, a combined window consisting of a Hanning window and a box can be used to extract the direct sound waveform from the time-domain waveform of the left ear BRIR. Taking N as 128 as an example, the combined window used to extract the direct sound waveform is shown below: ; window=1,maxpoint<n<maxpoint+128; ; Where n represents the number of sampling points, and window represents the combined window. Continuing with... Figure 9 For example, when the maxpoint is Figure 9Taking sampling point 901 as an example, a combined window 902 for extracting the direct sound waveform is determined based on sampling point 901. The combined window 902 consists of a Hanning window 902-1, a box 902-2, and a Hanning window 902-3. Accordingly, the portion of the time-domain waveform of the left ear BRIR that belongs to the combined window 902 is taken as the direct sound waveform.
[0121] As one implementation method, the process of dividing the late reverberation waveform from the time-domain waveform of the left ear BRIR is as follows: S1 converts the left ear BRIR into the corresponding energy curve.
[0122] In some embodiments, the formula can be used: ETC=20log 10 abs(BRIR(t)) is used to fit the BRIR curve of the left ear to the corresponding energy curve. For example, as... Figure 10 As shown, curve 1001 can be the energy curve derived from the left ear BRIR conversion.
[0123] S2, in the energy curve, obtain beginPoint.
[0124] Here, beginPoint (the second sampling point) can be the starting point of a line segment (target line segment) in the energy curve that conforms to the characteristics of a sloped descent. Understandably, a line segment conforming to the characteristics of a sloped descent refers to a line segment in the energy curve that exhibits a continuous, monotonous, and approximately linear smooth decay pattern. For example... Figure 10 In the diagram, sampling point 1002 is the beginPoint of the energy curve (i.e., curve 1001).
[0125] Furthermore, the left ear BRIR varies depending on the target distance, and consequently, the corresponding energy curves also differ. For example, in Figure 11 In the diagram, when the target distance is 1m, the corresponding energy curve is 1003. When the target distance is 2m, the corresponding energy curve is 1004. Energy curves 1003 and 1004 are different, and their corresponding beginPoints are different. For example, the beginPoint of energy curve 1003 is sampling point 1005. For example, the beginPoint of energy curve 1004 is sampling point 1006.
[0126] S3, based on beginPoint, extract the late reverberation waveform from the time-domain waveform of the left ear BRIR.
[0127] In some embodiments, the formula can be used based on beginPoint: ; Define a Hanning window for extracting the late reverberation waveform, such as Figure 10Hanning Window 1007-1 in the middle.
[0128] In some embodiments, the formula can also be used based on beginPoint: window=1, n>beginPoint; Define a box for truncating the late reverberation waveform, such as... Figure 10 Box 1007-2 in the middle.
[0129] The Hanning window 1007-1 and the box 1007-2 described above constitute a combined window 1007 for extracting the late reverberation waveform. Then, using the combined window 1007, the late reverberation waveform (second late reverberation waveform) is extracted from the time-domain waveform of the left ear BRIR.
[0130] One implementation involves dividing the early reflection waveform (second early reflection waveform) from the time-domain waveform of the left ear BRIR, including: dividing the waveform between the direct sound waveform and the late reverberation waveform in the time-domain waveform of the left ear BRIR into early reflection waveforms. For example, the early reflection waveform partially overlaps with both the direct sound waveform and the late reverberation waveform.
[0131] For example, a formula can also be used: ; window=1,maxpoint+192<n<beginPoint-64; ; Determine the combined window for the user to capture the early reflection waveform, such as Figure 12 As shown, combined with combined windows 902 and 1007, combined window 1008 is defined. Combined window 1008 consists of Hanning window 1008-1, box 1008-2, and Hanning window 1008-3. Then, using combined window 1008, the early reflection waveform is extracted from the time-domain waveform of the left ear BRIR.
[0132] S203, determine whether the energy percentage of the direct sound waveform exceeds the target percentage value.
[0133] In some embodiments, the total energy value corresponding to the left ear BRIR can be calculated based on the amplitudes corresponding to all sampling points in the left ear BRIR. For example, the total energy value can be calculated based on the sum of squares of the amplitudes of all sampling points. Then, the energy value corresponding to the direct sound waveform can be calculated based on the amplitudes of the sampling points in the direct sound waveform. For example, the energy value can be calculated based on the sum of squares of the amplitudes of the sampling points belonging to the direct sound waveform. In this way, the energy proportion of the direct sound waveform can be obtained based on the energy value corresponding to the direct sound waveform and the total energy value of the left ear BRIR.
[0134] In some embodiments, the target ratio value (first ratio threshold) can be a pre-configured empirical value, such as 75%, and in other examples, it can be configured to other values greater than 50%.
[0135] In some embodiments, when the energy percentage of the direct acoustic waveform exceeds a target percentage, the process proceeds to S204, and the process ends after S204 is completed. When the energy percentage of the direct acoustic waveform does not exceed the target percentage, the process proceeds to S205.
[0136] In a possible embodiment, S203 and S204 may be skipped, and S205 may be executed after S202.
[0137] Understandably, after dividing other BRIRs, S203 can also be executed based on these other BRIRs. For example, other BRIRs could be the right ear BRIRs corresponding to the same target distance. Alternatively, other BRIRs could be BRIRs corresponding to different target distances.
[0138] S204, Based on the direct acoustic waveform, generate a compensation filter a for calibrating the direct acoustic waveform.
[0139] In some embodiments, the implementation details of generating compensation filter a can be found in the foregoing embodiments for generating compensation filter 1, and will not be repeated here.
[0140] S205 generates compensation filters a, c, and d based on the direct sound waveform, early reflection waveform, and late reverberation waveform, respectively.
[0141] In some embodiments, the implementation details of generating compensation filter a can be found in the foregoing embodiments for generating compensation filter 1, and will not be repeated here.
[0142] In some embodiments, the compensation filter c is generated as follows: 1) Transform the early reflection waveform from the time domain to the frequency domain to obtain the corresponding spectrum information to be processed 2 (first spectrum information).
[0143] 2) Perform 1 / 4 octave band smoothing on the spectrum information to be processed to remove extreme peaks and troughs. For example, Figure 13 Curve 1101 in the diagram can be the spectral curve of the spectrum information to be processed 2. After the spectrum information to be processed 2 undergoes 1 / 4 octave band smoothing, the spectrum information to be processed 3 is obtained, as shown below. Figure 13 Curve 1102 is shown in the figure. Compared to curve 1101, curve 1102 is smoother.
[0144] 3) Based on the spectrum information 3 to be processed, fit the corresponding compensation filter c. For example, an inverse filter 2 (second spectrum information) corresponding to the spectrum information 3 to be processed can be generated, and the inverse filter 2 can be transformed to the time domain to obtain the corresponding compensation filter c.
[0145] Understandably, inverse filter 2 includes multiple intensity values corresponding to multiple frequency points. The intensity values in inverse filter 2 can be called calibration values, which are used to calibrate the intensity values of frequency points in the spectrum information 3 to be processed. For example, frequency point a corresponds to calibration value a in inverse filter 2. Thus, calibration value a can be used to calibrate the intensity value a of frequency point a in the spectrum information 3 to be processed.
[0146] In some examples, if frequency point a (first frequency point, second frequency point) belongs to preset frequency band 1 (first frequency band), the calibration value corresponding to frequency point a has the same absolute value but opposite phase as the intensity value a corresponding to frequency point a. Thus, after calibration, the intensity value of frequency point a in the spectrum information 3 to be processed is 0 dB. For example, preset frequency band 1 can be an empirical value, such as a frequency band less than 3000 Hz.
[0147] In some examples, if frequency point a (the third frequency point, the fourth frequency point) belongs to preset frequency band 2 (the second frequency band), the calibration value a corresponding to frequency point a does not exceed a preset upper limit intensity value (e.g., 5dB). The calibration value a is out of phase with the intensity value a. By limiting the calibration degree of the frequency points in preset frequency band 2, the resulting compensation filter c can preserve the spatial characteristics of the audio and ensure a natural listening experience. The frequency points in preset frequency band 2 are higher than those in preset frequency band 1.
[0148] Thus, if the absolute value of intensity value 'a' is less than or equal to the preset upper limit intensity value (also known as the preset intensity threshold), after calibration, the intensity value of frequency point 'a' in the spectrum information 3 to be processed is 0 dB. If the absolute value of calibration value 'a' is greater than the preset upper limit intensity value, after calibration, the intensity value of frequency point 'a' in the spectrum information 3 to be processed is: intensity value a - 5 dB. For example, the preset frequency band 2 can be an empirical value, such as a frequency band not less than 3000 Hz and not greater than 10000 Hz.
[0149] For example, based on the spectrum information 3 to be processed, the generated inverse filter 2 can be as follows: Figure 13 As shown in curve 1103. The frequency points corresponding to line segments 1104 in curve 1102 and 1105 in curve 1103 both belong to the preset frequency band 2. If the intensity value corresponding to line segment 1104 is greater than 5dB, then the calibration value corresponding to line segment 1105 is 5dB.
[0150] In some examples, if frequency point a belongs to preset frequency band 3, frequency point a does not correspond to a calibration value; that is, the intensity value of frequency point a in the spectrum information 3 to be processed is not calibrated. For example, the frequency points in preset frequency band 3 are higher than the frequency points in preset frequency band 2. Preset frequency band 3 can be an empirical value, such as a band greater than 10000Hz.
[0151] In other possible embodiments, the calibration value corresponding to any frequency point in the inverse filter 2 is the same as the absolute value of the intensity value of that frequency point in the quasi-processable spectrum information 3, but opposite in phase.
[0152] Thus, when processing the rendering signal using the compensation filter c, the amplitude of the preset frequency band 2 is calibrated to a limited extent, while the amplitude of the preset frequency band 3 is not calibrated.
[0153] In some embodiments, the method for generating compensation filter d can refer to the method for generating compensation filter c, which will not be elaborated here.
[0154] Thus, in practical applications, when electronic devices select different target distances, the rendered spatial audio can be calibrated according to the compensation filter 2 corresponding to that target distance. For example, if the selected target distance corresponds only to compensation filter a, then only the direct sound waveform of the rendered spatial audio is calibrated. Or, if the selected target distance corresponds to compensation filters a, c, and d, then the direct sound waveform, early reflection waveform, and late reverberation model of the rendered spatial audio are all calibrated.
[0155] In some embodiments, the electronic device may be configured with multiple sets of compensation filters 3. Different sets of compensation filters 3 correspond to different target distances. The target distance refers to the simulated distance between the sound source and the listener.
[0156] For example, a set of compensation filters 3 includes a compensation filter e for calibrating the direct sound signal and a compensation filter f for calibrating the reflected sound signal. The compensation filter for calibrating the reflected sound signal includes a compensation filter g for calibrating the early reflection signal and a compensation filter h for calibrating the late reverberation signal.
[0157] As one implementation method, such as Figure 14 As shown, the method for generating compensation filter 3 is as follows: S301, select multiple evenly distributed spherical points on the sphere.
[0158] In some embodiments, such as Figure 15As shown, multiple spherical points, such as spherical point 1301 and spherical point 1302, are uniformly selected on the sphere. Spherical point 1301 and spherical point 1302 are equidistant from the corresponding center point of the sphere. For example, the aforementioned multiple spherical points can be 492 spherical points.
[0159] S302 uses Ambisonics to render the BRIR corresponding to each spherical point.
[0160] In some embodiments, such as Figure 16 As shown, the implementation details of S302 above are as follows: (1) Obtain the FOA (or HOA) impulse response signal corresponding to each spherical point.
[0161] For example, the FOA (or HOA) impulse response information corresponding to each spherical point can be tested by simulating a real-world scenario. For specific implementation details, please refer to relevant technologies, which will not be elaborated here.
[0162] (2) The original audio signal (e.g., unit pulse signal) is encoded using Ambisonics to obtain the corresponding encoded signal.
[0163] (3) Convolution processing is performed on the encoded signal and the FOA (or HOA) impulse response signal to obtain the FOA (or HOA) signal containing spatial characteristics.
[0164] (4) Perform Ambisonics decoding on the FOA (or HOA) signal to obtain the dual-channel time-domain signal corresponding to the spherical point. For specific implementation details, please refer to relevant technologies, which will not be elaborated here.
[0165] (5) The BRIR of the spherical point is obtained by convolution processing based on the time-domain signal of the dual-channel and the head correlation transfer function.
[0166] The dual-channel time-domain signal includes a time-domain signal for the left ear and a time-domain signal for the right ear. Correspondingly, the BRIR of this spherical point includes the left ear BRIR and the right ear BRIR. That is, the left ear BRIR and right ear BRIR can be obtained by convolving the time-domain signals for the left ear and the right ear with the head correlation transfer function, respectively.
[0167] S303, fuse the BRIRs corresponding to multiple spherical points to obtain the target BRIR.
[0168] In some embodiments, S303 described above may be the fusion of a left ear BRIR with multiple spherical points and the fusion of a right ear BRIR with multiple spherical points, respectively. Furthermore, the implementation method for fusing multiple BRIRs can be found in related technologies and will not be elaborated here. For example, linear superposition can be performed, and the time-domain waveform of the target BRIR obtained after superposition can be found in the aforementioned embodiments. Figure 4 The waveform shown will not be described again here.
[0169] S304 performs signal segmentation on the target BRIR.
[0170] In some embodiments, the implementation details of S304 can be referred to S202 in the foregoing embodiments, and will not be repeated here.
[0171] S305, based on the partitioning results, generates compensation filter e, compensation filter g, and compensation filter h.
[0172] In some embodiments, the process of generating compensation filter e can refer to the process of generating compensation filter a, and will not be described in detail here. The process of generating compensation filter g can refer to the process of generating compensation filter c, and will not be described in detail here. The process of generating compensation filter h can refer to the process of generating compensation filter d, and will not be described in detail here.
[0173] In other embodiments, different spheres can be established with different target distances as radii, and S301~S305 can be executed based on each sphere to generate compensation filters 3 corresponding to different target distances.
[0174] Understandably, whether using HRIR, BRIR, or Ambisonics to render audio, different target distances are selected, and different compensation filters are used to calibrate the rendered audio signal after rendering.
[0175] In an exemplary scenario, an electronic device may display an audio configuration interface for changing the horizontal azimuth angle (also known as the horizontal angle), the vertical elevation angle (also known as the pitch angle), and the target distance.
[0176] For example, the audio configuration interface could be the system configuration interface. Figure 17 The audio configuration interfaces 1701 and 1711 are shown.
[0177] The audio configuration interface 1701 includes a distance adjustment bar 1702. A slider 1703 is included on the distance adjustment bar 1702. In response to the user dragging the slider 1703, the position of the slider 1703 relative to the distance adjustment bar 1702 can be adjusted. Different positions of the slider 1703 on the distance adjustment bar 1702 indicate different target distances. For example, at position 1704, the target distance is set to 1m. At position 1705, the target distance is set to 3m.
[0178] Different target distances indicate different distance values between the simulated sound source and the listener, also known as sound source distances. At different sound source distances, the audio effects rendered and calibrated using the method provided in this application's embodiments differ.
[0179] As described in the foregoing embodiments, the electronic device is configured with a default distance, such as 1m. Accordingly, the sound source distance (target distance) displayed in the audio configuration interface 1701 is 1m. Additionally, default horizontal azimuth angle, vertical elevation angle, etc., may also be included.
[0180] Without changing the horizontal azimuth (also known as the horizontal angle), vertical elevation (also known as the pitch angle), and target distance (sound source distance), the electronic device defaults to using the impulse response (e.g., HRIR or BRIR) corresponding to 1m to render the audio data. Then, a supplementary filter corresponding to that 1m is used to calibrate the timbre of the rendered signal. Alternatively, after rendering the audio data using Ambisonics, the supplementary filter corresponding to that 1m is used to calibrate the timbre of the rendered signal.
[0181] Taking an electronic device using HRIR rendering as an example, the electronic device displays a music playback interface in response to detecting an operation to open a music application. During the display of the music playback interface, in response to detecting an operation instructing the playback of music, it acquires the audio data 'a' of the music to be played. Before playing the audio data 'a' (the second raw audio), the electronic device can determine whether it is connected to headphones. If headphones are connected, it acquires the target HRIR (or the HRIR corresponding to 1m). The audio data 'a' is rendered using the target HRIR (or the HRIR corresponding to 1m) to obtain the corresponding rendered signal 'a' (third spatial audio, fourth spatial audio). Then, the sound quality of the rendered signal 'a' is calibrated using compensation filters 1 (fifth filter, sixth filter) corresponding to 1m. Finally, the calibrated rendered signal 'a' is played through the headphones.
[0182] Taking an electronic device using BRIR rendering as an example, the electronic device displays a music playback interface in response to detecting an operation to open a music application. During the display of the music playback interface, in response to detecting an operation to indicate that music should be played, the electronic device acquires the audio data 'a' of the music to be played. Before playing the audio data 'a', the electronic device can determine whether it is connected to headphones. If the electronic device is connected to headphones, then it acquires the BRIR corresponding to 1m. If the energy percentage of the direct sound waveform in the BRIR corresponding to 1m (first distance) is greater than 75%, after rendering the audio data 'a' (first original audio) using the BRIR at 1m, a rendered signal 'c' (first spatial audio) is obtained. The timbre of the first direct sound waveform in the rendered signal 'c' is calibrated using the compensation filter 'a' (first filter) corresponding to 1m, and then the calibrated rendered signal 'c' is played through headphones.
[0183] If the energy proportion of the direct sound waveform in the BRIR corresponding to 1m (the second distance) is not greater than 75%, such as Figure 18 As shown, the electronic device can separate the direct sound waveform (second direct sound waveform), the early reflection waveform (first early reflection waveform), and the late reverberation waveform (first early reflection waveform) from the rendered signal c (second spatial audio). The implementation details of separating the rendered signal c can be found in S202 of the aforementioned embodiment, and will not be repeated here. Then, the direct sound waveform in the rendered signal c is calibrated using compensation filter a (second filter) corresponding to 1m. The early reflection waveform in the rendered signal c is calibrated using compensation filter c (third filter) corresponding to 1m. The late reverberation waveform in the rendered signal c is calibrated using compensation filter d (fourth filter) corresponding to 1m. Finally, the calibrated direct sound waveform, early reflection waveform, and late reverberation waveform are superimposed to obtain the output audio, which is then played by headphones.
[0184] Furthermore, for electronic devices using Ambisonics rendering, after rendering audio data a using Ambisonics to obtain the rendered signal e, the rendered signal e is calibrated using compensation filter 3, and then the calibrated rendered signal e is played through headphones. The implementation details of calibrating the rendered signal e using compensation filter 3 (e.g., compensation filter e, compensation filter g, and compensation filter h) can be found in the implementation details of calibrating the rendered signal c using compensation filter 2 (e.g., compensation filter a, compensation filter c, and compensation filter d), and will not be elaborated upon here.
[0185] In some embodiments, during the display of the audio configuration interface, the electronic device can change the horizontal azimuth angle (also known as the horizontal angle), the vertical elevation angle (also known as the pitch angle), and the target distance according to the user's operation.
[0186] Continue as Figure 17As shown, the audio configuration interface 1701 includes a manual mode 1709 for manual configuration and an automatic mode 1710 for automatic configuration. In the audio configuration interface 1701, manual mode 1709 is selected.
[0187] When the audio configuration interface 1701 is displayed, in response to the user dragging the slider 1703, the slider 1703 is moved relative to the distance adjustment bar 1702. When it moves to position 1705, the target position is modified to 3m, and correspondingly, the sound source distance displayed in the audio configuration interface 1701 is 3m. Then, the audio data can be rendered using the impulse response (e.g., HRIR or BRIR) corresponding to 3m. Then, the timbre of the rendered signal is calibrated using the supplementary filter corresponding to 3m. Alternatively, after rendering the audio data using Ambisonics, the timbre of the rendered signal is calibrated using the supplementary filter corresponding to 3m.
[0188] Furthermore, the aforementioned embodiments primarily use a horizontal azimuth of 0° and a vertical elevation of 0° as examples. In practical applications, both the horizontal azimuth and vertical elevation can be adjusted. Continuing... Figure 17 As shown, the audio configuration interface 1701 also includes adjustment lines 1706 and 1707. Users can change the horizontal azimuth and / or vertical elevation angle of the sound source 1708 relative to the listener by dragging adjustment lines 1706 and 1707.
[0189] Continue as Figure 17 As shown, when the audio configuration interface 1701 is displayed, in response to an operation on the automatic mode 1710, such as a click, the audio configuration interface 1711 is switched to be displayed. In the audio configuration interface 1711, the manual mode 1709 is unselected, and the automatic mode 1710 is selected. During the display of the audio configuration interface 1711, the electronic device can automatically adjust the corresponding horizontal azimuth and vertical elevation angles according to changes in the user's head position. Additionally, the audio configuration interface 1711 also includes a distance adjustment bar 1712. The user can also adjust the target distance by operating on the distance adjustment bar 1712.
[0190] In some embodiments, such as Figure 17 As shown, both audio configuration interfaces 1701 and 1711 include an enable switch 1713 for spatial audio functionality. In response to an operation on the enable switch 1713, such as clicking it, the spatial audio function can be disabled, that is, audio rendering and tone calibration can be stopped, and the audio configuration interface 1714 will be displayed, indicating that the spatial audio function is disabled.
[0191] For example, the audio configuration interface can be an application interface provided by an application with playback functionality. Taking a music application as an example... Figure 19 As shown, when a music application is running in the foreground of an electronic device, the audio configuration interface 1801 of the music application can be displayed. The audio configuration interface 1801 includes a music playback icon 1802. When music 1 is playing, the music playback icon 1802 can display information indicating that music 1 is being played. During the display of the audio configuration interface 1801, the electronic device can respond to user operations by adjusting the source distance, horizontal angle, and pitch angle for music 1. For example, after adjustment, the source distance is 3m, and both the horizontal angle and pitch angle are 0°. When actually playing music 1, the audio data of music 1 is rendered using the impulse response corresponding to a source distance of 3m, a horizontal angle of 0°, and a pitch angle of 0°. Then, the timbre of the rendered signal is calibrated using a supplementary filter corresponding to 3m. Alternatively, after rendering the audio data using Ambisonics, the timbre of the rendered signal is calibrated using a supplementary filter corresponding to 3m.
[0192] Continue as Figure 19 As shown, in response to an operation on the music playback icon 1802, such as a click, a song list 1803 pops up. In the song list 1803, the playback control 1804 corresponding to music 1 is in a playing state, while the playback controls for other music are in a non-playing state; for example, the playback control 1805 corresponding to music 2 is in a non-playing state. Figure 19 As shown, in response to detecting an operation on the playback control 1805, the playback of music 2 is switched, and the audio configuration interface 1806 for music 2 is switched to be displayed.
[0193] In this case, the sound source distance, horizontal angle, and pitch angle for Music 2 are all set to default values. For example, in the audio configuration interface 1806, the sound source distance is displayed as 1m.
[0194] In some embodiments, the song list 1803 is still displayed on the audio configuration interface 1806. The song list 1803 can be de-displayed in response to a user clicking on an area outside of the song list 1803.
[0195] The foregoing embodiments mainly describe a scenario where an electronic device is connected to headphones. The following describes the implementation details of enabling the electronic device to play audio data through its built-in speaker: In some embodiments, the electronic device determines whether it is connected to headphones before playing audio data. If it is determined that the electronic device is not connected to headphones, a crosstalk cancellation algorithm is used to process the audio data to be played. For example, a crosstalk cancellation filter is used to process the audio data to be played.
[0196] like Figure 20As shown, the electronic device has a built-in left channel speaker (LS1) and a right channel speaker (LS2). The left channel audio (D1) is filtered by a pass-through path filter (C). 11 ) and cross-path filter (C 21 Processing is performed, such as convolution. The right channel audio (D2) is processed by a pass-through filter (C... 22 ) and cross-path filter (C 12 Processing is performed, such as convolution. After C... 11 D1 after processing and after C 12 After processing and superimposing D2, the audio to be played (1) is obtained. This is then processed by C. 22 D2 after processing and after C 21 After processing D1 and superimposing it, we obtain the audio to be played, 2. Where, C 11 C 12 C 21 and C 22 Both can be called crosstalk cancellation filters. Among them, D1 and D2 can be called the third original audio.
[0197] like Figure 20 As shown, the HRIR (High-Resolution Earth Detection) signal, delivered directly to the listener's left ear by the left channel speaker, is... L1 Render the audio to be played (Audio 1), obtain the left channel rendered audio 1, and use the left channel speaker to achieve HRIR (High-Resolution Interval) in the listener's right ear via crosstalk. R1 Render the audio to be played (audio 1) to obtain the left channel rendered audio 2. Then, play the left channel rendered audio 1 and left channel rendered audio 2 through the left channel speaker.
[0198] HRIR (H for short) is delivered directly to the listener's right ear via the right channel speaker. R2 Render the audio to be played, obtain the right channel rendered audio, and use the right channel speaker to achieve HRIR (High-Resolution Interval) in the listener's left ear via crosstalk. L2 Render the audio to be played, resulting in right channel rendered audio 2. Then, play the right channel rendered audio 1 and right channel rendered audio 2 through the right channel speaker. The left channel rendered audio 1, left channel rendered audio 2, right channel rendered audio 1, and right channel rendered audio 2 can be referred to as the fifth space audio.
[0199] In some embodiments, an electronic device may be configured with a crosstalk cancellation filter capable of compensating for timbre, thereby reducing timbre loss in subsequently rendered audio data. Exemplarily, the steps for generating this crosstalk cancellation filter are as follows: A1, obtain multiple HRIRs.
[0200] In some embodiments, the multiple HRIRs can be referred to as second HRIRs. These multiple second HRIRs are response signals formed when a unit pulse signal travels directly from multiple second sound source locations to the listener. The multiple second sound source locations are different points at the same distance from the listener. At the same target distance, multiple HRIRs corresponding to the electronic device are simulated respectively. L1 Multiple H L2 Multiple H R1 and multiple H R2 With multiple H L1 For example, multiple H L1 H can be measured at different locations where the electronic device is at the same distance from the listener. L1 The candidate HRIRs mentioned in the previous embodiments can be referred to, and the others are similar and will not be repeated here.
[0201] A2, merge multiple HRIRs to obtain the corresponding merged HRIR (second target HRIR).
[0202] In some embodiments, multiple H are fused respectively. L1 Multiple H L2 Multiple H R1 and multiple H R2 H L1 The corresponding fusion of HRIR yields H L2 The corresponding fusion of HRIR yields H R1 The corresponding fusion of HRIR and the resulting H R2 For details on the implementation of the corresponding fused HRIR, please refer to the implementation details of the fused target HRIR in the previous embodiments, which will not be repeated here.
[0203] A3 generates the corresponding inverse filter (seventh filter) based on the fused HRIR.
[0204] In some embodiments, H is respectively L1 H L2 H R1 and H R2 Inverting the fused HRIR yields H. L1 H L2 H R1 and H R2 The corresponding inverse filter.
[0205] A4 uses the inverse filter of HRIR to perform convolution processing on HRIR to obtain compensated HRIR.
[0206] In some embodiments, H L1 inverse filter and H L1 Convolution processing is performed to obtain the compensated HRIR. Among them, H participates in the convolution... L1H can be a pitch angle of 0 and a horizontal angle of 0. L1 Similarly, H can be generated. L2 H R1 and H R2 HRIR compensation.
[0207] A5 generates a crosstalk cancellation filter based on the compensated HRIR.
[0208] Understandably, the relationship between HRIR and crosstalk cancellation filters is as follows: ; Wherein, H is composed of H L1 H L2 H R1 and H R2 The matrix formed, C is C 11 C 12 C 21 and C 22 The matrix composed of crosstalk cancellation filters, where I is the identity matrix, can also be expressed as the following formula: ; In some embodiments, the optimal solution to the above formula is evaluated using the minimum variance, i.e., as shown in the following formula: ; in, A crosstalk cancellation filter that can calibrate timbre loss can also be called a compensated crosstalk cancellation filter. To compensate HRIR, Here, M is the regularization factor, and M is the model time delay. Angular frequency, for The conjugate transpose of .
[0209] This application also provides an electronic device that may include a memory and one or more processors. The memory and processors are coupled. The memory stores computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device can perform the steps described in the above embodiments.
[0210] Figure 21The diagram illustrates the hardware structure of electronic device 100. Electronic device 100 may include a processor 111, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, a camera 193, a display screen 194, an audio module 170, etc. The sensor module 180 may include a pressure sensor 180A, a touch sensor 180K, etc.; and the audio module 170 may include a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, etc.
[0211] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0212] Processor 111 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0213] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0214] In some embodiments, the processor 111 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, and / or a universal serial bus (USB) interface, etc.
[0215] The MIPI interface can be used to connect the processor 111 to peripheral devices such as the display 194 and the camera 193. The MIPI interface includes the camera serial interface (CSI) and the display serial interface (DSI).
[0216] In some embodiments, the processor 111 and the camera 193 communicate via a CSI interface to enable the electronic device 100 to capture images. The processor 111 and the display screen 194 communicate via a DSI interface to enable the electronic device 100 to display images.
[0217] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0218] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0219] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 111 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0220] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or more display screens 194. Pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals.
[0221] In some embodiments, pressure sensor 180A may be disposed on display screen 194. Pressure sensor 180A can be of many types, such as resistive pressure sensor, inductive pressure sensor, capacitive pressure sensor, etc. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 may also calculate the touch position based on the detection signal from pressure sensor 180A.
[0222] In some embodiments, touch operations applied to the same touch location but with different touch intensity can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS message is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS message is executed.
[0223] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.
[0224] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 111 through the external memory interface 120 to realize data storage functions. For example, music, video, and other files can be saved on the external memory card. The internal memory 121 can be used to store computer executable program code, which includes instructions. The processor 111 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0225] This application also provides a chip system that can be applied to the electronic devices described in the foregoing embodiments. The chip system includes at least one processor and at least one interface circuit. The processor may be the processor in the aforementioned electronic device. The processor and the interface circuit are interconnected via wiring. The processor can receive and execute computer instructions from the memory of the aforementioned electronic device through the interface circuit. When the computer instructions are executed by the processor, the electronic device can perform the various steps in the foregoing embodiments. Of course, the chip system may also include other discrete devices, and this application does not specifically limit this.
[0226] In some embodiments, as described above, those skilled in the art will clearly understand that, for the sake of convenience and brevity, the division of the functional modules described above is merely an example. In practical applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0227] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0228] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0229] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. An audio data processing method, characterized in that, Applied to electronic devices, the method includes: When the simulated sound source and the listener are at a first distance, in response to detecting that the electronic device is connected to headphones, a first spatial audio is rendered based on the first original audio. Using the first filter, the timbre of the first direct sound waveform in the first spatial audio is calibrated. When the simulated sound source and the listener are at a second distance, in response to detecting that the electronic device is connected to headphones, a second spatial audio is rendered based on the first original audio. The second filter is used to calibrate the timbre of the second direct sound waveform in the second spatial audio. The timbre of the first early reflection waveform in the second spatial audio is calibrated using a third filter. The timbre of the first late reverberation waveform in the second spatial audio is calibrated using a fourth filter.
2. The method according to claim 1, characterized in that, The step of rendering the first spatial audio based on the first original audio includes: The first original audio is rendered using binaural room impulse response (BRIR) or Ambisonics.
3. The method according to claim 2, characterized in that, The method further includes: Switch to enable Head-Related Impulse Response (HRIR) audio rendering; When the simulated sound source and the listener are at the first distance, in response to detecting that the electronic device is connected to headphones, a third spatial audio is rendered based on the second original audio. The third spatial audio is calibrated using the fifth filter; The calibrated third-space audio is played through the headphones; When the simulated sound source and the listener are at the second distance, in response to detecting that the electronic device has been connected to headphones, a fourth spatial audio is rendered based on the second original audio. The timbre of the fourth spatial audio is calibrated using the sixth filter; The calibrated fourth-space audio is played through the headphones.
4. The method according to claim 3, characterized in that, The method further includes: Multiple first HRIRs are acquired, wherein the multiple first HRIRs are response signals formed when a unit pulse signal directly reaches the listener from multiple first sound source locations; the multiple first sound source locations are different locations at the same distance from the listener. Based on the multiple first HRIRs, a first target HRIR is fused together; The fifth filter is generated based on the first target HRIR.
5. The method according to claim 1, characterized in that, After rendering the second spatial audio, the method further includes: In the second spatial audio, a second direct sound waveform, a first early reflection waveform, and a first late reverberation waveform are defined. The second direct sound waveform includes a first sampling point. In the second spatial audio, the amplitude corresponding to the first sampling point is greater than the amplitude corresponding to other sampling points. The first late reverberation waveform includes a second sampling point. The second sampling point is the starting point of the target line segment in the energy curve corresponding to the second spatial audio. The target line segment is a line segment that conforms to the characteristics of a ramp-like descent. The first early reflection waveform includes the waveform in the second spatial audio located between the second direct sound waveform and the first late reverberation waveform.
6. The method according to claim 1, characterized in that, The method further includes: Obtain the first BRIR corresponding to the first distance; The first BRIR is divided into a third direct sound waveform, a second early reflection waveform, and a second late reverberation waveform; The second filter is generated based on the third direct acoustic waveform; The third filter is generated based on the second early reflection waveform; The fourth filter is generated based on the second late reverberation waveform.
7. The method according to claim 6, characterized in that, The step of generating the third filter based on the second early reflection waveform includes: The first spectral information of the second early reflection waveform is smoothed. The smoothed first spectral information is inverted to obtain the corresponding second spectral information; The corresponding third filter is generated based on the second spectrum information.
8. The method according to claim 7, characterized in that, The first frequency point in the second spectrum information has the opposite phase to the second frequency point in the first spectrum information, and the corresponding absolute intensity values are the same; wherein, the first frequency point and the second frequency point are the same and both belong to the first frequency band; The third frequency point in the second spectrum information is out of phase with the fourth frequency point in the first spectrum information. The absolute intensity value corresponding to the fourth frequency point is greater than the absolute intensity value corresponding to the third frequency point, and the difference is not greater than a preset intensity threshold. The third frequency point and the fourth frequency point are the same and both belong to the second frequency band.
9. The method according to claim 1, characterized in that, The method further includes: In response to detecting that the electronic device is not connected to the headphones, a fifth spatial audio is rendered based on the third original audio and a compensated crosstalk cancellation filter; The fifth-space audio is played through the speaker built into the electronic device.
10. The method according to claim 9, characterized in that, The method further includes: Multiple second HRIRs are acquired, wherein the multiple second HRIRs are response signals formed by a unit pulse signal reaching the listener directly from multiple second sound source locations; the multiple second sound source locations are different location points at the same distance from the listener. Based on the multiple second HRIRs, a second target HRIR is fused together; Invert the second target HRIR to generate the corresponding seventh filter; The seventh filter and a second HRIR are convolved to generate a compensated HRIR. The compensated crosstalk cancellation filter is evaluated based on the compensated HRIR.
11. The method according to claim 6, characterized in that, Before performing timbre calibration on the first direct sound waveform in the first spatial audio using the first filter, the method further includes: It is determined that the energy percentage corresponding to the third direct sound waveform is greater than the first proportional threshold.
12. The method according to claim 1, characterized in that, Before performing timbre calibration on the second direct sound waveform in the second spatial audio using the second filter, the method further includes: Obtain the second BRIR corresponding to the second distance; It is determined that the energy percentage corresponding to the fourth direct acoustic waveform in the second BRIR is not greater than the first proportional threshold.
13. An electronic device, characterized in that, The electronic device includes: a memory and one or more processors; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-12.
14. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-12.
15. A computer program product, characterized in that, The computer program product includes a computer program or instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-12.