Audio Processing Method and Electronic Device
By generating and splicing HRTF and BRIR of head-moving tracking data, the problem of high computing complexity in the prior art is solved, and the fast audio rendering and accurate sound image positioning of headphone devices during head-moving tracking is realized, which improves the real-time response effect of audio playback.
Patent Information
- Application Number
- CN202211065962.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-09-01
AI Technical Summary
The existing spatial audio rendering method has high computational complexity when the user's head position or orientation changes, resulting in slow real-time rendering of electronic devices, affecting the real-time response effect of audio playback of headphone devices.
By acquiring the first and second header tracking data, a third HRTF is generated, and splicing it with the first BRIR into a second BRIR, the second BRIR is used to render the audio signal to be rendered, and a target audio signal is generated only by one convolution operation.
It reduces the computational complexity and memory usage, improves the speed of electronic devices to render target audio signals in real time, improves the real-time audio playback effect of headphone devices during head movement tracking, meets personalized needs and improves the accuracy of audio-visual positioning.
Smart Images

Figure CN117676002B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technologies, and in particular, to an audio processing method and an electronic device. Background Art
[0002] With the rapid development of electronic devices, users have put forward higher and higher requirements for audio experience. For example, when an electronic device establishes a communication connection with a headphone device and the user wears the headphone device to listen to audio, in order to improve the realism and presence of the sound heard by the user, the audio signal to be rendered can be subjected to spatial audio rendering, and the headphone device is used to play the audio signal after spatial audio rendering.
[0003] Currently, an electronic device can use a head related transfer function (HRTF) or a binaural room impulse response (BRIR) to perform spatial audio rendering on an audio signal to be rendered, so that a better spatial effect can be generated when the rendered audio signal is played through the headphone device.
[0004] However, when the user wears the headphone device to listen to audio, if the user's head position or orientation changes, in the process of real-time rendering of the audio signal based on the head movement tracking data, the current spatial audio rendering method has a high computational complexity, which slows down the real-time rendering speed of the electronic device for the audio signal, resulting in a poor real-time response effect when the headphone device plays audio. Summary of the Invention
[0005] An embodiment of this application provides an audio processing method and an electronic device, which reduce the computational complexity and improve the real-time rendering speed of the electronic device for the audio signal in the process of real-time rendering of the audio signal based on the head movement tracking data, so as to improve the real-time response effect when the headphone device plays audio.
[0006] In a first aspect, an embodiment of this application proposes an audio processing method. An electronic device obtains first head movement tracking data and second head movement tracking data, where the second head movement tracking data is the head movement tracking data obtained in the previous time of the first head movement tracking data; the electronic device obtains a first HRTF according to the first head movement tracking data, and obtains a second HRTF according to the second head movement tracking data; when the first head movement tracking data is different from the second head movement tracking data, the electronic device generates a third HRTF according to the first HRTF and the second HRTF; the electronic device splices the third HRTF with a first BRIR to obtain a second BRIR; the electronic device uses the second BRIR to render the audio signal to be rendered to obtain a target audio signal.
[0007] In this way, the target audio signal can be generated through a single convolution operation, thereby reducing the computational complexity and the memory occupied by the computation, improving the speed at which the electronic device renders the target audio signal in real time, and thus reducing the link delay of sending the target audio signal between the electronic device and the headphone device. In this way, when the headphone device performs head movement tracking, the played target audio signal can be quickly switched according to the first head movement tracking data and the second head movement tracking data, thereby improving the real-time response effect of the audio playback when the headphone device performs head movement tracking.
[0008] In a possible implementation manner, the electronic device generates a third HRTF according to the first HRTF and the second HRTF, including: the electronic device removes the first interaural time difference in the first HRTF to obtain a fourth HRTF, and removes the first interaural time difference in the second HRTF to obtain a fifth HRTF; the electronic device generates a sixth HRTF according to the fourth HRTF and the fifth HRTF; the electronic device obtains the second interaural time difference of the target user, and the second interaural time difference is related to the head parameters of the target user; the electronic device adds the second interaural time difference to the sixth HRTF to obtain the third HRTF. In this way, since the generated third HRTF includes the second interaural time difference of the target user, the subsequent second BRIR used for rendering the audio signal to be rendered also includes the second interaural time difference of the target user, and the second interaural time difference is related to the actual head parameters of the target user. Therefore, when the second BRIR is used to render the audio signal to be rendered to generate the target audio signal, it can meet the personalized needs of the target user, improve the immersion of the target user when the headphone device plays the target audio signal, improve the accuracy of sound image localization, and control the externalization degree of the sound image to be more in line with the actual externalization degree required by the target user. Moreover, by first performing parameter switching according to the fourth HRTF and the fifth HRTF, and then adding the second interaural time difference of the target user, the spatial localization accuracy of the finally rendered target audio signal can be improved.
[0009] In a possible implementation, after the electronic device obtains the first HRTF based on the first head movement tracking data and obtains the second HRTF based on the second head movement tracking data, it further includes: when the first head movement tracking data is the same as the second head movement tracking data, the electronic device removes the first interaural time difference in the first target HRTF to obtain a second target HRTF; the electronic device obtains the second interaural time difference of the target user, and the second interaural time difference is related to the head parameters of the target user; the electronic device adds the second interaural time difference to the second target HRTF to obtain a third HRTF. Among them, the first target HRTF is the first HRTF, and the second target HRTF is the fourth HRTF; or, the first target HRTF is the second HRTF, and the second target HRTF is the fifth HRTF. In this way, the present application is also applicable to the audio processing scenario when the head movement tracking data does not change, enriching the applicable scenarios of the audio processing method of the present application.
[0010] In a possible implementation, the electronic device obtains the second interaural time difference of the target user, including: the electronic device obtains the head parameters of the target user; the electronic device obtains the second interaural time difference of the target user according to the head parameters and the horizontal azimuth angle in the first head movement tracking data. In this way, the present application can generate the second interaural time difference according to the head parameters and the horizontal azimuth angle in the first head movement tracking data, simplifying the calculation complexity of the second interaural time difference.
[0011] In a possible implementation, the electronic device obtains the head parameters of the target user, including: the electronic device displays a first interface, and the first interface includes a head parameter input box; the electronic device receives the head parameters input into the head parameter input box. Among them, the first interface can be Figure 11 the interface shown, and the head parameter input box can be Figure 11 the head parameter input box 111 shown. In this way, the present application can flexibly adjust the head parameters of the target user to improve the accuracy of sound image localization and control the externalization degree of the sound image to better meet the actual needs of the target user.
[0012] In a possible implementation, the head parameter is the head radius; the electronic device obtains the second interaural time difference of the target user according to the head parameter and the horizontal azimuth angle in the first head movement tracking data, including: when the electronic device is in the far-field mode, the electronic device calculates the second interaural time difference of the target user through the following formula: ITD_new = 2×a×sinθ1 / c; when the electronic device is in the near-field mode, the electronic device calculates the second interaural time difference of the target user through the following formula: ITD_new = a×(sinθ1 + θ1) / c; where ITD_new is the second interaural time difference, a is the head radius, θ1 is the horizontal azimuth angle in the first head movement tracking data, and c is the speed of sound. In this way, different calculation formulas are used in the far-field mode and the near-field mode of the electronic device to calculate the second interaural time difference of the target user, making the calculated second interaural time difference more accurate.
[0013] In a possible implementation, the sixth HRTF is obtained through the following formula: HRTF_new = b / N×HRTF4 + (1 - b / N)×HRTF5; where HRTF4 is the fourth HRTF, HRTF5 is the fifth HRTF, b is the sequence number of the sampling points, N is the total number of sampling points, and HRTF_new is the sixth HRTF. Among them, N can be 256, and b can take 1, 2,..., 256 in sequence. In this way, the sixth HRTF is calculated by using a fade-in and fade-out parameter switching method, making the change between the calculated sixth HRTFs smoother; and because the number of sampling points of the HRTF is small, when using the HRTF for parameter switching, its calculation complexity is low, thereby reducing the calculation complexity of generating the target audio signal and improving the real-time response effect of audio playback when the headphone device performs head movement tracking.
[0014] In a possible implementation, the electronic device obtains the first HRTF according to the first head movement tracking data and obtains the second HRTF according to the second head movement tracking data, including: the electronic device obtains the seventh HRTF corresponding to the first head movement tracking data and the eighth HRTF corresponding to the second head movement tracking data from the HRTF database; the electronic device adjusts the seventh HRTF with the first gain to obtain the first HRTF, and adjusts the eighth HRTF with the first gain to obtain the second HRTF. In this way, by adjusting the seventh HRTF and the eighth HRTF with the first gain, the amplitude of the direct sound part in the subsequent generated second BRIR can be adjusted, thereby controlling the externalization effect achieved when the target audio signal is played through the headphone device.
[0015] In a possible implementation, before the electronic device splices the third HRTF with the first BRIR to obtain the second BRIR, it further includes: the electronic device extracts the reverberation part from the original BRIR to obtain the third BRIR; the electronic device determines the first BRIR according to the third BRIR. In this way, since the first BRIR includes the reverberation part, the subsequently generated second BRIR also includes the reverberation part. Therefore, when using the second BRIR including the reverberation part to render the audio signal to be rendered and playing the rendered target audio signal using the headphone device, the target user wearing the headphone device can perceive the externalization effect of the sound.
[0016] In a possible implementation, the electronic device determines the first BRIR according to the third BRIR, including: the electronic device processes the third BRIR using the reverberation parameters to obtain the first BRIR; the reverberation parameters include the reverberation time and / or the second gain. In this way, the reverberation time and / or the second gain can be used to enable the target audio signal played by the headphone device to achieve different externalization effects.
[0017] In a possible implementation, the electronic device splices the third HRTF with the first BRIR to obtain the second BRIR, including: the electronic device splices the first BRIR after the third HRTF to obtain the second BRIR. In this way, since when the user listens to the sound, the direct sound usually reaches the user's ears earlier than the reverberation, therefore, after splicing the first BRIR after the third HRTF and using the spliced second BRIR to render the audio signal to be rendered, the auditory experience during playback of the finally rendered target audio signal conforms to the auditory experience required by the user when played through the headphone device.
[0018] In a possible implementation, the electronic device is communicatively connected to the headphone device; the electronic device obtains the first head movement tracking data and the second head movement tracking data, including: the electronic device receives the first head movement tracking data and the second head movement tracking data sent by the headphone device; after the electronic device uses the second BRIR to render the audio signal to be rendered to obtain the target audio signal, it further includes: the electronic device sends the target audio signal to the headphone device. In this way, an implementation manner of the cooperation between the electronic device and the headphone device to achieve spatial audio rendering is provided.
[0019] In a second aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, the memory is used to store a computer program, and the processor is used to call the computer program to execute the above audio processing method.
[0020] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program or instructions, which, when run, implement the above-mentioned audio processing method.
[0021] In a fourth aspect, an embodiment of the present application provides a computer program product including a computer program, which, when run, causes a computer to execute the above-mentioned audio processing method.
[0022] The possible implementation manners of the second to fourth aspects have effects similar to those of the first aspect and the possible designs of the first aspect, and will not be elaborated herein. Description of the Drawings
[0023] Figure 1 It is a schematic diagram of an application scenario of the audio processing method provided by an embodiment of the present application;
[0024] Figure 2 It is a schematic diagram of an application scenario of an audio processing method provided by a related technology;
[0025] Figure 3 It is a schematic diagram of an application scenario of an audio processing method provided by another related technology;
[0026] Figure 4 It is a schematic diagram of the hardware system structure of an electronic device provided by an embodiment of the present application;
[0027] Figure 5 It is a schematic diagram of the hardware system structure of a headphone device provided by an embodiment of the present application;
[0028] Figure 6 It is a schematic diagram of the software system structure of an electronic device provided by an embodiment of the present application;
[0029] Figure 7 It is a schematic flowchart of an audio processing method provided by an embodiment of the present application;
[0030] Figure 8 It is a schematic diagram of head movement tracking data in a three-dimensional coordinate system provided by an embodiment of the present application;
[0031] Figure 9 It is a processing flowchart of generating a target audio signal provided by an embodiment of the present application;
[0032] Figure 10 It is a schematic curve diagram of ITD corresponding to different head radii provided by an embodiment of the present application;
[0033] Figure 11 It is a schematic interface diagram of obtaining a head radius provided by an embodiment of the present application;
[0034] Figure 12 Another schematic diagram of the interface for obtaining the head radius provided by the embodiment of the present application;
[0035] Figure 13 Schematic diagram of the amplitude curve of the pulse signals received by both ears when the horizontal azimuth angle is 0° provided by the embodiment of the present application;
[0036] Figure 14 Schematic diagram of the amplitude curve of the pulse signals received by both ears when the horizontal azimuth angle is 90° provided by the embodiment of the present application;
[0037] Figure 15 Schematic diagram of the principle for calculating the ITD in the far - field scenario and the near - field scenario provided by the embodiment of the present application;
[0038] Figure 16 Schematic diagram of the interface for enabling the spatial audio rendering function and the head - movement tracking rendering function provided by the embodiment of the present application;
[0039] Figure 17 Schematic diagram of the interface with different ear contours provided by the embodiment of the present application;
[0040] Figure 18 Schematic diagram of the interface after replacing the user wearing the headphone device provided by the embodiment of the present application;
[0041] Figure 19 Schematic diagram of the structure of an electronic device provided by the embodiment of the present application. Detailed implementation manners
[0042] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. For example, the first chip and the second chip are only used to distinguish different chips, and do not limit their sequence. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.
[0043] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0044] In the embodiments of the present application, "at least one" refers to one or more, and "plurality" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0045] For example, Figure 1 A schematic diagram of a scenario provided in an embodiment of the present application. Figure 1 In the corresponding embodiments, an electronic device is taken as a mobile phone as an example for illustration, and this example does not constitute a limitation on the embodiments of the present application.
[0046] like Figure 1 As shown, the scene may include an electronic device 100 and an earphone device 200 worn by a user. The earphone device 200 may be as follows: Figure 1 The headset shown, or the headphone device 200 may also be a true wireless stereo (TWS) or a wired headset, etc. The specific type of the headphone device 200 is not limited in the embodiment of the present application.
[0047] Exemplarily, when the electronic device 100 establishes a communication connection with the headphone device 200, the electronic device 100 can perform spatial audio rendering on the audio signal to be rendered, obtain a target audio signal, and send the target audio signal to the headphone device 200, and play the target audio signal through the headphone device 200 to achieve virtual 3D sound effect playback.
[0048] Virtual 3D playback technology based on headphone devices simulates the sound field generated by a point sound source in space in both ears, so that the user wearing the headphone device feels that the virtual sound source is emitted from the corresponding position in space. This technology mainly uses HRTF to perform azimuth rendering on the audio signal to simulate a point sound source in space to produce a directional rendering effect.
[0049] HRTF is a processing technology for sound localization, which is the acoustic transfer function between the sound source and the two ears under free-field conditions, and includes the influence of the head, earlobes, shoulders, and part of the torso on sound transmission. The spatial audio rendering function can be understood as: simulating the effect of listening to audio in different spaces, performing audio processing on the audio signal to be rendered, so that the processed audio signal has a sense of space and three-dimensionality.
[0050] Since the sound image externalization effect is limited when using HRTF to process audio signals, therefore, the electronic device can introduce reverberation control to add a reverberation effect to the audio signal to increase the sense of externalization of the audio signal.
[0051] When sound waves propagate indoors, they are reflected by obstacles such as walls, ceilings, and floors, and each reflection will absorb some by the obstacles. In this way, when the sound source stops emitting sound, the sound waves will undergo multiple reflections and absorptions in the room before disappearing. Therefore, users can feel that the sound continues for a period of time after the sound source stops emitting sound. This phenomenon is called reverberation.
[0052] The addition of the reverberation effect can be achieved through BRIR. BRIR considers the influence of environmental reflected sound on the sound source for playback. BRIR can be regarded as the impulse response of a system composed of the sound source, the indoor environment, and the two ears (including the head, torso, and earlobes), which consists of direct sound, early reflections, and late reverberation.
[0053] In a related technology, as Figure 2 shown, the electronic device can perform spatial audio rendering on the audio signal to be rendered in the following manner: after the electronic device obtains the first head movement tracking data and the second head movement tracking data sent by the headphone device, the electronic device obtains the first HRTF corresponding to the first head movement tracking data from the HRTF database, and obtains the second HRTF corresponding to the second head movement tracking data from the HRTF database; the electronic device convolves the audio signal to be rendered with the first HRTF to obtain a first rendered signal, and convolves the audio signal to be rendered with the second HRTF to obtain a second rendered signal; then, the electronic device uses the crossfading method to generate a third rendered signal according to the first rendered signal and the second signal to be rendered; finally, the electronic device can continue to convolve the third rendered signal with the first BRIR to obtain the target audio signal. Thus, when the headphone device performs sound effect playback, the sound image localization effect and externalization effect can be achieved.
[0054] Among them, when the user wears the headphone device, the first head movement tracking data is the relative position or orientation between the user's current head and the sound source collected by the headphone device, and the second head movement tracking data is the relative position or orientation between the user's head and the sound source collected by the headphone device last time. The first BRIR refers to the reverberation part in the original BRIR.
[0055] In this spatial audio rendering method, when the position or orientation of the user's head changes, three convolution operations are required to generate the target audio signal. Moreover, since there are a large number of sampling points for the signal during each convolution operation, the computational complexity of the target audio signal is high and the memory occupied by the calculation increases. As a result, the speed at which the electronic device renders the target audio signal in real time becomes slower, and further, the link delay for transmitting the target audio signal between the electronic device and the headphone device also increases, leading to a poor real-time response effect when the headphone device plays audio. For example, when the headphone device plays audio, due to the slow real-time response speed, there may be a pop sound similar to background noise (such sounds are manifested as "puff puff" sounds).
[0056] Exemplarily, audio signals are transmitted in units of frames. The duration required for transmitting one frame of audio signal to be rendered can be 10 ms. Taking the sampling rate of the audio signal to be rendered as 48 KHz as an example, 48 sampling points can be sampled in 1 ms. Then, the number of sampling points for one frame of audio signal to be rendered is 480, which results in a high computational complexity for the convolution calculation of the audio signal to be rendered with the first HRTF and the convolution calculation of the audio signal to be rendered with the second HRTF. Additionally, the number of sampling points of the first BRIR is also relatively large, which also leads to a high computational complexity for the convolution calculation of the third rendered signal with the first BRIR.
[0057] In another related technology, as Figure 3 shown, the electronic device can perform spatial audio rendering on the audio signal to be rendered in the following manner: After the electronic device obtains the first head movement tracking data and the second head movement tracking data sent by the headphone device, the electronic device obtains the fourth BRIR corresponding to the first head movement tracking data from the BRIR database, and obtains the fifth BRIR corresponding to the second head movement tracking data from the BRIR database; the electronic device convolves the audio signal to be rendered with the fourth BRIR to obtain the fourth rendered signal, and convolves the audio signal to be rendered with the fifth BRIR to obtain the fifth rendered signal; then, the electronic device uses the crossfading method to generate the target audio signal based on the fourth rendered signal and the fifth rendered signal.
[0058] In this spatial audio rendering method, when the position or orientation of the user's head changes, two convolution operations are required to generate the target audio signal. Moreover, since there are a large number of sampling points for the signal during each convolution operation, the computational complexity of the target audio signal is high and the memory occupied by the calculation increases. As a result, the speed at which the electronic device renders the target audio signal in real time becomes slower, and further, the link delay for transmitting the target audio signal between the electronic device and the headphone device also increases, leading to a poor real-time response effect when the headphone device plays audio.
[0059] Based on this, the embodiments of the present application provide an audio processing method. After obtaining the first HRTF according to the first head movement tracking data and obtaining the second HRTF according to the second head movement tracking data, the third HRTF is generated according to the first HRTF and the second HRTF. Then, the third HRTF is spliced with the first BRIR to obtain the second BRIR. Finally, the second BRIR is used to render the audio signal to be rendered, and the target audio signal is obtained. This spatial audio rendering method based on head movement tracking data only requires one convolution operation to generate the target audio signal when the head position or orientation of the user changes, thereby reducing the computational complexity and the memory occupied by the computation, improving the speed of the electronic device to render the target audio signal in real time, reducing the link delay of sending the target audio signal between the electronic device and the headphone device, and enhancing the real-time response effect of the headphone device during audio playback.
[0060] Moreover, the spatial audio rendering method based on head movement tracking data in the embodiments of the present application can achieve the 3D surround sound effect on the one hand and simulate the accuracy of sound image localization in the real scene on the other hand when playing the rendered target audio signal using the headphone device.
[0061] The accuracy of sound image localization refers to that the sound image position heard by the target user from the worn headphone device is consistent with the sound image position in the real physical space. The target user can refer to the user wearing the headphone device.
[0062] For example, after the target user wears the headphone device, if the electronic device is directly in front of the target user's face and the position of the electronic device remains unchanged, when the target user's head deflects 90° to the left, the target audio signals heard by the target user's left ear and right ear based on the headphone device change, so that the sound heard by the target user based on the headphone device is transmitted from the right ear side, making the sound image position heard by the target user from the worn headphone device consistent with the sound image position in the real physical space.
[0063] The spatial audio rendering method based on head movement tracking data in the embodiments of the present application can be applied to audio-visual technologies such as virtual reality (VR) and augmented reality (AR) to enhance the immersion and sense of space of VR / AR. Moreover, the audio signal to be rendered during the audio processing process can be the audio signal in scenarios such as video scenes, game scenes, call scenes, and music playback scenes.
[0064] It can be understood that the 3D in the above virtual 3D replay technology can refer to three degrees of freedom (DOF), and the 3DOF data can include three rotation angles of the head, such as the horizontal azimuth angle, the tilt angle, and the pitch angle.
[0065] Of course, the audio processing method in the embodiments of the present application can also be applied to scenarios such as virtual 4D replay technology, virtual 5D replay technology, and virtual 6D replay technology. 6D refers to 6DOF. In addition to including three rotation angles, 6DOF data also includes the translation distances in the front-back direction, left-right direction, and up-down direction generated by the user's body movement; 4D refers to 4DOF. In addition to including three rotation angles, 4DOF data also includes the translation distance in any one of the front-back direction, left-right direction, and up-down direction; 5D refers to 5DOF. In addition to including three rotation angles, 5DOF data also includes the translation distances in any two of the front-back direction, left-right direction, and up-down direction.
[0066] The electronic device provided in the embodiments of the present application can be a smart phone, a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), a personal digital assistant (PDA), a television (which can also be called a smart TV, a smart screen, or a large-screen device), a wearable device (such as a smart watch or a bracelet), a vehicle-mounted device, a VR device, and an AR device, etc. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the electronic device.
[0067] In order to better understand the embodiments of the present application, the structure of the electronic device in the embodiments of the present application will be introduced below.
[0068] Figure 4 FIG. shows a schematic structural diagram of an electronic device 100. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0069] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figures, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0070] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0071] The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0072] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0073] The USB interface 130 is an interface that conforms to the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used for data transmission between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio through the headphones. This interface can also be used to connect other electronic devices, such as AR devices, etc.
[0074] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is for illustrative purposes only and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0075] The charging management module 140 is used to receive a charging input from a charger. Herein, the charger can be a wireless charger or a wired charger. The power management module 141 is used to connect to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc.
[0076] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.
[0077] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. The antennas in the electronic device 100 can be used to cover single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0078] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, perform filtering, amplification, etc. on the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation.
[0079] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive the signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.
[0080] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change the display information.
[0081] The display screen 194 is used to display images, videos, and receive sliding operations, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include one or more display screens 194.
[0082] The electronic device 100 can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0083] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive component. The optical signal is converted into an electrical signal, and the camera photosensitive component transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye.
[0084] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive component. The photosensitive component converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard formats such as RGB and YUV. In some embodiments, the electronic device 100 may include one or more cameras 193.
[0085] The video codec is used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0086] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.
[0087] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.). The data storage area can store the data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a Universal Flash Storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0088] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. Such as music playback, recording, etc.
[0089] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0090] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A.
[0091] The receiver 170B, also known as an "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the receiver 170B can be placed close to the human ear to listen to the voice.
[0092] The microphone 170C, also known as a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak close to the microphone 170C with the mouth to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0093] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0094] The keys 190 include a power-on key, volume keys, etc. The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. The indicator 192 can be an indicator light, which can be used to indicate the charging status, power change, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device 100.
[0095] Exemplarily,Figure 5 A schematic diagram of the hardware structure of a headset device provided in an embodiment of the present application.
[0096] like Figure 5 As shown, the headphone device 200 includes one or more processors 210, one or more memories 220, a communication interface 230, an audio acquisition circuit and an audio playback circuit. The audio acquisition circuit may further include at least one microphone 240 and an analog-to-digital converter (ADC) 250. The audio playback circuit may further include a speaker 260 and a digital-to-analog converter (DAC) 270.
[0097] The headset may further include one or more sensors 280, such as an inertial measurement unit (IMU), a proximity sensor, a motion sensor, etc. The above hardware components may communicate on one or more communication buses.
[0098] In the embodiment of the present application, the IMU can be used to measure the motion posture of the headphone device 200. For example, the IMU can be used to determine the head movement tracking data when the user wears the headphone device 200. The IMU can be provided with a gyroscope sensor and an acceleration sensor.
[0099] The processor 210 is the control center of the headphone device 200. The processor 210 may also be referred to as a control unit, a controller, a microcontroller, or some other suitable term. The processor 210 uses various interfaces and lines to connect the various components of the headset. In a possible embodiment, the processor 210 may also include one or more processing cores. In a possible embodiment, a main control unit and a signal processing module may be integrated in the processor 210. The main control unit (MCU) is used to receive data collected by the sensor 280 or a monitoring signal from the signal processing module or a control signal from a terminal (such as a mobile phone APP), and finally controls the headphone device 200 through comprehensive judgment and decision-making.
[0100] The memory 220 may be coupled to the processor 210, or connected to the processor 210 via a bus, and is used to store various software programs and / or multiple sets of instructions and data. The memory 220 may also store a communication program, which may be used to communicate with a terminal. In one example, the memory 220 may also store data / program instructions, and the processor 210 may be used to call and execute the data / program instructions in the memory 220. Optionally, the memory 220 may be a memory external to the MCU, or may be a storage unit provided by the MCU.
[0101] The communication interface 230 is used to communicate with the terminal, and the communication method can be wired or wireless. When the communication method is wired communication, the communication interface 230 can be connected to the terminal through a cable. When the communication method is wireless communication, the communication interface 230 is used to receive and send radio frequency signals, and the supported wireless communication methods can be, for example, at least one of Bluetooth communication, wireless-fidelity (Wifi) communication, infrared communication, or cellular 2 / 3 / 4 / 5 generation (2G / 3G / 4G / 5G) communication, etc.
[0102] The microphone 240 can be used to collect sound signals (or called audio signals, and the audio signals are analog signals), and the analog-to-digital converter 250 is used to convert the analog signals collected by the microphone 240 into digital signals and send the digital signals to the processor 210 for processing. In a specific embodiment, it can be sent to the signal processing module for processing. The signal processing module can transmit the processed signals (such as mixed audio signals) to the digital-to-analog converter 270, and the digital-to-analog converter 270 can convert the received signals into analog signals and then transmit them to the speaker 260. The speaker 260 is used to play according to the analog signals, so that the user can hear the sound.
[0103] In the embodiment of the present application, the communication interface 230 can be used to send the head movement tracking data detected by the IMU to the electronic device 100, so that the electronic device 100 can obtain the corresponding HRTF from the HRTF database according to the head movement tracking data. And, the communication interface 230 can also be used to receive the target audio signal sent by the electronic device 100.
[0104] It can be understood that the above earphone device 200 may also be called earplugs, earphones, walkmans, audio players, media players, headset receivers, earphone devices, or some other suitable terms, and the embodiments of the present application do not limit this.
[0105] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture, etc. In the embodiment of the present application, taking the Android system with a layered architecture as an example, the software structure of the electronic device 100 is exemplarily described.
[0106] Figure 6It is a software structure block diagram of the electronic device 100 according to an embodiment of the present application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and the system library, and the kernel layer.
[0107] The application layer may include a series of application packages. Such as Figure 6 shown, the application packages may include a music application, a video application, a game application, and a headphone settings application, etc. Among them, the headphone settings application can be used to implement the setting of the audio playback effect when playing audio using the headphone device 200.
[0108] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0109] Such as Figure 6 shown, the application framework layer may include an activity manager, a notification manager, a resource manager, a phone manager, a view system, and an audio processing module, etc.
[0110] The activity manager is used to manage the life cycle of each application and the navigation back function. It is responsible for creating the main thread of Android and maintaining the life cycle of each application.
[0111] The notification manager enables applications to display notification information in the status bar. It can be used to convey messages of the notification type, and can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or a scroll bar text, such as the notification of a background running application, or a notification that appears on the screen in the form of a dialogue window. For example, prompt text information in the status bar, emit a prompt tone, the electronic device vibrates, the indicator light flashes, etc.
[0112] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0113] The phone manager is used to manage the functions of the mobile device, including: the mobile phone call status, obtaining phone information (device, SIM card, network information), monitoring the phone status, and calling the phone dialer to make a call.
[0114] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon can include a view for displaying text and a view for displaying pictures.
[0115] In some embodiments, the audio processing module is used to obtain a first HRTF according to first head movement tracking data, and obtain a second HRTF according to second head movement tracking data; in the case where the first head movement tracking data is different from the second head movement tracking data, generate a third HRTF according to the first HRTF and the second HRTF; splice the third HRTF with the first BRIR to obtain a second BRIR; use the second BRIR to render the audio signal to be rendered to obtain a target audio signal. That is to say, the audio processing module can be used to implement the execution process of generating the target audio signal in the embodiments of the present application.
[0116] Android runtime includes a core library and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system.
[0117] The core library consists of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.
[0118] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0119] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc.
[0120] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.
[0121] The media library supports the playback and recording of a variety of common audio and video formats, as well as static image files, etc. The media library can support a variety of audio and video coding formats, such as: MPEG2, H.262, MP3, AAC, AMR, JPG, PNG, etc.
[0122] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is a drawing engine for 2D drawing.
[0123] The kernel layer is the layer between hardware and software. The kernel layer includes display drivers, Bluetooth drivers, audio drivers, etc.
[0124] It should be noted that although the embodiments of this application are described with the Android system, the principle of the audio processing method is equally applicable to electronic devices operating systems such as iOS or Windows.
[0125] The technical solutions of this application and how the technical solutions of this application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be implemented independently or in combination with each other. For the same or similar concepts or processes, they may not be repeated in some embodiments.
[0126] Exemplarily, Figure 7 is a flowchart of an audio processing method provided by an embodiment of this application, which can be applied to Figure 1 the corresponding application scenarios. Referring to Figure 7 as shown, the audio processing method may specifically include the following steps:
[0127] S701, the electronic device establishes a communication connection with the headphone device.
[0128] In some embodiments, the electronic device and the headphone device can establish a communication connection by wire; alternatively, the electronic device and the headphone device can also establish a communication connection wirelessly. For example, the electronic device and the headphone device can establish a communication connection wirelessly through Bluetooth, WIFI, or connecting to the same cloud account, etc. The embodiments of this application do not limit the way the electronic device and the headphone device establish a communication connection.
[0129] S702, the electronic device receives the first head movement tracking data and the second head movement tracking data sent by the headphone device.
[0130] After the electronic device and the headphone device establish a communication connection, the electronic device can send an instruction message for obtaining head movement tracking data to the headphone device, and the headphone device collects the head movement tracking data of the target user wearing the headphone device according to the instruction message.
[0131] Exemplarily, the headphone device can collect the head movement tracking data of the target user in real time or periodically (such as 0.5ms or 1ms, etc.), and send the collected head movement tracking data to the electronic device through the established communication link, so that the electronic device obtains the head movement tracking data.
[0132] When the audio processing method of the embodiment of the present application supports head movement tracking rendering, after the target user wears the headphone device, when the position and orientation of the user's head change relative to the sound source, the head movement tracking data collected above will also change. Therefore, the head movement tracking data currently collected by the headphone device can be referred to as the first head movement tracking data, and the head movement tracking data collected by the headphone device last time can be referred to as the second head movement tracking data, that is, the second head movement tracking data is the head movement tracking data obtained the previous time of the first head movement tracking data.
[0133] It should be noted that the first head movement tracking data and the second head movement tracking data are collected at different time points. Each time the headphone device detects a head movement tracking data, it sends it to the electronic device. Therefore, in the actual scenario, the headphone sends the first head movement tracking data and the second head movement tracking data in two times, that is, it sends the second head movement tracking data first and then the first head movement tracking data.
[0134] Exemplarily, the head movement tracking data may include the horizontal azimuth angle and the pitch angle. Of course, in some actual usage scenarios, in addition to the horizontal azimuth angle and the pitch angle, the head movement tracking data may further include the tilt angle, and the translation distances in the front-back direction, left-right direction, and up-down direction generated by the user's body movement, etc.
[0135] Among them, the first head movement tracking data may include a first horizontal azimuth angle and a first pitch angle, and the second head movement tracking data may include a second horizontal azimuth angle and a second pitch angle.
[0136] As Figure 8 shown, taking the head reference point of the user as the origin, the head reference point may be the midpoint between the two ears, taking the direction where the front of the user's face is located as the X-axis direction, taking the direction of the user's left ear as the Y-axis direction, and taking the direction directly above the user's head as the Z-axis direction.
[0137] Therefore, the horizontal azimuth angle θ refers to the angle between the projection of the straight line from the sound source to the head reference point on the horizontal plane (that is, the plane formed by the X-axis and the Y-axis) and the X-axis. The value range of the horizontal azimuth angle θ is [0°, 360°], the azimuth angle corresponding to the front of the user's face is 0°, the azimuth angle corresponding to the right side of the user's face is 90°, the azimuth angle corresponding to the back of the user's face is 180°, and the azimuth angle corresponding to the left side of the user's face is 270°.
[0138] The pitch angle β refers to the angle between the straight line from the sound source to the head reference point and the horizontal plane. The value unit of the pitch angle β is [-90°, 90°], the pitch angle β corresponding to the horizontal plane is 0°, the pitch angle β corresponding to directly above the user's head is 90°, and the pitch angle β corresponding to directly below the user's head is -90°.
[0139] S703, The electronic device obtains the first HRTF based on the first head movement tracking data, and obtains the second HRTF based on the second head movement tracking data.
[0140] In some embodiments, after the electronic device obtains the second head movement tracking data, it can obtain the corresponding second HRTF according to the second head movement tracking data; correspondingly, when the electronic device obtains the first head movement tracking data, it can obtain the corresponding first HRTF according to the first head movement tracking data.
[0141] The HRTFs corresponding to different head movement tracking data are different. Therefore, when the first head movement tracking data is different from the second head movement tracking data, the first HRTF and the second HRTF are also different.
[0142] S704, When the first head movement tracking data is different from the second head movement tracking data, the electronic device generates a third HRTF according to the first HRTF and the second HRTF.
[0143] In the embodiments of the present application, after the electronic device obtains the first HRTF and the second HRTF, the electronic device does not directly convolve the audio signal to be rendered with the first HRTF and the second HRTF respectively, but performs parameter switching according to the first HRTF and the second HRTF, such as using crossfading to perform parameter switching to generate a third HRTF.
[0144] Since the number of sampling points of the HRFT is 256, which is less than the number of sampling points of one frame of audio signal (for example, when the sampling rate of the audio signal is 48KHz, the number of sampling points of the audio signal is 480), the computational complexity during parameter switching according to the first HRTF and the second HRTF is small.
[0145] S705, The electronic device splices the third HRTF with the first BRIR to obtain the second BRIR.
[0146] S706, The electronic device uses the second BRIR to render the audio signal to be rendered to obtain the target audio signal.
[0147] When the spatial audio rendering reaches a certain externalization effect, in addition to sound image localization, a certain amount of reverberation needs to be added. However, the sound image externalization effect of processing the audio signal with HRTF is limited, and HRTF generally includes a direct sound part and an obvious early reflection sound part. Therefore, the electronic device needs to add a certain amount of late reverberation part based on the third HRTF. That is, the electronic device can splice the third HRTF with the first BRIR to obtain the second BRIR, and use the second BRIR to render the audio signal to be rendered to obtain the target audio signal, so as to add a reverberation effect to the target audio signal to increase the externalization sense of the target audio signal.
[0148] Among them, the first BRIR refers to the BRIR obtained by adjusting the reverberation part in the original BRIR using reverberation parameters, or the first BRIR can also refer to the reverberation part in the original BRIR.
[0149] An achievable way is that the electronic device splices the first BRIR after the third HRTF to achieve the splicing of the third HRTF and the first BRIR, thereby obtaining the second BRIR.
[0150] Moreover, the electronic device can convolve the second BRIR with the audio signal to be rendered to achieve the rendering of the audio signal to be rendered using the second BRIR, thereby obtaining the target audio signal.
[0151] It should be noted that the first HRTF, the second HRTF, and the third HRTF can refer to parameters represented in the time domain; or, the first HRTF, the second HRTF, and the third HRTF can refer to parameters represented in the frequency domain. The electronic device can convert the third HRTF represented in the frequency domain into the time domain, and then splice the third HRTF represented in the time domain with the first BRIR to obtain the second BRIR.
[0152] S707, the electronic device sends the target audio signal to the headphone device.
[0153] S708, the headphone device plays the target audio signal through the speaker.
[0154] In some embodiments, after the electronic device renders the target audio signal, the electronic device can send the target audio signal to the headphone device through the established communication link. The headphone device plays the target audio signal through the speaker after receiving the target audio signal sent by the electronic device.
[0155] It should be noted that the target audio signal can include the target audio signal corresponding to the left channel and the target audio signal corresponding to the right channel, so that the headphone device can play the target audio signal corresponding to the left channel through the left speaker and play the target audio signal corresponding to the right channel through the right speaker.
[0156] In summary, in the audio processing method according to the embodiments of the present application, since only one convolution operation is required to generate the target audio signal, the computational complexity and the memory occupied by the computation are reduced, the speed of the electronic device for rendering the target audio signal in real time is increased, and the link delay of sending the target audio signal between the electronic device and the headphone device is reduced. In this way, when the headphone device performs head movement tracking, the played target audio signal can be quickly switched according to the first head movement tracking data and the second head movement tracking data, thereby improving the real-time response effect of the audio playback when the headphone device performs head movement tracking.
[0157] Moreover, since the embodiments of the present application also use HRTF and BRIR to render the audio signal to be rendered, when the headphone device plays the rendered target audio signal, the target user wearing the headphone device can perceive the sound localization effect of the virtual sound source in the three-dimensional space and the externalization effect of the sound.
[0158] In some scenarios, when the HRTF in the current HRFT database is recorded, a series of HRTFs are recorded by simplifying and simulating the head, auricle, shoulders and part of the torso of the actual user through an artificial head or a specific small number of subjects. The head of the artificial head or a small number of subjects cannot truly represent the head of the target user wearing the headphone device, resulting in poor generality of the recorded HRFT.
[0159] When the head parameters of the target user wearing the headphone device match the head parameters of the human head during recording, a relatively good spatial rendering effect can be felt when using the recorded HRTF to render the audio to be rendered. However, when there are significant differences between the head parameters of the target user wearing the headphone device and the head parameters of the human head during recording, when using the recorded HRTF to render the audio to be rendered, the immersion of the target user when the headphone device plays the rendered target audio signal is not strong, and it is easy to confuse the localization direction of the sound source, and there is a difference between the externalization degree of the sound image and the externalization degree required by the target user.
[0160] Therefore, the electronic device in the embodiments of the present application can further modify the first interaural time difference included in the HRTF in the HRTF database according to the actual head parameters of the target user, so that the HRTF finally used by the electronic device to render the audio to be rendered includes the second interaural time difference related to the head parameters of the target user, thereby improving the immersion of the target user when the headphone device plays the rendered target audio signal, and can control the externalization degree of the sound image and improve the accuracy of sound image localization.
[0161] In Figure 7Based on the corresponding embodiments, the electronic device can also control the generated target audio signal according to the actual head parameters of the target user. To further illustrate the specific implementation process of the electronic device generating the target audio signal, reference can be made to Figure 9 the corresponding embodiments.
[0162] Exemplarily, Figure 9 FIG. is a processing flow chart for generating a target audio signal provided by an embodiment of the present application. Referring to Figure 9 as shown, the method for generating a target audio signal may specifically include the following steps:
[0163] S901, the electronic device obtains the seventh HRTF corresponding to the first head movement tracking data and the eighth HRTF corresponding to the second head movement tracking data from the HRTF database.
[0164] In the process of generating the target audio signal in the embodiment of the present application, a general HRTF database can be used, and the HRTF database includes HRTFs corresponding to different orientations. The embodiment of the present application does not limit the adopted HRTF database, which may be a CIPIC (center for image processing and integrated computing) database, a SADIE database, etc.
[0165] Therefore, after the electronic device obtains the second head movement tracking data, it can obtain the eighth HRTF corresponding to the second head movement tracking data from the HRTF database; correspondingly, after the electronic device obtains the first head movement tracking data, it can obtain the seventh HRTF corresponding to the first head movement tracking data from the HRTF database.
[0166] S902, the electronic device adjusts the seventh HRTF with a first gain to obtain a first HRTF, and adjusts the eighth HRTF with the first gain to obtain a second HRTF.
[0167] In some embodiments, the electronic device can adjust the amplitude of the seventh HRTF with the first gain in the direct mixing ratio control parameter to obtain a first HRTF, and adjust the amplitude of the eighth HRTF with the first gain to obtain a second HRTF.
[0168] Specifically, the amplitude of the seventh HRTF is multiplied by the first gain to obtain the first HRTF; correspondingly, the amplitude of the eighth HRTF is multiplied by the first gain to obtain the second HRTF.
[0169] In another implementation manner, the electronic device may not adjust the seventh HRTF and the eighth HRTF with the first gain. In this case, the first HRTF is the seventh HRTF, and the second HRTF is the eighth HRTF.
[0170]
[0170] In S903, the electronic device determines whether the first head movement tracking data is different from the second head movement tracking data.
[0171]
[0171] To more accurately control the head movement tracking rendering effect, the electronic device may determine whether the first head movement tracking data is different from the second head movement tracking data.
[0172]
[0172] Among them, the first head movement tracking data may include a first horizontal azimuth angle and a first pitch angle, and the second head movement tracking data may include a second horizontal azimuth angle and a second pitch angle. In this case, the electronic device determines whether the first horizontal azimuth angle is different from the second horizontal azimuth angle, and whether the first pitch angle is different from the second pitch angle.
[0173]
[0173] When the first horizontal azimuth angle is different from the second horizontal azimuth angle, and / or the first pitch angle is different from the second pitch angle, the electronic device determines that the first head movement tracking data is different from the second head movement tracking data. At this time, the electronic device executes S904 below. When the first horizontal azimuth angle is the same as the second horizontal azimuth angle, and the first pitch angle is the same as the second pitch angle, the electronic device determines that the first head movement tracking data is the same as the second head movement tracking data. At this time, the electronic device executes S909 below.
[0174]
[0174] It should be noted that the electronic device may determine whether the first head movement tracking data is different from the second head movement tracking data after executing S902. Alternatively, after the electronic device obtains the first head movement tracking data and the second head movement tracking data sent by the headphone device, it may also first determine whether the first head movement tracking data is different from the second head movement tracking data. When the first head movement tracking data is different from the second head movement tracking data, it executes S901, S902, S904 and subsequent steps. When the first head movement tracking data is the same as the second head movement tracking data, the electronic device may directly obtain the HRTF corresponding to the first head movement tracking data or the second head movement tracking data from the HRTF database, adjust the obtained HRTF with the first gain, and then execute S909 and subsequent steps. The specific execution order of the determination steps of the first head movement tracking data and the second head movement tracking data in the embodiments of the present application is not limited.
[0175]
[0175] In S904, when the first head movement tracking data is different from the second head movement tracking data, the electronic device removes the first interaural time difference in the first HRTF to obtain a fourth HRTF, and removes the first interaural time difference in the second HRTF to obtain a fifth HRTF.
[0176] In some embodiments, the HRTF is related to parameters such as the sound source position and frequency, and includes the interaural level difference (ILD), the interaural time difference (ITD), and the inter-aural coherence (IC), etc. The interaural time difference refers to the time interval for the audio signal to reach the user's left ear and right ear, and is related to the user's head parameters.
[0177] Since the HRTF in the HRTF database is recorded based on an artificial head or the heads of a small number of subjects, there is a difference between the first interaural time difference in the seventh HRTF and the eighth HRTF obtained by the electronic device from the HRTF database and the actual second interaural time difference of the target user wearing the headphone device. Correspondingly, there is also a difference between the first interaural time difference in the first HRTF and the second HRTF obtained after the first gain adjustment and the actual second interaural time difference of the target user wearing the headphone device.
[0178] In order to accurately control the externalization degree of the sound image and improve the accuracy of sound image localization in the future, the electronic device can remove the first interaural time difference in the first HRTF to obtain the fourth HRTF, and remove the first interaural time difference in the second HRTF to obtain the fifth HRTF. That is, the fourth HRTF lacks the first interaural time difference compared to the first HRTF, and the fifth HRTF also lacks the first interaural time difference compared to the second HRTF.
[0179] S905, the electronic device generates a sixth HRTF according to the fourth HRTF and the fifth HRTF.
[0180] After obtaining the fourth HRTF and the fifth HRTF, the electronic device performs parameter switching according to the fourth HRTF and the fifth HRTF to generate a sixth HRTF.
[0181] In some alternative implementation manners, the electronic device can perform parameter switching in the way of crossfading. In this case, the sixth HRTF is obtained through the following formula:
[0182] HRTF_new = b / N × HRTF4 + (1 - b / N) × HRTF5;
[0183] where HRTF4 is the fourth HRTF, HRTF5 is the fifth HRTF, b is the sequence number of the sampling points, N is the total number of sampling points, and HRTF_new is the sixth HRTF. Usually, the number of sampling points of the HRTF is 256, so N can be taken as 256, and b can be taken as 1, 2,..., 256 in sequence.
[0184] Since the number of sampling points of the HRTF is small, when performing parameter switching using the HRTF, its computational complexity is low, thereby reducing the computational complexity of generating the target audio signal, improving the real-time response effect of audio playback when the headphone device performs head movement tracking, and reducing the possibility of the headphone device generating pop sounds.
[0185] S906, the electronic device obtains the head parameters of the target user.
[0186] Since in the above process, the electronic device has removed the first interaural time difference in the first HRTF to obtain the fourth HRTF, and removed the first interaural time difference in the second HRTF to obtain the fifth HRTF, and generated the sixth HRTF according to the fourth HRTF and the fifth HRTF, the sixth HRTF also lacks the interaural time difference.
[0187] In order to accurately control the degree of externalization of the sound image and improve the accuracy of sound image localization subsequently, the electronic device can obtain the second interaural time difference of the target user and add it to the sixth HRTF, so that subsequently when rendering the audio signal to be rendered using the second BRIR including the second interaural time difference, the degree of externalization of the sound image can be accurately controlled and the accuracy of sound image localization can be improved.
[0188] The second interaural time difference is related to the actual head parameters of the target user. When the horizontal azimuth angle is the same, when the actual head parameters of the target user are different, the calculated second interaural time difference is also different. Therefore, the electronic device first needs to obtain the head parameters of the target user. Among them, the head parameters of the target user can refer to parameters such as the head radius or head diameter of the target user. The head radius can also be called the head radius of the human head, and the head diameter can also be called the head diameter of the human head.
[0189] Exemplarily, Figure 10 It is a schematic diagram of the ITD curve corresponding to different head radii provided by the embodiment of the present application, which shows the ITD sample delays corresponding to three different head radii. It is an ITD sample delay curve obtained under the sampling rate of 48KHz. The abscissa is the incident angle, which can be understood as the horizontal azimuth angle, and the ordinate represents the ITD sample delay.
[0190] It can be seen that in the range of 0° to 90° of the horizontal azimuth angle, the ITD sample delays corresponding to different head radii are different. For example, when the horizontal azimuth angle is 90°, the head radius of 0.1075m corresponds to 42 ITD sample delays, while the head radius of 0.0875 corresponds to 34 ITD sample delays.
[0191] There are various ways for the electronic device to obtain the head parameters of the target user. The following gives two feasible ways to obtain the head parameters of the target user.
[0192] In one way, an electronic device displays a first interface, and the first interface includes a head parameter input box; the electronic device receives the head parameters input into the head parameter input box.
[0193] As Figure 11 shown, a head parameter input box 111 is displayed on the first interface, and the head parameter input box 111 can be a head radius input box.
[0194] As Figure 11 shown in (a) of [], the default input in the head parameter input box 111 is the head radius of an artificial head, such as a = 0.0875m. The target user can perform a touch operation on the Figure 11 head parameter input box 111 shown in (a) of [], such as a click operation, etc., and input the actual head radius of the target user into the head parameter input box 111. For example, the actual head radius a = 0.1075m of the target user input into the head parameter input box 111 causes the electronic device to display the Figure 11 interface shown in (b) of [].
[0195] Moreover, in the interfaces shown in (a) and (b) of Figure 11 [], a head effect schematic image is also displayed. Figure 11 The interface shown in (a) of [] shows a first head effect schematic image 112a, which is used to represent the effect schematic diagram when the head radius a = 0.0875m; Figure 11 The interface shown in (b) of [] shows a first head effect schematic image 112a and a second head effect schematic image 112b. The second head effect schematic image 112b can be used to represent the effect schematic diagram when the modified head radius a = 0.1075m.
[0196] After modifying the head radius in the head parameter input box 111, the first head effect schematic image 112a corresponding to the head radius before modification and the second head effect schematic image 112b corresponding to the modified head radius are shown in the Figure 11 interface shown in (b) of [], so as to prompt the user to observe the modification effect of the head radius.
[0197] In another way, an electronic device displays a second interface, and the second interface includes a head parameter adjustment control; the electronic device receives a touch operation of the user on the head parameter adjustment control to adjust the head parameters of the target user.
[0198] As Figure 12As shown, a head parameter adjustment control 1211 is displayed on the second interface. The head parameter adjustment control 1211 can be a head radius adjustment control. The head radius adjustment control is used to indicate the specific size of the currently set head radius, the minimum head radius allowed to be set, and the maximum head radius allowed to be set, etc. For example, Figure 12 the minimum head radius allowed to be set by the head radius adjustment control in
[0199] is 5 cm, and the maximum head radius allowed to be set is 15 cm. Figure 12 As shown in (a) of Figure 12 , the head radius currently indicated by the head parameter adjustment control 1211 is the default head radius a = 8.75 cm. The target user can perform touch operations on the head parameter adjustment control 1211 shown in (a) of
[0200] , such as a swipe operation. When the swipe direction of the swipe operation is to the right, the head radius can be increased. When the swipe direction of the swipe operation is to the left, the head radius can be decreased. Figure 12 After the target user swipes the head parameter adjustment control 1211 shown in (a) of Figure 12 to the right, the electronic device can display the interface shown in (b) of
[0201] . At this time, the head radius indicated by the head parameter adjustment control 1211 is the adjusted head radius. For example, the adjusted head radius is 12 cm. Figure 12 And, a head effect schematic image is also displayed in the interfaces shown in (a) and (b) of Figure 12 . The third head effect schematic image 122a shown in the interface of (a) of Figure 12 is used to represent the schematic diagram of the effect when the head radius is 8.75 cm; the fourth head effect schematic image 122b shown in the interface of (b) of
[0202] can be used to represent the schematic diagram of the effect when the modified head radius is 12 cm.
[0203] It can be understood that in addition to the above two ways for the electronic device to obtain the head parameters of the target user, other ways can also be used to obtain the head parameters of the target user. For example, the head parameters of the target user can be automatically collected through sensors in the headphone device and sent to the electronic device. Figure 12The head parameter adjustment control 1211 in the shown interface further adjusts the head parameters of the target user. The adjusted head parameters can be greater than or less than the actual head parameters of the target user, so that the target audio signal finally rendered by the electronic device can achieve different spatial audio rendering effects.
[0204] S907, the electronic device obtains the second interaural time difference of the target user according to the head parameters and the horizontal azimuth angle in the first head movement tracking data.
[0205] In some embodiments, when calculating the second interaural time difference of the target user, the second interaural time difference is related not only to the head parameters of the target user but also to the horizontal azimuth angle.
[0206] Such as Figure 13 and Figure 14 shown, Figure 13 shows a schematic diagram of the amplitude curve of the pulse signal received by both ears when the horizontal azimuth angle is 0°, Figure 14 shows a schematic diagram of the amplitude curve of the pulse signal received by both ears when the horizontal azimuth angle is 90°. The abscissa is the sampling point, the ordinate is the amplitude of the audio signal received by both ears, and the curves represented by the solid line and the dashed line respectively represent the amplitudes of the pulse signals received by the left ear and the right ear.
[0207] Comparing Figure 13 and Figure 14 it can be seen that near the 200th sampling point, the amplitudes of the pulse signals received by both ears when the horizontal azimuth angle is 0° are very close, while the amplitudes of the pulse signals received by both ears when the horizontal azimuth angle is 90° are very different.
[0208] Therefore, the electronic device can obtain the second interaural time difference of the target user according to the head parameters of the target user and the horizontal azimuth angle in the first head movement tracking data.
[0209] In the actual process of calculating the second interaural time difference, it is necessary to first select whether the electronic device is in the far-field mode or the near-field mode according to the distance between the sound source and the human ear. For example, when the distance between the sound source and the human ear is greater than or equal to 1 m, the electronic device can be set to be in the far-field mode, and when the distance between the sound source and the human ear is less than 1 m, the electronic device can be set to be in the near-field mode.
[0210] Such as Figure 15 shown in (a) of, when the electronic device is in the far-field mode, the influence of the human head on the interaural time difference can be ignored, the incident wave can be approximated as a parallel wave, and the interaural time difference is only related to the radius of the human head. The incident wave can refer to the transmission signal in the process of the sound source transmitting to the human ear.
[0211] Therefore, when the electronic device is in the far-field mode, the electronic device calculates the second interaural time difference of the target user through the following formula: ITD_new = 2 × a × sinθ1 / c; where ITD_new is the second interaural time difference, a is the head radius of the target user (i.e., the human head radius), θ1 is the horizontal azimuth angle in the first head movement tracking data (i.e., the first horizontal azimuth angle), and c is the speed of sound.
[0212] As Figure 15 shown in (b) of , when the electronic device is in the near-field mode, it is necessary to consider the influence of the human head on the interaural time difference, that is, to consider the influence of the curved surface of the head on the interaural time difference.
[0213] Therefore, when the electronic device is in the near-field mode, the electronic device calculates the second interaural time difference of the target user through the following formula: ITD_new = a × (sinθ1 + θ1) / c, where θ1 can be in radians greater than or equal to 0 and less than π / 2.
[0214] It should be noted that the electronic device can display a far-field mode selection control and a near-field mode selection control on the interface. The user can perform a touch operation on the far-field mode selection control or the near-field mode selection control according to the distance between the sound source and the human ear as required to set whether the electronic device is in the far-field mode or the near-field mode.
[0215] S908, the electronic device adds the second interaural time difference to the sixth HRTF to obtain the third HRTF.
[0216] After the electronic device obtains the sixth HRTF and the second interaural time difference of the target user, it can add the second interaural time difference to the sixth HRTF to obtain the third HRTF.
[0217] Since the related technology first convolves the audio signal to be rendered (which includes the first interaural time difference) with the first HRTF and the second HRTF respectively, and then performs parameter switching on the convolved information, this method will blur some ITD information, resulting in a lower spatial localization accuracy of the finally rendered target audio signal. In the embodiment of the present application, parameter switching is performed first and then the second interaural time difference of the target user is added, which can improve the spatial localization accuracy of the finally rendered target audio signal.
[0218] S909, when the first head movement tracking data is the same as the second head movement tracking data, the electronic device removes the first interaural time difference in the first target HRTF to obtain the second target HRTF.
[0219] In practical applications, the first head movement tracking data and the second head movement tracking data obtained by the electronic device may also be equal. Therefore, when the first head movement tracking data is the same as the second head movement tracking data, the electronic device can remove the first interaural time difference in the first HRTF to obtain a fourth HRTF, or the electronic device can remove the first interaural time difference in the second HRTF to obtain a fifth HRTF.
[0220] Among them, the first target HRTF is the first HRTF, and the second target HRTF is the fourth HRTF; or the first target HRTF is the second HRTF, and the second target HRTF is the fifth HRTF.
[0221] S910, the electronic device adds a second interaural time difference to the second target HRTF to obtain a third HRTF.
[0222] After the electronic device obtains the second interaural time difference of the target user and the second target HRTF (the fourth HRTF or the fifth HRTF), it adds the second interaural time difference to the second target HRTF to obtain a third HRTF.
[0223] S911, the electronic device extracts the reverberation part from the original BRIR to obtain a third BRIR.
[0224] After the electronic device generates the third HRTF, in order to make the finally rendered target audio signal have an externalization effect when played through the headphone device, it can obtain the first BRIR, splice the third HRTF with the first BRIR, and use the spliced second BRIR to render the audio signal to be rendered to obtain the target audio signal.
[0225] When the electronic device generates the first BRIR, it can first extract the reverberation part from the original BRIR to obtain a third BRIR.
[0226] The original BRIR may include a direct sound part, an early reflection sound part, and a reverberation part. The reverberation part provides room-related information for the target user wearing the headphone device. The electronic device removes the direct sound part and the early reflection part in the original BRIR, and uses the remaining reverberation part as the third BRIR.
[0227] When the electronic device removes the direct sound part and the early reflection part in the original BRIR, the removed duration can be determined according to the number of sampling points of the third HRTF. For example, if the number of sampling points of the third HRTF is 256, the electronic device can remove 256 sampling points after the start stage of the original BRIR to obtain the third BRIR.
[0228] The S912 electronic device processes the third BRIR using reverberation parameters to obtain the first BRIR; the reverberation parameters include reverberation time and / or a second gain.
[0229] In some embodiments, the electronic device can use the reverberation time and / or the second gain to adjust the third BRIR to obtain the first BRIR, so that after the second BRIR spliced with the third HRTF and the first BRIR is used to render the audio signal to be rendered, the reverberation effects of different rooms can be simulated.
[0230] In one way, the electronic device can use the reverberation time to adjust the third BRIR to control the externalization effect achieved when the target audio signal is played through the headphone device. Among them, the reverberation time refers to the time required for the sound pressure level to decrease by 60 dB after the sound source stops emitting sound.
[0231] The length of the reverberation time of a room depends on the volume of the room and the sound absorption. If the room volume is large and the sound absorption is small, the reverberation time will be long. If the room volume is small and the sound absorption is strong, the reverberation time will be short. A reverberation time within 1 s will increase the clarity of the sound and make the sound mellow and pleasant. If the reverberation time is too short, the sound will be dry and unnatural. If the reverberation time is too long, the sound will become unclear.
[0232] That is to say, when different reverberation times are selected, the sound effects of the room sizes simulated by the target audio signal finally played by the headphone device are different, so that the externalization degree of the simulated sound is also different. Therefore, the electronic device can process the third BRIR by controlling different reverberation times to simulate the sound effects of different rooms to adjust different externalization effects of the target audio signal. When the selected reverberation time is longer, it can be used to simulate the sound effects of a larger room. When the selected reverberation time is shorter, it can be used to simulate the sound effects of a smaller room.
[0233] When recording the audio signal to be rendered, the mixing duration caused by room reflection during recording can be called the original reverberation time. The electronic device can select a preset reverberation time to process the third BRIR, and this reverberation time is less than the original reverberation time. For example, if the original reverberation time is 200 ms, then this reverberation time can be selected as 50 ms.
[0234] It should be noted that the reverberation time for adjusting the third BRIR can be a preset fixed value. For example, the reverberation time can be selected as 50 ms. Alternatively, the reverberation time for adjusting the third BRIR can also be set by the user. For example, a reverberation time adjustment control is displayed on the interface of the electronic device, and the user performs a touch operation on the reverberation time adjustment control to adjust the reverberation time. Subsequently, when performing spatial audio rendering, the electronic device can use the adjusted reverberation time to process the third BRIR to generate the first BRIR.
[0235] In another way, the electronic device can process the third BRIR with the second gain to control the externalization effect achieved when the target audio signal is played through the headphone device. Among them, the ratio of the first gain to the second gain is called the direct-reverberant ratio control parameter, that is, the direct-reverberant ratio control parameter K = g1 / g2, where g1 is the first gain and g2 is the second gain. For example, the first gain can be 0.8 and the second gain can be 0.9, then the direct-reverberant ratio control parameter is 8 / 9; or, the first gain can be 0.6 and the second gain can be 0.9, then the direct-reverberant ratio control parameter is 2 / 3.
[0236] Specifically, when the electronic device processes the third BRIR with the second gain, it multiplies the amplitude of the third BRIR by the second gain to obtain the first BRIR.
[0237] It can be understood that the first gain is used to adjust the amplitude of the direct sound part in the subsequently generated second BRIR, and the second gain is used to adjust the amplitude of the reverberation part in the subsequently generated second BRIR.
[0238] In the embodiments of the present application, the electronic device can adjust the seventh HRTF and the eighth HRTF with the first gain to control the externalization effect achieved when the target audio signal is played through the headphone device; the electronic device can also adjust the third BRIR with the second gain to control the externalization effect achieved when the target audio signal is played through the headphone device; the electronic device can also control the externalization effect achieved when the target audio signal is played through the headphone device by adjusting the first gain and the second gain simultaneously.
[0239] Generally, if the second gain is set to be smaller, the amplitude of the reverberation part in the subsequently generated second BRIR decays. Correspondingly, that is, the amplitude of the direct sound part in the subsequently generated second BRIR is relatively larger, making the direct-reverberant ratio control parameter larger. As a result, when the finally generated target audio signal is played through the headphone device, the simulated sound effect is that of a relatively close sound source, without much sense of externalization.
[0240] If the second gain is set to be relatively large, the amplitude of the reverberation part in the subsequently generated second BRIR increases. Correspondingly, the amplitude of the direct sound part in the subsequently generated second BRIR becomes relatively small, resulting in a relatively small direct-to-reverberation ratio control parameter. As a result, when the finally generated target audio signal is played through a headphone device, the simulated sound effect is that of a relatively distant sound source, thereby simulating a more obvious externalization effect.
[0241] Therefore, the electronic device can control the obviousness of the externalization effect of the target audio signal when played through a headphone device by controlling the direct-to-reverberation ratio control parameter.
[0242] It should be noted that the above-mentioned direct-to-reverberation ratio control parameter can be a preset fixed value. For example, the direct-to-reverberation ratio control parameter can be selected as 8 / 9. Alternatively, the above-mentioned direct-to-reverberation ratio control parameter can also be set by the user. For example, a direct-to-reverberation ratio control parameter adjustment control is displayed on the interface of the electronic device, and the user performs a touch operation on the direct-to-reverberation ratio control parameter adjustment control to adjust the direct-to-reverberation ratio control parameter. Subsequently, when performing spatial audio rendering, the electronic device can use the adjusted direct-to-reverberation ratio control parameter to process the third BRIR to generate the first BRIR.
[0243] It can be understood that the electronic device may not process the third BRIR using the reverberation parameter, but directly use the third BRIR as the first BRIR.
[0244] It should be noted that the seventh HRTF, eighth HRTF, first HRTF, second HRTF, fourth HRTF, fifth HRTF, sixth HRTF, and third HRTF generated in the above process can refer to parameters represented in the time domain; or, the seventh HRTF, eighth HRTF, first HRTF, second HRTF, fourth HRTF, fifth HRTF, sixth HRTF, and third HRTF generated in the above process can also refer to parameters represented in the frequency domain. The electronic device can convert the third HRTF represented in the frequency domain to the time domain, and then splice the third HRTF represented in the time domain with the first BRIR to obtain the second BRIR.
[0245] S913, the electronic device splices the third HRTF with the first BRIR to obtain the second BRIR.
[0246] S914, the electronic device renders the audio signal to be rendered using the second BRIR to obtain the target audio signal.
[0247] In summary, in the audio processing method according to the embodiments of the present application, since only one convolution operation is required to generate the target audio signal, the computational complexity and the memory occupied by the computation are reduced, the speed of the electronic device for real-time rendering of the target audio signal is increased, and thus the link delay of sending the target audio signal between the electronic device and the headphone device is reduced. In this way, when the headphone device performs head movement tracking, the played target audio signal can be quickly switched according to the first head movement tracking data and the second head movement tracking data, thereby improving the real-time response effect of the audio playback when the headphone device performs head movement tracking.
[0248] Moreover, since the generated third HRTF includes the second interaural time difference of the target user, the second BRIR used for rendering the audio signal to be rendered also includes the second interaural time difference of the target user, and the second interaural time difference is related to the actual head parameters of the target user. Therefore, the target audio signal generated by rendering the audio signal to be rendered using the second BRIR can meet the personalized needs of the target user, improve the immersion of the target user when the headphone device plays the target audio signal, enhance the accuracy of sound image localization, and control the externalization degree of the sound image to be more in line with the actual externalization degree required by the target user.
[0249] In addition, by first performing parameter switching according to the fourth HRTF and the fifth HRTF and then adding the second interaural time difference of the target user, the spatial localization accuracy of the finally rendered target audio signal can be improved.
[0250] In some implementable ways, as Figure 16 shown, a spatial audio rendering control 161 and a head movement tracking control 162 are provided on the third interface of the electronic device. The spatial audio rendering control 161 is used to control whether to enable the spatial audio rendering function, and the head movement tracking control 162 is used to control whether to perform audio rendering according to the head movement tracking data collected by the headphone device. The functions corresponding to the spatial audio rendering control 161 and the head movement tracking control 162 can be functions in the headphone settings application.
[0251] When the electronic device receives an enabling operation from the user for the spatial audio rendering control 161 and the head movement tracking control 162, the electronic device can execute Figure 7 or Figure 9Corresponding embodiments. When the electronic device receives an operation to turn on the spatial audio rendering control 161 and an operation to turn off the head movement tracking control 162 from the user, the headphone device no longer collects head movement tracking data in real time or periodically, and the electronic device also does not perform spatial audio rendering on the audio signal to be rendered based on the head movement tracking data collected by the headphone device. At this time, the electronic device can perform spatial audio rendering based on the default head movement tracking data. Exemplarily, the electronic device can remove the first interaural time difference in the HRTF corresponding to the default head movement tracking data, and then sequentially execute the above steps S910 and the subsequent steps. When the electronic device receives an operation to turn off both the spatial audio rendering control 161 and the head movement tracking control 162 from the user, the electronic device does not execute the audio processing method of the embodiments of the present application.
[0252] In addition, as Figure 16 shown, a head radius setting control 163 and an ear contour setting control 164 are also displayed on the third interface of the electronic device.
[0253] When the electronic device receives a touch operation on the head radius setting control 163 from the user, the electronic device can jump to the first interface as shown in (a) of Figure 11 , or the electronic device can jump to the second interface as shown in (a) of Figure 12 .
[0254] In some embodiments, in addition to the head radius, the ear contour also affects the accuracy of sound image localization among the factors affecting sound image localization.
[0255] When sound waves reach the auricle, a part of the sound waves directly enter the ear canal, and another part enters the ear canal after being reflected by the auricle. There will be phenomena of in-phase addition and anti-phase subtraction among the reflected waves that are refracted by the auricle towards the ear canal, thus forming peaks and valleys in the frequency spectrum. Therefore, different auricle shapes correspond to different sound wave reflections, resulting in different peak and valley values in the high-frequency signal. The auricle shape plays an important role in the vertical localization of the sound source.
[0256] When the electronic device receives a touch operation on the ear contour setting control 164 from the user, the electronic device can jump to the fourth interface as shown in Figure 17 . The fourth interface displays controls corresponding to multiple different auricle features, such as auricle feature 701, auricle feature 702, and auricle feature 703, etc.
[0257] When the electronic device receives a touch operation on any auricle feature from the user, if the pitch angle in the first head movement tracking data is not equal to 0, when using Figure 7 or Figure 9When generating the target audio signal corresponding to the implementation manner, an equalizer (EQ) can be used to process the target audio signal and filter the signals at specific frequency points in the target audio signal to a certain extent. For example, the valley value of the signal at a preset frequency point in the target audio signal is reduced by a preset amplitude. Finally, the filtered target audio signal is sent to the headphone device for playback.
[0258] When the auricle features touched by the user are different, the above-mentioned preset frequency points and preset amplitudes are different. For example, when the auricle feature touched by the user is auricle feature 701, the corresponding preset frequency point is 4.2KHz and the preset amplitude is 11.8dB; when the auricle feature touched by the user is auricle feature 702, the corresponding preset frequency point is 9.6KHz and the preset amplitude is 4.8dB.
[0259] In addition, Figure 17 The fourth interface shown may also include a prompt message. For example, the prompt message is: If you do not select any auricle feature, a default auricle feature will be configured for you, and the default auricle feature can be any one of auricle feature 701, auricle feature 702, and auricle feature 703.
[0260] In some scenarios, after target user 1 wears the headphone device for a period of time, the headphone device can be handed over to target user 2, and target user 2 wears the headphone device to continue listening to the audio signal. Therefore, in this scenario, when target user 2 wears the headphone device for a certain duration (such as 30s), or when target user 2's head turns a certain angle in a certain direction after wearing the headphone device, the headphone device can send an indication message for triggering the recalculation of the second interaural time difference to the electronic device to instruct the electronic device to recalculate the second interaural time difference of target user 2.
[0261] A possible implementation manner is that after the electronic device receives the indication message for triggering the recalculation of the second interaural time difference sent by the headphone device, the electronic device can display the fifth interface as shown in Figure 18 The fifth interface displays a prompt message. For example, the prompt message is: It is detected that the user wearing the headphones has changed. Please confirm whether to use the head radius of the current user wearing the headphones for audio rendering.
[0262] In addition, the fifth interface also displays a confirmation control 181 and a cancellation control 182. When the electronic device receives a touch operation on the confirmation control 181, the second interaural time difference is recalculated according to the head radius of the current user wearing the headphone device (such as target user 2) for spatial audio rendering; when the electronic device receives a touch operation on the cancellation control 182, the spatial audio rendering is still performed according to the second interaural time difference corresponding to the head radius of target user 1.
[0263] In some embodiments, the processes performed by the above S901, S902, S906 and S907 can be implemented through a customized HRTF module; the processes performed by the above S903, S904, S905, S908, S909 and S910 can be implemented through a parameter switching module; the processes performed by the above S911, S912, S913 and S914 can be implemented through a reverberation control module.
[0264] The above-mentioned customized HRTF module, parameter switching module and reverberation control module can all be integrated into an electronic device; the above-mentioned customized HRTF module, parameter switching module and reverberation control module can also all be integrated into a headphone device, that is, the steps shown in S901 to S914 can be performed in the headphone device; part of the above-mentioned customized HRTF module, parameter switching module and reverberation control module can be integrated into an electronic device, while the other part is integrated in the headphone device, that is, part of the steps shown in S901 to S914 can be performed in the electronic device, while the other part is performed in the headphone device. This is not limited in the embodiments of the present application.
[0265] Figure 19 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 100 includes: a memory 1901, a processor 110, and a communication interface 1902, wherein the memory 1901, the processor 110, and the communication interface 1902 can communicate; illustratively, the memory 1901, the processor 110, and the communication interface 1902 can communicate via a communication bus.
[0266] The memory 1901 may be a read only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM). The memory 1901 may store a computer program, which is controlled and executed by the processor 110 and communicated by the communication interface 1902, thereby implementing the audio processing method provided in the above embodiment of the present application.
[0267] The communication interface 1902 in the chip may be an input / output interface, a pin or a circuit, etc.
[0268] The electronic device 100 of this embodiment can be used to execute the steps executed in the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0269] The embodiments of the present application also provide a computer-readable storage medium. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over a computer-readable medium as one or more instructions or codes. The computer-readable medium can include a computer storage medium and a communication medium, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0270] In a possible implementation, the computer-readable medium can include RAM, ROM, a compact disc read-only memory (CD-ROM), or other optical disc storage, a magnetic disk storage, or other magnetic storage device, or any other medium targeted to carry or store the required program code in the form of instructions or data structures and accessible by a computer. Moreover, any connection is properly termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include optical discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. The above combinations should also be included within the scope of the computer-readable medium.
[0271] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0272] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0273] The above specific implementation manners further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above are only specific implementation manners of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present application should be included in the protection scope of the present application.
Claims
1. An audio processing method, characterized in that, Including: The electronic device obtains first head movement tracking data and second head movement tracking data; The second head movement tracking data is the head movement tracking data obtained in the previous time of the first head movement tracking data; The electronic device obtains a first head-related transfer function HRTF according to the first head movement tracking data, and obtains a second HRTF according to the second head movement tracking data; When the first head movement tracking data is different from the second head movement tracking data, the electronic device performs parameter switching according to the first HRTF and the second HRTF to generate a third HRTF; The electronic device splices the third HRTF represented in the time domain with a first binaural room impulse response BRIR to obtain a second BRIR; the first BRIR refers to the BRIR obtained by adjusting the reverberation part in the original BRIR with reverberation parameters, or the reverberation part in the original BRIR; The electronic device renders the audio signal to be rendered by using the second BRIR to obtain a target audio signal.
2. The method according to claim 1, characterized in that The electronic device performs parameter switching according to the first HRTF and the second HRTF to generate a third HRTF, including: The electronic device removes a first interaural time difference in the first HRTF to obtain a fourth HRTF, and removes the first interaural time difference in the second HRTF to obtain a fifth HRTF; The electronic device generates a sixth HRTF according to the fourth HRTF and the fifth HRTF; The electronic device obtains a second interaural time difference of a target user; the second interaural time difference is related to the head parameters of the target user; The electronic device adds the second interaural time difference to the sixth HRTF to obtain the third HRTF.
3. The method according to claim 1, wherein After the electronic device obtains a first HRTF according to the first head movement tracking data and obtains a second HRTF according to the second head movement tracking data, it further includes: When the first head movement tracking data is the same as the second head movement tracking data, the electronic device removes a first interaural time difference in a first target HRTF to obtain a second target HRTF; The electronic device obtains a second interaural time difference of a target user; the second interaural time difference is related to the head parameters of the target user; The electronic device adds the second interaural time difference to the second target HRTF to obtain the third HRTF; Wherein, the first target HRTF is the first HRTF, and the second target HRTF is the fourth HRTF; or, the first target HRTF is the second HRTF, and the second target HRTF is the fifth HRTF.
4. The method according to claim 2 or 3, characterized in that, The electronic device obtains a second interaural time difference of a target user, including: The electronic device obtains the head parameters of the target user; The electronic device obtains the second interaural time difference of the target user according to the head parameters and the horizontal azimuth angle in the first head movement tracking data.
5. The method according to claim 4, characterized in that The electronic device obtains the head parameters of the target user, including: The electronic device displays a first interface; the first interface includes a head parameter input box; The electronic device receives the head parameters input into the head parameter input box.
6. The method according to claim 4, characterized in that The head parameter is the head radius; the electronic device obtains the second interaural time difference of the target user according to the head parameter and the horizontal azimuth angle in the first head movement tracking data, including: When the electronic device is in the far-field mode, the electronic device calculates the second interaural time difference of the target user through the following formula: ; When the electronic device is in the near-field mode, the electronic device calculates the second interaural time difference of the target user through the following formula: ; Wherein, ITD_new is the second interaural time difference, and a is the head radius is the horizontal azimuth angle in the first head movement tracking data, and c is the speed of sound.
7. The method according to claim 2, wherein The sixth HRTF is obtained through the following formula: ; Where HRTF4 is the fourth HRTF, HRTF5 is the fifth HRTF, b is the sequence number of sampling points, N is the total number of sampling points, and HRTF_new is the sixth HRTF.
8. The method according to claim 1, characterized in that The electronic device obtains a first HRTF according to the first head movement tracking data and a second HRTF according to the second head movement tracking data, including: The electronic device obtains the seventh HRTF corresponding to the first head movement tracking data and the eighth HRTF corresponding to the second head movement tracking data from the HRTF database; The electronic device adjusts the seventh HRTF with a first gain to obtain the first HRTF, and adjusts the eighth HRTF with the first gain to obtain the second HRTF.
9. The method according to claim 1, characterized in that, Before the electronic device splices the third HRTF and the first BRIR to obtain the second BRIR, it further includes: The electronic device extracts the reverberation part from the original BRIR to obtain a third BRIR; The electronic device determines the first BRIR according to the third BRIR.
10. The method according to claim 9, characterized in that, The electronic device determines the first BRIR according to the third BRIR, including: The electronic device processes the third BRIR with reverberation parameters to obtain the first BRIR; the reverberation parameters include reverberation time and / or a second gain.
11. The method according to claim 1, characterized in that, The electronic device splices the third HRTF and the first BRIR to obtain the second BRIR, including: The electronic device splices the first BRIR after the third HRTF to obtain the second BRIR.
12. The method according to claim 1, wherein The electronic device is communicatively connected to the headphone device; the electronic device obtains first head movement tracking data and second head movement tracking data, including: The electronic device receives the first head movement tracking data and the second head movement tracking data sent by the headphone device; After the electronic device renders the audio signal to be rendered with the second BRIR to obtain a target audio signal, it further includes: The electronic device sends the target audio signal to the headphone device.
13. An electronic device, characterized in that, It includes a memory and a processor, the memory is used to store a computer program, and the processor is used to call the computer program to execute the audio processing method according to any one of claims 1 to 12.
14. A computer-readable storage medium, characterized in that, A computer program or instruction is stored in the computer-readable storage medium, and when the computer program or instruction is run, the audio processing method according to any one of claims 1 to 12 is implemented.
15. A computer program product, characterized in that, Comprising a computer program which, when run, causes a computer to execute the audio processing method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Headtracking for pre-rendered binaural audio
CN112954582A
An active monitoring headphone and a binaural method for the same
US20190130927A1