Head related transfer function generation method and apparatus

By acquiring audio signals through the relative motion of electronic devices and head-mounted devices, and combining them with location information to generate personalized HRTFs, the problem of convenient HRTF generation is solved, and realistic three-dimensional sound effects are achieved on small devices.

WO2026066193A1PCT designated stage Publication Date: 2026-04-02HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing technologies make it difficult to easily generate personalized head-related transfer functions (HRTFs), resulting in the inability to achieve realistic 3D sound effects on small mobile devices such as wireless headphones or smart glasses.

Method used

By using human-computer interaction and the relative motion between electronic devices and head-mounted devices, test audio signals and feedback audio signals are collected and combined with location information to generate personalized HRTFs.

Benefits of technology

It improves the accuracy of HRTF and achieves realistic 3D sound effects on small mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025098051_02042026_PF_FP_ABST
    Figure CN2025098051_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A head related transfer function (HRTF) generation method and apparatus, relating to the technical field of audio processing. The method comprises: an electronic device outputting first prompt information, the first prompt information being used for instructing a target individual wearing a head-mounted device to perform a first operation, and the first operation causing relative movement between the electronic device and the head-mounted device; playing test audio signals during the first operation, and receiving feedback audio signals and position information from the head-mounted device, wherein the feedback audio signals are test audio signals received by the head-mounted device at a plurality of different spatial positions, the spatial positions are the spatial positions of the electronic device relative to the head-mounted device, and the position information is used for determining the spatial positions; and on the basis of the test audio signals, the feedback audio signals and the position information, determining an HRTF corresponding to the target individual.
Need to check novelty before this filing date? Find Prior Art

Description

A head related transfer function generation method and device

[0001] Cross-reference to related applications

[0002] The present application claims priority to the Chinese patent application No. 202411368586.7, filed on September 27, 2024, and entitled "A head related transfer function generation method and device", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0003] The present application relates to the field of audio processing, and in particular to a head related transfer function generation method and device. BACKGROUND

[0004] To meet the needs of immersive sound experience, traditional stereo sound effects can present three-dimensional sound effects based on multi-channel technology. Although the traditional stereo sound effects have good three-dimensional sound presentation effect, they cannot be applied to small mobile devices such as wireless earphones or smart glasses due to the need to configure multiple loudspeakers.

[0005] A person can distinguish the direction of a sound source in space through both ears, and the principle lies in that the human brain can distinguish the direction of the sound source through the subtle differences in sound between both ears. The head related transfer function (HRTF) can describe the changes that occur when sound is reflected and diffracted by human body parts such as pinna, head, and torso during the process of sound being transmitted into the human ear from a specific direction. Therefore, the use of HRTF on devices such as earphones and smart glasses can simulate the sound information heard by both ears of a person, thereby presenting a realistic three-dimensional sound effect.

[0006] HRTF is closely related to individuals, and the HRTF of different individuals has great differences. How to conveniently generate personalized HRTF is a problem to be solved at present. SUMMARY

[0007] The embodiments of the present application provide a head related transfer function generation method and device to conveniently generate personalized head related transfer function based on human-computer interaction mode.

[0008] In a first aspect, a method for generating HRTF is provided. The method can be applied to an electronic device, such as a mobile phone, a tablet computer, a notebook computer, and the like. The method can also be applied to a wearable device, such as a watch, a bracelet, and the like. The method can also be applied to a vehicle-mounted device, and the like. The method comprises the following steps: outputting first prompt information, the first prompt information being used to instruct a target individual wearing a head-mounted device to perform a first operation, the first operation causing a relative movement between the electronic device and the head-mounted device; playing a test audio signal during the first operation, and receiving a feedback audio signal and position information from the head-mounted device; wherein the feedback audio signal is the test audio signal received by the head-mounted device at a plurality of different spatial positions, the spatial positions being spatial positions of the electronic device relative to the head-mounted device, and the position information being used to determine the spatial positions; and determining an HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information.

[0009] The above implementation manner can generate an HRTF by using an electronic device and a head-mounted device based on human-computer interaction when the target individual wears the head-mounted device, thereby conveniently generating a personalized HRTF.

[0010] In a possible implementation manner, the first operation is that the target individual holds the electronic device and rotates the electronic device around the head while keeping the head still, or the first operation is that the target individual rotates the head while keeping the electronic device held still.

[0011] In a possible implementation manner, the first prompt information is further used to instruct one or more of the following: a distance between the electronic device and the head-mounted device, a rotation direction, a rotation angle, a rotation range, and a duration.

[0012] In a possible implementation manner, the outputting of the first prompt information comprises: outputting the first prompt information after establishing a communication connection between the electronic device and the head-mounted device, the communication connection being used to transmit the feedback audio signal and the position information; or outputting the first prompt information after starting a first application program, the first application program being used to obtain an HRTF; or outputting the first prompt information in response to a first voice instruction, the first voice instruction being used to trigger the obtaining of the HRTF; or outputting the first prompt information when it is detected that a wearing state of the head-mounted device is wearing.

[0013] In a possible implementation manner, before the first prompt information is output, the method further includes: outputting second prompt information, the second prompt information being used to instruct the target individual to wear the head-mounted device; and the outputting of the first prompt information includes: outputting the first prompt information when it is detected that the wearing state of the head-mounted device is wearing.

[0014] In a possible implementation manner, the method further includes: ending the playing of the test audio signal after a set time length; or ending the playing of the audio signal based on a first user operation; or ending the playing of the test audio signal in response to a second voice instruction.

[0015] In a possible implementation manner, the method further includes: outputting third prompt information, the third prompt information being used to instruct the target individual to end the first operation and / or being used to indicate that the HRTF has been generated.

[0016] In a possible implementation manner, the outputting of the first prompt information includes one or more of the following: playing voice prompt information; or displaying text prompt information; or displaying an image or an animation, the image or the animation being used to show a manner in which relative motion is generated between the electronic device and the head-mounted device.

[0017] In a possible implementation manner, the method further includes: determining a spatial position of the electronic device relative to the head-mounted device according to position information of the electronic device and position information from the head-mounted device; wherein the position information of the head-mounted device is detected by a spatial attitude sensor in the head-mounted device, and the position information of the electronic device is detected by a spatial attitude sensor in the electronic device.

[0018] In a possible implementation manner, the feedback audio signal includes a first feedback audio signal collected by a first microphone of the head-mounted device and a second feedback audio signal collected by a second microphone of the head-mounted device; and the determining of the HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information includes: determining the HRTF corresponding to the target individual according to the test audio signal, the first feedback audio signal, the second feedback audio signal, and the position information.

[0019] In the implementation manners described above, the feedback audio signal of the test audio signal can be collected by using multiple microphones in the head-mounted device, and the HRTF can be generated based on the collected feedback audio signal, so that the accuracy of the HRTF can be improved, and the audio listening experience of the user can be improved.

[0020] In a possible implementation, the spatial position comprises a first spatial position; and the determining the HRTF corresponding to the target individual according to the test audio signal, the first feedback audio signal and the second feedback audio signal, and the spatial position comprises: determining a first HRTF corresponding to the first spatial position according to the test audio signal played by the electronic device and the first feedback audio signal collected by the first microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; determining a second HRTF corresponding to the first spatial position according to the test audio signal played by the electronic device and the second feedback audio signal collected by the second microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; and correcting the first HRTF based on the second HRTF to obtain the HRTF corresponding to the first spatial position.

[0021] In a possible implementation, after the HRTF corresponding to the target individual is determined, one or more of the following operations are further performed: correcting the HRTF according to a characteristic parameter of the head-mounted device; or correcting the HRTF according to a human body feature of the target individual; or correcting the HRTF according to an HRTF database.

[0022] In the implementations described above, the calculated HRTF can be further corrected according to the characteristic parameter of the head-mounted device, the human body feature of the target individual, or the HRTF database, so that the accuracy of the HRTF can be improved.

[0023] In a possible implementation, the correcting the HRTF according to the human body feature of the target individual comprises: acquiring an image of the target individual object collected by the electronic device; identifying the image to obtain an auricle structure feature and / or a head structure feature of the target individual; and correcting the HRTF according to the auricle structure feature and / or the head structure feature of the target individual.

[0024] In a possible implementation, the spatial position of the electronic device relative to the head-mounted device comprises a first spatial position, and the HRTF determined according to the test audio signal, the feedback audio signal and the position information comprises a first HRTF corresponding to the first spatial position; and the correcting the HRTF according to the HRTF database comprises: acquiring, from the HRTF database, a second HRTF matching the human body feature information and the first spatial position according to the human body feature information of the target individual and the first spatial position; and correcting the first HRTF according to the second HRTF.

[0025] In a possible implementation manner, the HRTF determined according to the test audio signal, the feedback audio signal, and the position information includes HRTFs corresponding to at least two spatial positions in a first spatial range; and the method further includes: acquiring, from an HRTF database, HRTFs corresponding to at least one spatial position in a second spatial range that matches a human body feature of the target individual according to the human body feature; and determining the HRTFs corresponding to the at least two spatial positions in the first spatial range and the HRTFs corresponding to the at least one spatial position in the second spatial range as the HRTF corresponding to the target individual.

[0026] In a second aspect, an HRTF generation method is provided, which can be applied to a head-mounted device, for example, a headset, or smart glasses, or a virtual reality (VR) helmet or VR glasses, or an augmented reality (AR) helmet or AR glasses, or a mixed reality (MR) helmet or MR glasses, etc. The method can include the following steps: collecting a feedback audio signal, the feedback audio signal being a test audio signal played by an electronic device and received by the head-mounted device at a plurality of different spatial positions, the spatial positions being spatial positions of the electronic device relative to the head-mounted device; receiving position information from the electronic device, the position information being used to determine the spatial positions; and determining an HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information.

[0027] In a possible implementation manner, the method further includes: determining the spatial positions of the electronic device relative to the head-mounted device according to position information of the head-mounted device and position information from the electronic device, wherein the position information of the head-mounted device is detected by a spatial attitude sensor in the head-mounted device, and the position information of the electronic device is detected by a spatial attitude sensor in the electronic device.

[0028] In a possible implementation manner, the feedback audio signal includes a first feedback audio signal collected by a first microphone of the head-mounted device and a second feedback audio signal collected by a second microphone of the head-mounted device; and the determining of the HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information includes: determining the HRTF corresponding to the target individual according to the test audio signal, the first feedback audio signal, the second feedback audio signal, and the position information.

[0029] In a possible implementation, the spatial position includes a first spatial position; and the determining the HRTF corresponding to the target individual according to the test audio signal, the first feedback audio signal and the second feedback audio signal, and the spatial position includes: determining a first HRTF corresponding to the first spatial position according to the test audio signal played by the electronic device and the first feedback audio signal collected by the first microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; determining a second HRTF corresponding to the first spatial position according to the test audio signal played by the electronic device and the second feedback audio signal collected by the second microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; and correcting the first HRTF based on the second HRTF to obtain the HRTF corresponding to the first spatial position.

[0030] In a possible implementation, after the HRTF corresponding to the target individual is determined, the method further includes one or more of the following: correcting the HRTF according to a characteristic parameter of the head-mounted device; or correcting the HRTF according to a human body feature of the target individual; or correcting the HRTF according to an HRTF database.

[0031] In a possible implementation, the spatial position of the electronic device relative to the head-mounted device includes a first spatial position, and the HRTF determined according to the test audio signal, the feedback audio signal and the position information includes a first HRTF corresponding to the first spatial position; and the correcting the HRTF according to the HRTF database includes: obtaining, from the HRTF database, a second HRTF matching the human body feature information of the target individual and the first spatial position according to the human body feature information of the target individual and the first spatial position; and correcting the first HRTF according to the second HRTF.

[0032] In a possible implementation, the HRTF determined according to the test audio signal, the feedback audio signal and the position information includes HRTFs corresponding to at least two spatial positions in a first spatial range; and the method further includes: obtaining, from an HRTF database, HRTFs corresponding to at least one spatial position in a second spatial range matching a human body feature of the target individual according to the human body feature of the target individual; and determining the HRTFs corresponding to the at least two spatial positions in the first spatial range and the HRTFs corresponding to the at least one spatial position in the second spatial range as the HRTF corresponding to the target individual.

[0033] In a third aspect, a method for generating HRTF is provided. The method can be applied to an electronic device. The method comprises: outputting first prompt information, the first prompt information being used to instruct a target individual wearing a head-mounted device to perform a first operation, the first operation causing relative movement between the electronic device and the head-mounted device; playing a test audio signal during the first operation and receiving a feedback audio signal and position information from the head-mounted device, wherein the feedback audio signal is the test audio signal received by the head-mounted device at a plurality of different spatial positions, the spatial positions being spatial positions of the electronic device relative to the head-mounted device, and the position information being used to determine the spatial positions; sending the test audio signal, the feedback audio signal, and information of the spatial positions to a server, and receiving an HRTF from the server, the HRTF being determined based on the test audio signal, the feedback audio signal, and the information of the spatial positions.

[0034] In a possible implementation, the first operation is that the target individual holds the electronic device and rotates the electronic device around the head while keeping the head still, or the first operation is that the target individual rotates the head while keeping the electronic device held in the hand still.

[0035] In a possible implementation, the first prompt information is further used to instruct one or more of the following: a distance between the electronic device and the head-mounted device, a rotation direction, a rotation angle, a rotation range, and a duration.

[0036] In a possible implementation, the outputting of the first prompt information comprises: outputting the first prompt information after establishing a communication connection between the electronic device and the head-mounted device, the communication connection being used to transmit the feedback audio signal and the position information; or outputting the first prompt information after starting a first application, the first application being used to obtain the HRTF; or outputting the first prompt information in response to a first voice instruction, the first voice instruction being used to trigger the obtaining of the HRTF; or outputting the first prompt information when it is detected that the head-mounted device is worn.

[0037] In a possible implementation, before the outputting of the first prompt information, the method further comprises: outputting second prompt information, the second prompt information being used to instruct the target individual to wear the head-mounted device; and the outputting of the first prompt information comprises: outputting the first prompt information when it is detected that the head-mounted device is worn.

[0038] In a possible implementation, the method further includes: ending playing the test audio signal after a set time duration; or ending playing the audio signal based on a first user operation; or ending playing the test audio signal in response to a second voice instruction.

[0039] In a possible implementation, the method further includes: outputting third prompt information, the third prompt information being used to instruct the target individual to end the first operation and / or being used to indicate that the HRTF has been generated.

[0040] In a possible implementation, the outputting the first prompt information includes one or more of the following: playing voice prompt information; or displaying text prompt information; or displaying an image or an animation, the image or the animation being used to show a manner in which relative motion is generated between the electronic device and the head-mounted device.

[0041] In a possible implementation, the method further includes: determining a spatial position of the electronic device relative to the head-mounted device according to position information of the electronic device and position information from the head-mounted device, wherein the position information of the head-mounted device is detected by a spatial attitude sensor in the head-mounted device, and the position information of the electronic device is detected by a spatial attitude sensor in the electronic device.

[0042] In a fourth aspect, a system is provided, the system including an electronic device and a head-mounted device, and the electronic device can implement the method in any one of the first aspect.

[0043] In a fifth aspect, a system is provided, the system including an electronic device and a head-mounted device, and the head-mounted device can implement the method in any one of the second aspect.

[0044] In a sixth aspect, a system is provided, the system including an electronic device, a head-mounted device, and a server, and the electronic device can implement the method in any one of the third aspect.

[0045] In a seventh aspect, an apparatus is provided, the apparatus including units or modules for performing the method in any one of the first aspect, or performing the method in any one of the second aspect, or performing the method in any one of the third aspect.

[0046] In an eighth aspect, an apparatus is provided, the apparatus including one or more processors configured to perform the method in any one of the first aspect, or perform the method in any one of the second aspect, or perform the method in any one of the third aspect.

[0047] In a ninth aspect, a readable storage medium is provided, which stores a program or instructions, when the program or instructions are run on an apparatus, cause the apparatus to perform the method in any one of the first aspect, or perform the method in any one of the second aspect, or perform the method in any one of the third aspect.

[0048] In a tenth aspect, a chip system is provided, which comprises a processor for supporting a computer device to implement the method in any one of the first aspect, or implement the method in any one of the second aspect, or implement the method in any one of the third aspect.

[0049] In an eleventh aspect, a program product is provided, which comprises a program; when the program is run on a computer, cause the computer to perform the method in any one of the first aspect, or perform the method in any one of the second aspect, or perform the method in any one of the third aspect. BRIEF DESCRIPTION OF DRAWINGS

[0050] FIG. 1 is a schematic diagram of a hardware structure of an electronic device according to an embodiment of the present application;

[0051] FIG. 2 is a schematic diagram of a software structure of an electronic device according to an embodiment of the present application;

[0052] FIG. 3 is a schematic diagram of a structure of an earphone and a wearing state of the earphone according to an embodiment of the present application;

[0053] FIG. 4 is a schematic diagram of an application scenario according to an embodiment of the present application;

[0054] FIG. 5 is an example interface for prompting a user to perform a first operation through an interface according to an embodiment of the present application;

[0055] FIG. 6 is a schematic diagram of human-computer interaction according to an embodiment of the present application;

[0056] FIG. 7 is a schematic diagram of another human-computer interaction according to an embodiment of the present application;

[0057] FIG. 8 is a schematic diagram of a head-centered spatial coordinate system according to an embodiment of the present application;

[0058] FIG. 9 is a schematic diagram of a flow of a HRTF generation method according to an embodiment of the present application;

[0059] FIG. 10 is a schematic diagram of a flow of another HRTF generation method according to an embodiment of the present application;

[0060] FIG. 11 is a schematic diagram of a flow of another HRTF generation method according to an embodiment of the present application;

[0061] FIG. 12 is a schematic diagram of another application scenario according to an embodiment of the present application;

[0062] FIG. 13 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0063] In the following, some terms in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.

[0064] At least one of the embodiments of the present application includes one or more, and the plurality refers to two or more. In addition, it should be understood that the terms "first", "second", and the like used in the description of the specification are only used to distinguish the description and cannot be understood as indicating or implying relative importance or indicating or implying order. For example, the first operation and the second operation do not represent the importance or order of the two, but only for the purpose of distinguishing the description. In the embodiments of the present application, "and / or" is only used to describe the relationship between the two, which means that there are three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper generally represents that the front and rear associated objects are a "or" relationship.

[0065] In the description of the embodiments of the present application, it should be noted that, unless otherwise explicitly specified and limited, the terms "mounting", "connecting" should be understood in a broad sense, for example, "connecting" can be detachable connection, or can be non-detachable connection; can be direct connection, or can be indirect connection through intermediate medium. The orientation terms mentioned in the embodiments of the present application, such as "upper", "lower", "left", "right", "inner", "outer" and the like, are only the direction of the drawing, therefore, the orientation terms used are for better, clearer description and understanding of the embodiments of the present application, and are not intended to indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore, it cannot be understood as a limitation of the embodiments of the present application.

[0066] In the description of the present application, the reference to "one embodiment" or "some embodiments" and the like means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in other some embodiments" and the like appearing in different places in the present application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "including but not limited to", unless otherwise specifically emphasized.

[0067] Currently, the method for obtaining HRTF needs to use special measuring equipment and experimental environment to measure the target individual (i.e. user), which is complicated.

[0068] Therefore, the embodiments of the present application provide a method for generating HRTF and related devices and systems that can implement the method, so as to generate personalized HRTF based on human-computer interaction mode. In the method, the target individual wears a head-mounted device, the electronic device plays a test audio signal, the head-mounted device collects a feedback audio signal of the test audio signal, and based on the test audio signal and the corresponding feedback audio signal, the personalized HRTF corresponding to the target individual can be obtained.

[0069] The head-mounted device described above can be worn on the head of the user, has an audio playing function and an audio signal collecting function. Optionally, it also has a function of determining HRTF according to the test audio signal and the feedback audio signal. In some embodiments, the head-mounted device can be earphones, or smart glasses, or a virtual reality (VR) helmet or VR glasses, or an augmented reality (AR) helmet or AR glasses, or a mixed reality (MR) helmet or MR glasses, etc. In general, the type of head-mounted device is not limited in the present application.

[0070] The electronic device described above has a function of playing a test audio signal, and optionally, it also has a function of determining HRTF according to the test audio signal and the feedback audio signal collected from the head-mounted device. In some embodiments, the electronic device can be a mobile phone, a tablet computer, a notebook computer, etc. portable electronic device; it can also be a watch, a bracelet, etc. wearable device; or it can also be a vehicle-mounted device, etc. In general, the type of electronic device is not limited in the present application.

[0071] FIG. 1 shows a hardware structure diagram of an electronic device according to an embodiment of the present application. As shown in FIG. 1, the electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, a space attitude sensor, etc. The space attitude sensor can be an inertial measurement unit (IMU), for example.

[0072] An IMU is a space attitude sensor for detecting a space attitude of a carrier, for example, can detect three-axis angular velocity and acceleration of the carrier, etc. An IMU is a device for measuring three-axis attitude angles (or angular rates) and accelerations of an object. Generally, an IMU includes three single-axis accelerometers and three single-axis gyroscopes. The accelerometers detect acceleration signals of an object on independent three axes of a carrier coordinate system, and the gyroscopes detect angular velocity signals of the carrier relative to a navigation coordinate system. The IMU measures angular velocity and acceleration of an object in a three-dimensional space and calculates a posture of the object based on the angular velocity and the acceleration.

[0073] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated in one or more processors. The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching instructions and executing instructions. The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can directly call from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system. The execution of the audio playing method in the embodiments of the present application can be controlled by the processor 110 or other components to complete, for example, calling the processing program of the embodiments of the present application stored in the internal memory 121, or calling the processing program of the embodiments of the present application stored in the third party device through the external memory interface 120, so that the electronic device 100 can automatically switch the audio playing mode, solving the problem of low efficiency caused by manually switching the audio playing mode in the prior art. In addition, in the embodiments of the present application, the electronic device 100 can adaptively adjust the audio playing state or the audio playing volume when it is determined that the user has the intention to switch the audio playing mode (for example, the electronic device 100 is close to or away from the user's head), for example, when switching the audio playing mode (for example, from the first audio playing mode to the second audio playing mode), pausing the playing of the audio, or gradually reducing the volume of the audio played in the first audio playing mode and gradually reducing the volume of the audio played in the second audio playing mode, so that the user experience is not affected when the audio playing mode is switched due to the sudden change of the audio playing volume, and the user's needs are adapted.

[0074] The internal memory 121 can be used to store computer executable program codes including instructions. The processor 110 performs various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system and software codes of at least one application program (e.g., an iQiyi application, a WeChat application, etc.), etc. The data storage area can store data (e.g., images, videos, etc.) generated during use of the electronic device 100, etc. In addition, the internal memory 121 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one of a magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0075] The external memory interface 120 can be used to connect an external memory card such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to perform data storage functions. For example, files such as pictures and videos can be saved in the external memory card.

[0076] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface 130, etc.

[0077] The I2C interface is a bidirectional synchronous serial bus including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can include multiple sets of I2C buses. The processor 110 can be coupled to a touch sensor 180K, a charger, a flash, a camera 193, etc. via different I2C bus interfaces, respectively.

[0078] I2S interface can be used for audio communication. In some embodiments, processor 110 can include multiple sets of I2S bus. Processor 110 can be coupled with audio module 170 through I2S bus to enable communication between processor 110 and audio module 170. In some embodiments, audio module 170 can deliver audio signals to wireless communication module 160 through I2S interface to enable the function of answering phone through Bluetooth earphone.

[0079] PCM interface can also be used for audio communication, which samples, quantizes and encodes analog signals. In some embodiments, audio module 170 can be coupled with wireless communication module 160 through PCM bus interface. In some embodiments, audio module 170 can also deliver audio signals to wireless communication module 160 through PCM interface to enable the function of playing music through Bluetooth earphone. Both I2S interface and PCM interface can be used for audio communication.

[0080] UART interface is a universal serial bus for asynchronous communication. The bus can be a bidirectional communication bus. It converts data to be transmitted between serial communication and parallel communication. In some embodiments, UART interface is usually used to connect processor 110 and wireless communication module 160. For example, processor 110 communicates with Bluetooth module in wireless communication module 160 through UART interface to enable Bluetooth function. In some embodiments, audio module 170 can deliver audio signals to wireless communication module 160 through UART interface to enable the function of playing music through Bluetooth earphone.

[0081] MIPI interface can be used to connect processor 110 and peripheral devices such as display screen 194 and camera 193. MIPI interface includes camera serial interface (CSI), display serial interface (DSI) and the like. In some embodiments, processor 110 and camera 193 communicate through CSI interface to enable the shooting function of electronic device 100. Processor 110 and display screen 194 communicate through DSI interface to enable the display function of electronic device 100.

[0082] GPIO interface can be configured by software. GPIO interface can be configured as control signal or data signal. In some embodiments, GPIO interface can be used to connect processor 110 and camera 193, display screen 194, wireless communication module 160, audio module 170, sensor module 180 and the like. GPIO interface can also be configured as I2C interface, I2S interface, UART interface, MIPI interface and the like.

[0083] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection modes or a combination of multiple interface connection modes in the above embodiments.

[0084] The electronic device 100 can realize the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc. Among them, the ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing, and converts it into an image visible to the naked eye. The ISP can also algorithmically optimize the noise and brightness of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be arranged in the camera 193.

[0085] The electronic device 100 can realize the audio function through the audio module 170, the speaker 170A, the earpiece 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, music playing, recording, etc.

[0086] The audio module 170 is used to convert digital audio signals into analog audio signals and is also used to convert analog audio signals into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or some functional modules of the audio module 170 can be arranged in the processor 110.

[0087] The speaker 170A, also known as a "loudspeaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to an external call through one or more speakers 170A in an external call scene.

[0088] In some embodiments, the speaker 170A and / or the earpiece 170B can include a single channel or multiple channels. In some embodiments, the multiple channels are used to provide the effect of stereo sound. In some other embodiments, the multiple channels can be combined. Taking a dual-channel as an example, the left channel and the right channel can play the same sound, which can increase the volume size.

[0089] The microphone 170C, also referred to as a "microphone" or a "microphone", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, a user can speak into the microphone 170C by placing his mouth close to the microphone 170C to input a sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones, which can implement a noise reduction function in addition to collecting sound signals. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones to collect sound signals, reduce noise, and also identify the source of the sound to implement a directional recording function, etc.

[0090] The earphone interface 170D is used to connect a wired earphone. The earphone interface can be a USB interface, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0091] It can be understood that the components shown in FIG. 1 do not constitute a specific limitation on the electronic device. The electronic device in the embodiments of the present application can include more or fewer components than those in FIG. 1. In addition, the combination / connection relationship between the components in FIG. 1 can also be adjusted and modified.

[0092] FIG. 2 shows a software structure diagram of an electronic device according to an embodiment of the present application. As shown in FIG. 2, the software structure of the electronic device can be a layered architecture, for example, the software can be divided into several layers, each layer has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the operating system is divided into four layers, from top to bottom, the application layer, the application framework layer (FWK), the runtime and the system library, and the kernel layer.

[0093] The application layer can include a series of application packages. As shown in FIG. 2, the application layer can include a camera, a setting, a skin module, a user interface (UI), a third-party application, etc. Among them, the third-party application can include a gallery, a calendar, a call, a map, a navigation, a wireless local area network (WLAN), Bluetooth, music, video, short message, etc.

[0094] The application framework layer provides an application programming interface (API) and a programming framework for the applications of the application layer. The application framework layer can include some pre-defined functions. As shown in FIG. 2, the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, and a notification manager.

[0095] The runtime includes a core library and a virtual machine. The runtime is responsible for scheduling and management of the operating system.

[0096] The core library contains two parts: one part is the function function that the java language needs to call, and the other part is the core library of the operating system. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java file of the application layer and the application framework layer into a binary file. The virtual machine is used to perform the management of the object life cycle, the stack management, the thread management, the security and the exception management, and the garbage collection and the like.

[0097] The system library can include a plurality of functional modules. For example: a surface manager, media libraries, a three-dimensional graphics processing library (for example: OpenGL ES), a 2D graphics engine (for example: SGL) and the like.

[0098] The surface manager is used for managing the display subsystem, and provides a fusion of 2D and 3D layers for a plurality of applications.

[0099] The media library supports a plurality of commonly used audio, video format playback and recording, and static image files and the like. The media library can support a plurality of audio and video coding formats, for example: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG and the like.

[0100] The three-dimensional graphics processing library is used for realizing three-dimensional graphics drawing, image rendering, synthesis, and layer processing and the like.

[0101] The 2D graphics engine is a drawing engine for 2D drawing.

[0102] The kernel layer is a layer between hardware and software. The kernel layer at least contains a display driver, a camera driver, an audio driver, and a sensor driver.

[0103] The hardware layer can include various sensors, such as an acceleration sensor, a gravity sensor, a touch sensor and the like.

[0104] Based on the architecture shown in FIG. 2, in one possible implementation of the present application, an application program can be installed in the electronic device, and the application program can be used to implement the functions of the HRTF generation method provided by the embodiments of the present application. For example, the first application program can be included in the application program layer of the electronic device, and the first application program can be used to implement the HRTF generation method provided by the embodiments of the present application. Correspondingly, the related layers in the operating system can provide an interface for the first application program, so that the first application program can call the system functions through the interface to generate the HRTF.

[0105] Based on the architecture shown in FIG. 2, in another possible implementation of the present application, the operating system can be enhanced to have the functions of the HRTF generation method provided by the embodiments of the present application. For example, a service module or a function module can be added to one or more of the application program framework layer, the system library layer, and the kernel layer, and the service module or the function module can be used to implement the functions of the HRTF generation method provided by the embodiments of the present application.

[0106] Taking the head-mounted device as an example, FIG. 3 shows the structure of an earphone and a schematic diagram of the wearing state of the earphone. As shown in FIG. 3, the earphone 30 for the left ear is provided with a processor 31, a loudspeaker 32, and a first microphone 33. The loudspeaker 32 and the first microphone 33 can be arranged close to the ear canal in the earphone 30.

[0107] The loudspeaker 32 can play the audio signal processed by the HRTF under the control of the processor 31, and the first microphone 33 can collect the test audio signal emitted by the electronic device.

[0108] As shown in FIG. 3, when the user wears the earphone 30, the audio signal played by the electronic device is reflected and / or diffracted by the human body parts such as the pinna, the head, and the torso of the user, and the first microphone 33 in the earphone 30 can collect the audio signal after the reflection and / or diffraction, which is referred to as the feedback audio signal in the embodiments of the present application.

[0109] Optionally, the earphone 30 can further be provided with a second microphone 34, and the second microphone 34 can be arranged on the outside of the earphone. When the user wears the earphone, the second microphone 34 is closer to the mouth of the user than the first microphone 33, so as to collect the speech signal of the user.

[0110] It should be understood that the right ear earphone and the left ear earphone have similar structures, and the description of the earphone below can be explained as the left ear earphone or the right ear earphone, or as the left and right ear earphones.

[0111] It should be understood that the right ear earphone and the left ear earphone have similar structures, and the description of the earphone below can be explained as the left ear earphone or the right ear earphone, or as the left and right ear earphones.

[0112] In addition, other devices or modules, such as a communication module, a battery module, etc., can also be arranged in the earphone.

[0113] It should be understood that, although the above is described by taking an in-ear earphone as an example, the present application can also be applied to a semi-in-ear earphone or an open earphone, or a similar head-mounted device.

[0114] Referring to FIG. 4, it is a schematic diagram of an application scenario in an embodiment of the present application. The head 301 of a user wears an earphone 302, and the left earphone and the right earphone of the earphone 302 are respectively provided with independent microphones (also referred to as microphones). The user holds a mobile phone 303, and the distance between the mobile phone 303 and the head 301 of the user is approximately the length of an arm, such as about 40-50 cm. A first communication connection is established between the mobile phone 303 and the earphone 302. The first communication connection can be a wired connection or a wireless connection, and the wireless connection can be a short-range communication connection, such as a Bluetooth connection, a Wi-Fi connection, or a near link connection, etc., which is not limited in the present application.

[0115] The HRTF generation process can include a data acquisition stage and a calculation stage. In the data acquisition stage, the earphone 302 worn on the head 301 can acquire the feedback audio signal of the test audio signal emitted by the mobile phone 303, and in the calculation stage, the HRTF can be calculated according to the acquired data.

[0116] As shown in (a) of FIG. 4, in the data acquisition stage, the head 301 of the user remains stationary, and the user holds the mobile phone 302 and slowly moves around the head 301. The moving direction can be approximately in the horizontal direction. In order to obtain more acquisition data in more positions to improve the accuracy of the HRTF, the moving direction can include the horizontal direction and the vertical direction. The moving range in the horizontal direction can be the maximum turning range that the user can reach, and can even include the rear of the user's head. For example, taking the head 301 as the origin of the coordinate system and the front direction of the head 301 as the positive direction of the coordinate axis, the user can be required to have a moving range of not less than [-60 degrees, 60 degrees] in the horizontal direction. The turning range in the vertical direction can be the maximum turning range that the user can reach. For example, taking the head 301 of the user as the origin of the coordinate system and the front direction of the head 301 as the positive direction of the coordinate axis, the user can be required to have a turning range of not less than [-60 degrees, 60 degrees] in the vertical direction.

[0117] In the data acquisition stage, the speaker of the mobile phone 303 plays a test audio signal, and the microphone (also referred to as a microphone) of the earphone 302 acquires the feedback signal of the test audio signal. After the earphone 302 performs analog-to-digital conversion on the feedback signal, the digital information of the feedback signal is sent to the mobile phone 303 through the first communication connection between the earphone 302 and the mobile phone 303.

[0118] In the calculation stage, the mobile phone 303 generates the HRTF corresponding to the earphone 302 according to the test audio signal and the feedback audio signal from the earphone 302.

[0119] During the data collection stage, the calculation stage can be performed, or the calculation stage can be performed after the data collection stage.

[0120] The specific implementation process of the above-mentioned data collection stage and calculation stage can refer to the flow shown in FIG. 9.

[0121] The HRTF is a set of filters that uses interaural time delay (HDITD), interaural amplitude difference (IAD), and pinna frequency vibration to produce a stereo sound effect, so that when the sound is transmitted to the pinna, ear canal and eardrum in the human ear, the listener will have a surround sound effect. Through the digital signal processor (DSP) in the earphone, the HRTF can process the sound source in real time.

[0122] After generating the HRTF, when the user uses the earphone 302 to listen to the audio played by the mobile phone 303, the mobile phone 303 can process the audio signal by using the HRTF corresponding to the earphone 302 stored in the mobile phone 303, and play the processed audio signal, so that the audio signal received by the earphone 302 can present a realistic three-dimensional sound effect to the user. In another possible manner, the mobile phone 303 sends the generated HRTF to the earphone 302 through the first communication connection, and the speaker of the earphone 302 can process the received audio signal by using the HRTF when playing the audio signal, so as to simulate the sound information heard by the human ears, and present a realistic three-dimensional sound effect to the user.

[0123] In another possible application scenario shown in (b) of FIG. 4, different from the application scenario shown in (a) of FIG. 4 is that, in the data collection stage, the user holds the mobile phone 303 and keeps still, and the head 301 of the user slowly rotates, for example, can rotate left and right or up and down, or randomly. The rotation range can be the maximum rotation range that the user can reach. For example, taking the head 301 of the user as the origin of the coordinate system and the front direction of the head 301 as the positive direction of the coordinate axis, the user can be required to rotate the head within a range of [-60 degrees, 60 degrees]. The specific implementation manner of the data calculation stage in this scenario is the same as the implementation manner of the above-mentioned data calculation stage.

[0124] In another possible application scenario, different from the above application scenario, in the calculation phase, the microphone of the earphone 302 collects the feedback audio signal of the test audio signal, and then calculates the HRTF according to the test audio signal and the feedback audio signal. Optionally, the frequency response range of the test audio signal can be pre-set, or the mobile phone 303 can send the digital information of the test audio signal to the earphone 302 through the first communication connection. After the HRTF is generated, when the user uses the earphone 302 to listen to the audio played by the mobile phone 303, the earphone 302 can process the received audio signal by using the generated HRTF, so as to simulate the sound information heard by human ears and present realistic three-dimensional sound effects to the user.

[0125] In another possible application scenario, different from the above application scenario, in the calculation phase, the mobile phone 303 sends the feedback audio signal from the earphone 302 to the cloud platform (or cloud server), and the cloud platform generates the HRTF according to the test audio signal and the feedback audio signal, and sends the HRTF to the mobile phone 303. Optionally, the mobile phone 303 can send the HRTF to the earphone 302.

[0126] Optionally, the frequency response range of the test audio signal can be pre-set, or the mobile phone 303 can send the test audio signal to the cloud platform.

[0127] Taking the scenario shown in FIG. 4 as an example, in order to obtain the HRTF at multiple spatial positions based on the movement of the mobile phone 303 relative to the head 301 (or the earphone 302), in the embodiment of the present application, the electronic device can output first prompt information, the first prompt information prompting the user to perform a first operation, and the first operation can cause relative movement between the mobile phone 303 and the earphone 302.

[0128] For example, the first prompt information can prompt the user to perform the first operation as shown in (a) of FIG. 4: the user rotates the mobile phone 303 around the head 301 while keeping the head still (i.e., keeping the earphone 302 still).

[0129] For another example, the first prompt information can prompt the user to perform the first operation as shown in (b) of FIG. 4: the user rotates the head 301 while keeping the mobile phone 303 still.

[0130] In a possible implementation, the first prompt information is further used to indicate one or more of the following:

[0131] The distance between the mobile phone 303 and the earphone 302 (or the head 301). For example, the first prompt information can prompt that the distance between the mobile phone 303 and the head 301 is about one arm.

[0132] - rotation direction. For example, in the scenario shown in (a) of FIG. 4, the first prompt information can prompt the user to rotate the handset around the head to the left and / or to the right, or to rotate the handset first to the left and then to the right. For another example, in the scenario shown in (b) of FIG. 4, the first prompt information can prompt the user to hold the handset still and rotate the head to the left or to the right, or to rotate the head first to the left and then to the right.

[0133] - rotation angle or rotation range. For example, in the scenario shown in (a) of FIG. 4, the first prompt information can prompt the user to rotate the handset around the head to the left by 60 degrees and to the right by 60 degrees. For another example, in the scenario shown in (b) of FIG. 4, the first prompt information can prompt the user to hold the handset still and rotate the head to the left by 45 degrees and to the right by 45 degrees.

[0134] - duration. The first prompt information can prompt the user to perform the first operation for a duration.

[0135] In the embodiments of this application, the electronic device can adopt a single prompt mode or a combination of multiple prompt modes to instruct the user to perform the first operation. For example, the electronic device can adopt the following prompt modes to prompt the user to perform the first operation:

[0136] - voice prompt mode: the electronic device can play voice prompt information to instruct the user to perform the first operation.

[0137] - interface prompt mode: the electronic device can display first prompt information on an interface to instruct the user to perform the first operation. Optionally, the first prompt information can include one or a combination of text, image, and animation. The image and animation can show the relative motion between the electronic device and the head-mounted device (or the head), such as the relative motion shown in (a) or (b) of FIG. 4.

[0138] For example, FIG. 5 shows several examples of instructing the user to perform the first operation by the interface prompt mode provided in the embodiments of this application.

[0139] As shown in (a) of FIG. 5, the interface 510 includes a text prompt area 511, a first control 512, and optionally a second control 513. The text prompt area 511 displays information prompting the user to perform a first operation, for example, the following content can be prompted: "Please hold the mobile phone parallel to the line of sight, and keep a one-arm distance from the head. After clicking the first control, keep the mobile phone still, and slowly turn the head to the left and then to the right". The first control 512 is used to trigger the HRTF generation process, for example, "start test" can be displayed on the first control 512, after the first control 512 is triggered, the mobile phone enters the data collection stage, and the mobile phone can send instructions to the earphone through the first communication connection, so that the earphone also enters the data collection stage. The second control 513 is used to end the HRTF generation process, for example, "end test" can be displayed on the second control 513, after the second control 513 is triggered, the mobile phone ends the data collection stage and the calculation stage, and the mobile phone can send instructions to the earphone through the first communication connection, so that the earphone also ends the data collection stage and the calculation stage.

[0140] Optionally, the interface 510 can also include a third control 514 for canceling the operation of obtaining the HRTF, or in other words, for closing the interface 510. Considering that the same earphone can be used by different users, if the user has generated an HRTF before using the earphone this time, and the HRTF is the personalized HRTF of the user, the third control 514 can be triggered to cancel the operation of generating the personalized HRTF this time, and directly enter the use stage of the earphone.

[0141] As shown in (b) of FIG. 5, the interface 520 includes a text prompt area 521, an animation area 522, a first control 523, and optionally a second control 524. The text prompt area 521 displays information prompting the user to perform a first operation, and the animation area 522 displays an animation example of the first operation. For example, the animation area 522 displays an animation example of the first operation as shown in (b) of FIG. 4, and the text prompt area 521 can prompt the following content: "After clicking the first control, keep the mobile phone still, and slowly turn the head to the left and then to the right according to the following figure". The first control 523 is used to trigger the HRTF generation process, for example, "start test" can be displayed on the first control 523, and the second control 524 is used to end the HRTF generation process, for example, "end test" can be displayed on the second control 524. The functions of the first control 523 and the second control 524 can be referred to the above embodiments.

[0142] Optionally, the interface 510 or the interface 520 can also include a third control 525 for canceling the operation of obtaining the HRTF, or in other words, for closing the interface 520.

[0143] In another possible implementation, the first control and the second control can not be included in the interface 510 and the interface 520. The timing starts after the interface 510 or the interface 520 is opened, and the data collection stage is automatically entered after the first time length, and the data collection stage and the calculation stage are automatically ended after the second time length after entering the data collection stage.

[0144] It should be understood that FIG. 5 is only an example of several possible interfaces, and the present application is not limited thereto.

[0145] In a possible implementation, before outputting the first prompt information, the electronic device can first output the second prompt information in a voice manner and / or an interface manner, and the second prompt information is used to instruct the user to wear the head-mounted device. When the electronic device detects that the wearing state of the head-mounted device is “wearing”, the first prompt information is outputted.

[0146] Taking the application scenario of the above mobile phone and earphone as an example, the triggering manner of the data collection stage can include the following:

[0147] Triggering manner 1:

[0148] When the mobile phone and the earphone establish the first communication connection, the mobile phone and the earphone enter the data collection stage.

[0149] For example, when the mobile phone and the earphone establish the first communication connection, the mobile phone and the earphone can automatically enter the data collection stage.

[0150] For another example, when the mobile phone and the earphone establish the first communication connection, the mobile phone can output the first prompt information, and enter the data collection stage in response to the user operation (for example, the operation of triggering the first control).

[0151] Wherein, when the distance between the mobile phone and the earphone meets the distance specified by the first communication protocol, the first communication connection between the mobile phone and the earphone can be automatically established, or the mobile phone can establish the first communication connection with the earphone based on the user's setting operation.

[0152] Triggering manner 2:

[0153] The user opens the first application (which is used to obtain the HRTF) on the mobile phone, and the user interface of the first application is displayed on the screen of the mobile phone. The interface can display the above-mentioned first prompt information. For example, after the first application is started, the above-mentioned interface 510 or interface 520 can be displayed.

[0154] In another possible implementation, when the first application is started, the mobile phone can output the first prompt information in a voice playing manner

[0155] Triggering manner 3:

[0156] The user can issue a first voice instruction for triggering data collection, or in other words, the first voice instruction is for triggering the acquisition of the HRTF. The mobile phone can enter the data collection stage in response to the first voice instruction, and send an instruction to the earphone through the first communication connection, so that the earphone also enters the data collection stage.

[0157] Optionally, after the mobile phone responds to the first voice instruction, it can first prompt the user to perform the above-mentioned first operation through voice broadcast or by displaying a first prompt information on the screen.

[0158] Triggering method 4:

[0159] The mobile phone enters the data collection stage when it detects that the wearing state of the earphone is worn.

[0160] Illustratively, when the mobile phone detects that the wearing state of the earphone is worn, it can first output a first prompt information, and then enter the data collection stage in response to a user operation (such as triggering a first control operation), or enter the data collection stage after a set time period, and send an instruction to the earphone through the first communication connection, so that the earphone also enters the data collection stage.

[0161] It can be understood that the above exemplary lists several possible triggering methods of the data collection stage, and the present application does not limit this.

[0162] Taking the application scenario of the above-mentioned mobile phone and earphone as an example, the ways to end the data collection stage can include the following:

[0163] Ending method 1:

[0164] After the mobile phone enters the data collection stage, it automatically ends the data collection stage after a set time period, for example, the mobile phone ends playing the test audio signal. For example, the mobile phone can be provided with a timer, which is started when the mobile phone enters the data collection stage, and when the timer times out, the data collection stage ends. Optionally, the countdown time of the timer can be displayed on the screen of the mobile phone, or the progress of the data collection stage can be displayed through a progress bar.

[0165] Optionally, after the earphone enters the data collection stage, it automatically ends the data collection stage after the set time period.

[0166] Optionally, when the mobile phone automatically ends the data collection stage, it can send an instruction to the earphone through the first communication connection, for instructing the earphone to end the data collection stage.

[0167] Ending method 2:

[0168] The data collection stage is ended based on a first user operation, for example, the mobile phone ends playing the test audio signal.

[0169] For example, after the mobile phone enters the data collection stage, a user interface is displayed on the screen of the mobile phone, and the user interface displays a "test end" control. When the control is triggered, the mobile phone ends the data collection stage and sends an instruction to the earphone through the first communication connection, to instruct the earphone to end the data collection stage.

[0170] For another example, as shown in FIG. 5, after the user presses the first control of "test start", the first control is kept pressed during the data collection stage. When the user releases the first control, the mobile phone ends the data collection stage and sends an instruction to the earphone through the first communication connection, to instruct the earphone to end the data collection stage.

[0171] Ending mode 3:

[0172] The user can issue a second voice instruction, and the second voice instruction is used to end the data collection. The mobile phone can end the data collection stage in response to the second voice instruction, and send an instruction to the earphone through the first communication connection, so that the earphone also enters the data collection stage.

[0173] In a possible implementation, after the electronic device obtains the HRTF, or after the head-mounted device notifies the electronic device that the HRTF is obtained, the electronic device can output third prompt information, and the third prompt information is used to instruct the user to end the first operation, and / or is used to indicate that the HRTF has been generated. Optionally, the electronic device can output the third prompt information in the form of playing a voice and / or in the form of an interface prompt.

[0174] It should be understood that the above-mentioned triggering mode and stopping mode of data collection can be combined with each other. The embodiments of the present application do not limit the combination of the triggering mode and the stopping mode. For example, the triggering mode 1 and the ending mode 1 can be combined; for another example, the triggering mode 2 and the ending mode 2 can be combined; for yet another example, the triggering mode 3 and the ending mode 3 can be combined.

[0175] In the following, several user interface change processes involved in HRTF detection based on human-computer interaction are given in combination with FIG. 6 and FIG. 7.

[0176] Referring to FIG. 6, a schematic diagram of human-computer interaction provided by an embodiment of the present application is shown. As shown in the figure, a user selects the earphone 302 from the list of available devices in the communication connection setting interface 610 of the mobile phone 303. In response to the user operation, the first communication connection is established between the mobile phone 303 and the earphone 302. After the first communication connection is established, the mobile phone displays a prompt window 620, in which the second prompt information is displayed, to prompt the user to wear the earphone. The user can wear the earphone according to the second prompt information. When the mobile phone detects that the wearing state of the earphone is "wearing", the mobile phone displays the interface 520. The user can click the first control 523 according to the prompt information in the interface 520, and perform the first operation according to the prompt information after clicking the first control 523. In response to the event that the first control 523 is triggered, the mobile phone plays the test audio signal, receives the feedback audio signal and the position information from the earphone, and determines the HRTF corresponding to the user according to the test audio signal, the feedback audio signal and the position information.

[0177] Optionally, in the process of obtaining the HRTF, the mobile phone can display the motion of the head of the target individual and / or the motion of the hand holding the mobile phone in the user interface in real time based on the detected position information and the position information from the earphone, so as to facilitate the user to adjust the motion of the head or the motion of the hand in time.

[0178] Optionally, in the process of obtaining the HRTF, the mobile phone can output new prompt information to instruct the user to perform corresponding operation according to the new prompt information. For example, if the mobile phone determines that the rotation angle of the mobile phone relative to the head is less than the rotation angle indicated by the first prompt information according to the position information from the earphone, the mobile phone can output new prompt information to instruct the user to rotate a larger angle. Optionally, the new prompt information can be prompted by playing voice and / or interface display.

[0179] Optionally, after the mobile phone obtains the HRTF, a window can be displayed, in which the third prompt information can be displayed, which can be used to indicate that the HRTF has been generated, and the user can use the earphone to listen to audio. The display window can be automatically closed after a set time, or the display window is closed based on user operation, for example, the display window includes a control for "confirmation" or "close", and when the control is triggered, the display window is closed.

[0180] It should be understood that the "window" described above can be a floating window, which is replaced by "interface", and the present application does not limit this.

[0181] Referring to FIG. 7, another schematic diagram of human-computer interaction provided by an embodiment of the present application is shown. If the HRTF has been saved in the mobile phone, the mobile phone can directly use the HRTF to process the played audio signal. If a new HRTF is needed, the related function of obtaining the HRTF can be called out through voice, or the related function of obtaining the HRTF can be started through the setting interface, or the related function of obtaining the HRTF can be started through other manners. FIG. 7 takes the related function of obtaining the HRTF started through the setting interface as an example for description.

[0182] As shown in FIG. 7, the mobile phone and the earphone have established the first communication connection, and the mobile phone has detected that the wearing state of the earphone is “wearing”. The interface 520 is displayed when the user opens the setting interface 710 on the mobile phone and triggers the control of “obtaining HRTF” in the setting interface. The subsequent operation can refer to the related content in FIG. 6, which is not repeated here.

[0183] In the application scenario shown in FIG. 4, the relative spatial position between the earphone 302 (or the head 301 of the user) and the mobile phone 303 changes, for example, relative rotation or relative movement. The relative spatial position change can be described based on a spatial coordinate system with the head 301 as the center. FIG. 8 shows a schematic diagram of a spatial coordinate system with the head as the center. The black dot in the coordinate system represents the spatial position of the mobile phone relative to the head, which can be represented by the distance from the head, the pitch angle relative to the head, the horizontal angle relative to the head, and the like. Taking the spatial position A in FIG. 8 as an example, the distance between the spatial position and the head center point O is r, the pitch angle relative to the head is and the horizontal angle relative to the head is θ.

[0184] In a possible implementation, the mobile phone 303 and the earphone 302 are respectively provided with a spatial posture sensor (for example, an IMU).

[0185] In the scenario of calculating the HRTF by the mobile phone 303, the earphone 302 can send the detection data of the spatial posture sensor in the earphone to the mobile phone 303 through the first communication connection. The mobile phone 303 determines the spatial position of the mobile phone 303 relative to the head 301 based on the detection data of the spatial posture sensor in the mobile phone and the detection data of the spatial posture sensor in the earphone 302, and can calculate the HRTF based on the spatial position, or optimize the calculated HRTF based on the spatial position.

[0186] In the scenario where the earphone 302 calculates the HRTF, the mobile phone 303 can send the detection data of the spatial posture sensor in the mobile phone to the earphone 302 through the first communication connection, and the earphone 302 determines the spatial position of the mobile phone 303 relative to the head 301 according to the detection data of the spatial posture sensor in the earphone and the detection data of the spatial posture sensor in the mobile phone 303, and can calculate the HRTF according to the spatial position, or optimize the calculated HRTF according to the spatial position.

[0187] It can be understood that in some other application scenarios, the earphone can be replaced by other head-mounted devices, for example, it can be replaced by smart glasses, or a VR helmet or VR glasses, or an AR helmet or AR glasses, or an MR helmet or MR glasses, etc. Those skilled in the art can understand that in the case of replacing the mobile phone with other head-mounted devices, the prompt information of the electronic device (such as the mobile phone) can be adjusted accordingly. For example, in the case of replacing the earphone with smart glasses, the user interface displayed by the mobile phone can display prompt information prompting the user to wear the glasses correctly, and when it is detected that the wearing state of the glasses is "correct", the mobile phone displays prompt information prompting the user to rotate the head or hold the mobile phone around the head.

[0188] It can also be understood that in some other application scenarios, the mobile phone can be replaced by other electronic devices, for example, it can be replaced by a tablet computer, or a smart watch, or a smart bracelet, or other wearable devices.

[0189] Based on the above application scenarios, FIG. 9 shows a flowchart of a HRTF generation method according to an embodiment of the present application. The electronic device in the flowchart can be a mobile phone or the like in the above application scenarios, and the head-mounted device in the flowchart can be an earphone or the like in the above application scenarios. Before executing the following flow, the head-mounted device has been worn on the head of the target individual (i.e. the user), and the electronic device and the head-mounted device enter the data collection phase in the manner described in the above application scenarios. In the data collection phase, the electronic device and the head-mounted device move relative to each other in the manner described in the above application scenarios.

[0190] As shown in FIG. 9, the method can include the following steps:

[0191] Step 901: The electronic device plays a test audio signal.

[0192] In the present embodiment, the test audio signal can also be referred to as a probe audio signal or a reference audio signal, etc., which is not limited in the present application.

[0193] Optionally, the test audio signal can be an audio signal with a frequency response range. For example, the frequency response range can include the frequency range of sounds that can be heard by human beings, specifically, 20 Hz to 20 KHz; or a middle frequency band in the frequency range, for example, 50 MHz to 150 MHz.

[0194] Optionally, the test audio signal can be an audio signal with a frequency response range gradually increasing in order from low frequency to high frequency, or an audio signal with a frequency response range gradually decreasing in order from low frequency to high frequency, or an audio signal with a frequency continuously circulating from high to low, or a piece of music melody with a frequency range in the above-mentioned frequency response range. The application does not limit the test audio signal.

[0195] In this step, in the data collection phase, the electronic device plays the test audio signal through the loudspeaker thereof.

[0196] Optionally, in the data collection phase, the electronic device can continuously play the test audio signal.

[0197] Step 902: The head-mounted device worn on the head of the target individual collects a feedback audio signal of the test audio signal.

[0198] The head-mounted device can collect the feedback audio signal of the test audio signal generated at the ear of the target individual. In other words, the audio signal formed after the test audio signal played by the electronic device is reflected and / or diffracted by the ear and / or head of the target individual can be collected by the microphone of the head-mounted device.

[0199] In one possible implementation, taking earphones as an example, the microphone in the left earphone and the microphone in the right earphone respectively collect the feedback audio signal of the test audio signal. Since the earphones are worn on the head of the target individual, the microphone in the left earphone and the microphone in the right earphone are located at the entrance of the ear canal of the target individual, and thus the feedback audio signal collected by the microphone in the left earphone and the microphone in the right earphone can be a head related impulse response (HRIR), that is, the feedback audio signal is an audio signal generated after the test audio signal is reflected and / or diffracted by the ear, head and other human body parts of the target individual.

[0200] Step 903: The head-mounted device sends the collected feedback audio signal and the position information of the head-mounted device to the electronic device.

[0201] Optionally, the head-mounted device can collect the feedback audio signal according to a preset sampling frequency to obtain a time domain sequence of the feedback audio signal. Embodiments of the present application take the feedback audio signal corresponding to N sampling points in the time domain sequence of the feedback audio signal collected in the data collection stage as an example for description, where N is an integer greater than 1. Optionally, the N sampling points can be equal time intervals or can not be equal time intervals.

[0202] Optionally, the head-mounted device can send the feedback audio signal to the electronic device after the end of the data collection stage, or can send the feedback audio signal to the electronic device immediately after collecting the feedback audio signal, which is not limited in the present application.

[0203] In specific implementation, the head-mounted device can perform analog-to-digital conversion on the collected feedback audio signal of the analog signal type to obtain digital information of the feedback audio signal, and send the digital information of the feedback audio signal to the electronic device through the first communication connection between the head-mounted device and the electronic device.

[0204] Optionally, the position information of the head-mounted device is detected by a spatial attitude sensor in the head-mounted device. Optionally, the position information further includes a time stamp.

[0205] Optionally, the head-mounted device can send the position information to the electronic device after the end of the data collection stage, or can send the position information to the electronic device immediately after obtaining the position information, which is not limited in the present application.

[0206] Step 904: The electronic device determines the HRTF corresponding to the spatial position of the electronic device relative to the head-mounted device according to the spatial position of the electronic device relative to the head-mounted device, and the test audio signal and the feedback audio signal corresponding to the spatial position.

[0207] In this step, taking the earphone as an example, the electronic device can determine the HRTF of the left earphone at the spatial position of the electronic device relative to the left earphone according to the spatial position of the electronic device relative to the left earphone, and the test audio signal and the feedback audio signal corresponding to the spatial position. Similarly, the electronic device determines the HRTF of the right earphone at the spatial position of the electronic device relative to the right earphone according to the spatial position of the electronic device relative to the right earphone, and the test audio signal and the feedback audio signal corresponding to the spatial position.

[0208] For example, taking the spatial position of the electronic device relative to the head-mounted device as the first spatial position, the test audio signal corresponding to the first spatial position refers to the test audio signal played by the electronic device when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; and the feedback audio signal corresponding to the first spatial position refers to the feedback audio signal collected by the head-mounted device when the spatial position of the electronic device relative to the head-mounted device is the first spatial position.

[0209] In the data collection phase, the electronic device and the head-mounted device move relatively, and in the time-domain sequence of the feedback audio signals collected by the head-mounted device, the feedback audio signal corresponding to each sampling point is related to the spatial position of the electronic device relative to the head-mounted device at the corresponding time, or in other words, the feedback audio signal corresponding to each sampling point is associated with the spatial position of the electronic device relative to the head-mounted device at the time. The electronic device can determine, for each sampling point, the HRTF at the spatial position corresponding to the time according to the test audio signal and the feedback audio signal at the time.

[0210] In a possible implementation, the electronic device can determine, based on the clock of the electronic device, the time stamp of the test audio signal corresponding to each sampling point in the time-domain sequence of the test spectral signal played by the electronic device, and the feedback audio signal received by the electronic device from the head-mounted device can include the time stamp of the feedback audio signal corresponding to each sampling point in the time-domain sequence of the feedback audio signal. In addition, the position information detected by the spatial posture sensor of the electronic device includes a time stamp, and the position information received by the electronic device from the head-mounted device includes a time stamp, so that the electronic device can determine, according to the time stamp, the test audio signal, the feedback audio signal, and the spatial position corresponding to each sampling point, and thus can calculate the HRTF at the spatial position according to the test audio signal, the feedback audio signal, and the spatial position corresponding to each sampling point.

[0211] In a possible implementation, the spatial position of the electronic device relative to the head-mounted device can be calculated according to the position information detected by the spatial posture sensor of the electronic device and the position information received from the head-mounted device. Optionally, the spatial position of the electronic device relative to the head-mounted device can be calculated in a coordinate system with the head of the target individual as the center.

[0212] In a possible implementation, the electronic device can perform convolution operation on the test audio signal and the feedback audio signal, so as to obtain the HRTF.

[0213] For example, for a first spatial position on the relative motion path of the electronic device and the head-mounted device, the electronic device can perform Fourier transform on the time-domain signal of the test audio signal and the time-domain signal of the feedback audio signal of the left earphone at the first spatial position respectively, to obtain corresponding frequency-domain signals, and then multiply the two frequency-domain signals, and the result can be used as the HRTF of the left earphone at the spatial position. Similarly, for the first spatial position, the electronic device can perform Fourier transform on the time-domain signal of the test audio signal and the time-domain signal of the feedback audio signal of the right earphone at the first spatial position respectively, to obtain corresponding frequency-domain signals, and then multiply the two frequency-domain signals, and the result can be used as the HRTF of the right earphone at the spatial position.

[0214] In some embodiments, after obtaining the HRTF corresponding to a plurality of spatial positions according to the test audio signal and the feedback audio signal, the electronic device can calculate the HRTF corresponding to more spatial positions based on the interpolation algorithm and the calculated HRTF corresponding to the plurality of spatial positions, so as to optimize the HRTF, thereby improving the listening experience of the user when playing audio based on the optimized HRTF. For example, the spatial range of the relative motion between the electronic device and the head-mounted device is a spatial range centered on the head of the target individual, and the spatial range is 45 degrees left and right and 30 degrees up and down in front. After interpolation, the HRTF of the spatial range of 90 degrees left and right and 60 degrees up and down in front of the target individual can be obtained, or the HRTF of a larger spatial range can be obtained.

[0215] In some embodiments, the electronic device can collect the test audio signal through the microphone of the head-mounted device, and the head-mounted device can feed back the feedback audio signal collected by the microphone of the head-mounted device to the electronic device. For a certain spatial position in the relative motion between the electronic device and the head-mounted device, the electronic device can determine the HRTF corresponding to the spatial position according to the test audio signal and the feedback audio signal collected by the microphone of the head-mounted device.

[0216] For example, the electronic device can calculate a first HRTF (also referred to as a preliminary HRTF) according to the test audio signal and the feedback audio signal collected by the first microphone (e.g., a feedback microphone) in the left earphone, calculate a second HRTF according to the test audio signal and the feedback audio signal collected by the second microphone (e.g., a talk microphone) in the left earphone, and correct the first HRTF by using the second HRTF to obtain the HRTF corresponding to the left earphone. The accuracy of the HRTF obtained in this way is usually higher than that of the HRTF without correction.

[0217] Similarly, the HRTF corresponding to the right earphone can also be calculated in the above manner.

[0218] It should be understood that although the above describes the case where the first microphone and the second microphone are arranged in the right earphone or the left earphone, in actual applications, if more microphones are arranged in the earphone, the HRTF can also be calculated in the above manner.

[0219] The test audio signal is collected and fed back by the microphones at different positions in the head-mounted device to serve as the basis for calculating the HRTF. Compared with collecting and feeding back the test audio signal by using a single microphone, the accuracy of the HRTF can be improved, and thus the user's audio listening experience can be improved when the head-mounted device plays audio based on the HRTF.

[0220] Considering that the characteristics (or referred to as geometric characteristics) of the head-mounted device can affect the HRTF, in some embodiments, after the HRTF is calculated, the calculated HRTF can be corrected based on the characteristic parameters of the head-mounted device to improve the accuracy of the HRTF.

[0221] Optionally, the characteristic parameters of the head-mounted device can include one or more of the following:

[0222] The size of the head-mounted device;

[0223] The shape or structure of the head-mounted device;

[0224] The material of the head-mounted device;

[0225] The position of the microphone in the head-mounted device;

[0226] The acoustic response characteristic parameters of the head-mounted device, for example, the acoustic characteristic parameters can include frequency response, phase, etc. The acoustic characteristic parameters can be obtained from the relevant manual of the head-mounted device, or obtained by detecting the head-mounted device by professional equipment. In this way, the HRTF detected by the microphone can be compensated according to the frequency response and other characteristics of the microphone in the head-mounted device.

[0227] In order to improve the accuracy of the HRTF, in some embodiments, after the HRTF corresponding to the target individual is calculated, the HRTF calculated can be corrected based on the human body characteristics (such as ear structure characteristics and / or head structure characteristics) of the target individual, so as to improve the accuracy of the HRTF.

[0228] Optionally, the image of the target individual can be collected by using the front camera on the electronic device, and the ear structure characteristics (such as the shape of the auricle, the distance between the two ears, etc.) and / or the head structure characteristics of the target individual can be obtained by analyzing or recognizing the image.

[0229] Optionally, in the case where the user authorizes the use of the front camera in the data collection stage, the electronic device can automatically start the front camera and use the front camera to collect the image of the target individual. Optionally, in another scenario, the electronic device can obtain the user's use permission of the front camera through voice or interface, and start the front camera to collect the image of the target individual after obtaining the authorization.

[0230] Optionally, the inter-aural distance affects the HRTF, and in a possible implementation manner in which the HRTF is corrected based on the ear features and / or structural features of the target individual, the difference between the HRTFs of the two ears can be corrected according to the inter-aural distance, for example, the inter-aural time difference (ITD) and / or the inter-aural level difference (ILD).

[0231] To improve the accuracy of the HRTF, in some embodiments, after the HRTF is calculated according to the test audio signal and the feedback audio signal, the HRTF can be corrected based on the HRTF database.

[0232] The HRTF database can be obtained by performing acoustic-related tests on users with different human features. For example, the HRTF database can include HRTF models corresponding to a plurality of human features. The HRTF model includes HRTFs corresponding to a plurality of spatial positions, for example, HRTFs corresponding to spatial positions in the spatial range shown in FIG. 8. Each human feature can correspond to a set of human feature information, for example, including pinna shape, inter-aural distance, head circumference, head shape, and the like.

[0233] Optionally, the HRTF database can be obtained in various ways, for example, it can be downloaded from the network side or pre-set in the electronic device, which is not limited in the present application.

[0234] Taking the first HRTF corresponding to the first spatial position calculated according to the test audio signal and the feedback audio signal as an example, a possible implementation manner for correcting the HRTF based on the HRTF database is as follows: the electronic device can query the HRTF database according to the human feature information of the target individual, obtain a matched HRTF model, and obtain a second HRTF corresponding to the first spatial position in the HRTF model; if the error between the first HRTF and the second HRTF is within a set range, the first HRTF does not need to be corrected; if the error between the first HRTF and the second HRTF exceeds the set range, the first HRTF can be corrected according to the second HRTF, so that the error between the corrected first HRTF and the second HRTF is within the set range, for example, the first HRTF is corrected to the second HRTF.

[0235] Optionally, the human features of the target individual can be obtained in the following manner:

[0236] Manner 1: providing a user interface for a user to input his / her own anthropometric information based on the user interface, or to select anthropometric information matching the user from the anthropometric information options in the user interface.

[0237] Manner 2: using a camera of the electronic device to take a photo of the user, and obtaining the anthropometric information of the user based on image analysis.

[0238] For example, the electronic device can automatically collect the image of the target individual in the data collection stage.

[0239] Manner 3: prompting the user to input anthropometric information in a voice manner, and obtaining the anthropometric information input by the user in the voice manner.

[0240] With the above implementation manners, the HRTF obtained by the embodiments of the present application can present similar stereo effect as the HRTF obtained in the laboratory environment in the medium and high frequency bands.

[0241] With the above implementation manners, the accuracy of the HRTF can be improved. For example, the HRTF obtained by the embodiments of the present application can have similar stereo effect as the HRTF obtained in the laboratory environment in the medium and high frequency bands.

[0242] It should be understood that the above only exemplarily shows one possible implementation manner of correcting the calculated HRTF according to the HRTF queried from the HRTF database, and the implementation manner of how to correct the HRTF is not limited in the present application.

[0243] In one possible implementation manner, after the HRTF corresponding to the target individual is calculated, the HRTF of a larger spatial range can be obtained based on the HRTF database.

[0244] Exemplarily, taking a first spatial range (for example, a spatial range of 45 degrees left and right, and 30 degrees up and down in front of the target individual) centered on the head of the target individual as an example, based on the test audio signal and the feedback audio signal, HRTFs corresponding to a plurality of spatial positions in the first spatial range can be obtained. Based on the HRTF database, the electronic device can query the HRTF database according to the human feature information of the target individual, and obtain an HRTF model matched with the human feature information; and then obtain, from the HRTF model, HRTFs corresponding to at least one spatial position in a second spatial range centered on the head of the target individual. The second spatial range is non-overlapping with the first spatial range. Exemplarily, the electronic device can obtain, from the HRTF model, HRTFs corresponding to a plurality of spatial positions other than the first spatial range in the spatial range centered on the head of the target individual, so that the HRTF corresponding to the target individual includes HRTFs in all spatial ranges within a certain distance and centered on the head of the target individual.

[0245] In some embodiments, in the data collection stage, the target individual can be instructed to perform relative motion in the maximum spatial range as possible, such as being instructed to turn the mobile phone to the back of the head for testing, or being instructed to turn the head at the maximum angle for testing. In this way, the HRTF obtained by using the embodiments of the present application can have similar stereo sound effects as the HRTF obtained in the laboratory environment at a large angle, that is, the accuracy of the HRTF can be improved.

[0246] According to simulation experiments, the left ear HRTF and the right ear HRTF obtained based on the flow shown in FIG. 9 are relatively close to the HRTF obtained in the laboratory environment in the low and medium frequency band of 2KHz or less, and the HRTF obtained in the frequency band of 2KHz to 6KHz has a high correlation with the HRTF obtained in the laboratory environment. The HRTF obtained in the above process can be corrected by combining related modeling, simulation and fitting, etc. In the frequency band of 6KHz or more, the accuracy of the HRTF can be improved by the above-mentioned correction of the calculated HRTF.

[0247] In the above embodiments of the present application, the HRTF of the head-mounted device can be obtained by using the electronic device and the head-mounted device to perform simple testing, without the need for testing in a laboratory environment to obtain the HRTF, thereby simplifying the process of obtaining the HRTF and improving the convenience of obtaining the HRTF. In addition, the HRTF obtained by using the method provided in the embodiments of the present application does not require the user to provide user privacy information such as human feature parameters, and therefore the safety of the user privacy information can be protected.

[0248] Based on the above application scenarios, FIG. 10 shows a flowchart of a HRTF generation method according to an embodiment of the present application. The electronic device in the flowchart can be a mobile phone or the like in the above application scenarios, and the head-mounted device in the flowchart can be a headset or the like in the above application scenarios. Before the following flow is executed, the head-mounted device has been worn on the head of a target individual (i.e., a user), and the electronic device and the head-mounted device start entering a data collection phase in the manner described above. In the data collection phase, the electronic device and the head-mounted device move relative to each other in the manner described above.

[0249] As shown in FIG. 10, the method can include the following steps:

[0250] Step 1001: The electronic device plays a test audio signal.

[0251] The specific implementation of this step can refer to step 901 in FIG. 9.

[0252] Step 1002: The head-mounted device worn on the head of the target individual collects a feedback audio signal of the test audio signal.

[0253] The specific implementation of this step can refer to step 902 in FIG. 9.

[0254] Step 1003: The head-mounted device determines a HRTF corresponding to the spatial position of the electronic device relative to the head-mounted device according to the spatial position and the test audio signal and the feedback audio signal corresponding to the spatial position.

[0255] The implementation of the head-mounted device to determine the HRTF can refer to the implementation of the electronic device to determine the HRTF in FIG. 9.

[0256] Optionally, the test audio signal can be pre-set in the head-mounted device, or the electronic device can send digital information of the test audio signal to the head-mounted device through the first communication connection between the electronic device and the head-mounted device.

[0257] Based on the above application scenarios, FIG. 11 shows a flowchart of a HRTF generation method according to an embodiment of the present application. The electronic device in the flowchart can be a mobile phone or the like in the above application scenarios, and the head-mounted device in the flowchart can be a headset or the like in the above application scenarios. Before the following flow is executed, the head-mounted device has been worn on the head of a target individual (i.e., a user), and the electronic device and the head-mounted device start entering a data collection phase in the manner described above. In the data collection phase, the electronic device and the head-mounted device move relative to each other in the manner described above.

[0258] As shown in FIG. 11, the method can include the following steps:

[0259] Step 1101: The electronic device plays a test audio signal.

[0260] The specific implementation of this step can refer to step 901 in FIG. 9.

[0261] Step 1102: The head-mounted device worn on the head of the target individual collects a feedback audio signal of the test audio signal.

[0262] The specific implementation of this step can refer to step 902 in FIG. 9.

[0263] Step 1103: The head-mounted device sends the collected feedback audio signal and the position information of the head-mounted device to the electronic device.

[0264] The specific implementation of this step can refer to step 903 in FIG. 9.

[0265] Step 1104: The electronic device sends the test audio signal, the feedback audio signal, and the information of the spatial position of the electronic device relative to the head-mounted device to the server.

[0266] Step 1105: The server determines the HRTF corresponding to the spatial position of the head-mounted device relative to the electronic device, and the test audio signal and the feedback audio signal corresponding to the spatial position.

[0267] The implementation of the server determining the HRTF can refer to the implementation of the electronic device determining the HRTF in FIG. 9.

[0268] Step 1106: The server sends the HRTF to the electronic device.

[0269] The embodiments of the present application also provide a method for generating an HRTF and related devices and systems that can implement the method, so as to conveniently generate personalized HRTFs based on human-computer interaction. In the method, the electronic device instructs the head of the target individual to move relative to the electronic device based on human-computer interaction, the electronic device collects multi-directional images of the head and / or ears of the target individual, models the head based on the images, and obtains the personalized HRTF corresponding to the target individual based on the head modeling. The structure and type of the electronic device can refer to the foregoing embodiments.

[0270] Referring to FIG. 12, it is a schematic diagram of an application scenario in the embodiments of the present application. A user holds a mobile phone 1202, and the distance between the mobile phone 1202 and the head 1201 of the user is approximately the length of an arm.

[0271] The HRTF generation process can include a data collection stage and a calculation stage. In the data collection stage, the mobile phone 1202 (e.g., a front-facing camera in the mobile phone) can collect multi-directional images of the head of the target individual. In the calculation stage, a three-dimensional model of the head of the target individual can be generated based on the collected multi-directional images of the head of the target individual, and the personalized HRTF of the target individual can be determined based on the three-dimensional model.

[0272] As shown in (a) of FIG. 12, in the data collection stage, the head 1201 of the user remains stationary, and the user holds the mobile phone 1202 to slowly move around the head 1201. As shown in (b) of FIG. 12, in the data collection stage, the head 1201 of the user rotates, and the user holds the mobile phone 1202 to remain stationary. The movement direction can refer to the related content in FIG. 4.

[0273] In the calculation stage, the mobile phone 1202 generates a head model of the target individual based on the collected images, and then generates the HRTF based on the model.

[0274] In a possible implementation, in the data collection stage, the spatial pose sensor in the mobile phone 1202 can detect the position information of the mobile phone (which can reflect the relative position relationship between the mobile phone and the head) during the rotation of the mobile phone 1202 around the head, and model the head based on the collected images and the detected position information. Since the detection data of the spatial pose sensor can reflect the relative position relationship between the mobile phone and the head, that is, can reflect the depth of field of the head object in the collected images, combining the position information and the images together for analysis and modeling can improve the modeling accuracy, and thus can improve the accuracy of the HRTF.

[0275] In another possible implementation, the head 1201 of the target individual can wear earphones, and other processing can refer to the scenario shown in (a) of FIG. 12. In this scenario, the image collected by the mobile phone 1202 can be partially blocked by the earphones to the ears, and the structure of the entire pinna can be reconstructed based on the algorithm through the unblocked pinna part.

[0276] In another possible implementation, the head 1201 of the target individual can wear earphones, and the spatial pose sensor in the earphones can send the detected data to the mobile phone 1202. The mobile phone 1202 can determine the relative position between the mobile phone and the earphones (or the head of the target individual) based on the position information detected by the spatial pose sensor in the mobile phone and the position information received from the earphones, and combine the relative position and the images together for analysis and modeling, which can improve the modeling accuracy, and thus can improve the accuracy of the HRTF.

[0277] In the above process, the mobile phone 1202 can output prompt information to prompt the user to operate, or to guide the user to complete the generation of the HRTF. The prompting manner of the mobile phone 1202 can refer to the foregoing embodiments.

[0278] It should be understood that the electronic device can also be replaced by other devices, and the types of the other devices can refer to the foregoing embodiments. When the mobile phone is replaced by other electronic devices, those skilled in the art can make necessary adjustments or improvements, for example, when the electronic device is a large-screen device (for example, a smart television), the electronic device can guide the user to rotate the head to complete the HRTF generation process based on the human-computer interaction manner.

[0279] It should be understood that the head-mounted device can be replaced by other devices, and the specific devices can refer to the foregoing embodiments. When the earphone is replaced by other electronic devices, those skilled in the art can make necessary adjustments or improvements.

[0280] In a possible implementation, after the HRTF is obtained based on the head three-dimensional model in the foregoing manner, a HRTF of a larger spatial range can be obtained based on the HRTF database, or a HRTF corresponding to a part of positions (for example, a HRTF with a large difference from the HRTF in the HRTF database) can be corrected based on the HRTF database or the characteristics (or referred to as geometric characteristics) of the head-mounted device. The HRTF of the larger spatial range is obtained based on the HRTF database. The specific implementation can refer to the foregoing embodiments.

[0281] In the foregoing embodiments of the present application, the HRTF of the head-mounted device can be obtained by using the electronic device to perform a simple test, without the need to test in a laboratory environment to obtain the HRTF, thereby simplifying the process of obtaining the HRTF and improving the convenience of obtaining the HRTF.

[0282] Based on the foregoing embodiments and the same concept, the embodiments of the present application further provide an electronic device, which is used to implement the method performed by the electronic device provided in the embodiments of the present application.

[0283] As shown in FIG. 13, the electronic device 1300 can include a memory 1301, one or more processors 1302, and one or more computer programs (not shown in the figure). The above devices can be coupled through one or more communication buses 1303. Optionally, when the electronic device 1300 is used to implement the method performed by the electronic device provided in the embodiments of the present application, the electronic device 1300 can further include a display screen 1304.

[0284] The memory 1301 stores one or more computer programs (codes) including computer instructions, and the one or more processors 1302 invoke the computer instructions stored in the memory 1301 to enable the electronic device 1300 to perform the method provided in the embodiments of the present application. The display screen 1304 is used to display images, videos, application interfaces, and other related user interfaces.

[0285] In specific implementations, the memory 1301 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 1301 can store an operating system (hereinafter referred to as a system), such as an ANDROID, IOS, WINDOWS, or LINUX embedded operating system. The memory 1301 can be used to store the implementation program of the embodiments of the present application. The memory 1301 can also store a network communication program, which can be used to communicate with one or more additional devices, one or more user devices, and one or more network devices. The one or more processors 1302 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.

[0286] It should be noted that FIG. 13 is only one implementation of the electronic device 1300 provided by the embodiments of the present application, and in actual applications, the electronic device 1300 can also include more or fewer components, which are not limited here.

[0287] Based on the above embodiments and the same concept, the embodiments of the present application also provide a computer readable storage medium storing a computer program, when the computer program runs on a computer, the computer program enables the computer to perform the method performed by the electronic device in the method provided in the above embodiments.

[0288] Based on the above embodiments and the same concept, the embodiments of the present application also provide a computer program product including a computer program or instructions, when the computer program or instructions run on a computer, the computer program or instructions enable the computer to perform the method performed by the electronic device in the method provided in the above embodiments.

[0289] Those skilled in the art will appreciate that embodiments of the present application can be readily used as software, hardware, or a combination of software and hardware. In a software embodiment, various software modules in accordance with embodiments of the present application are stored in a memory such as a computer memory or disk storage for use by, or in connection with, the software on the computer system. The software can provide for programs to be transferred to another computer readable medium (e.g., a removable medium, or a medium conveyed through a computer network) for use in a different system.

[0290] The present application is described in reference to the flow diagrams and / or block diagrams of the methods, apparatus (systems) and computer program products according to this application. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks.

[0291] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flow diagrams and / or block diagrams block or blocks.

[0292] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow diagrams and / or block diagrams block or blocks.

[0293] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A head-related transfer function (HRTF) generation method, characterized by, The method is applied to an electronic device, and the method comprises: outputting first prompt information, the first prompt information being used to instruct a target individual wearing a head-mounted device to perform a first operation, the first operation causing relative motion between the electronic device and the head-mounted device; playing a test audio signal during performance of the first operation, and receiving a feedback audio signal and position information from the head-mounted device; wherein the feedback audio signal is the test audio signal received by the head-mounted device at a plurality of different spatial positions, the spatial positions being spatial positions of the electronic device relative to the head-mounted device, and the position information being used to determine the spatial positions; determining an HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information.

2. The method of claim 1, wherein, the first operation is that the target individual holds the electronic device and rotates the electronic device around the head while keeping the head still; or the first operation is that the target individual rotates the head while keeping the electronic device held still.

3. The method of claim 2, wherein, the first prompt information is further used to instruct one or more of the following: a distance between the electronic device and the head-mounted device, a rotation direction, a rotation angle, a rotation range, and a duration.

4. The method according to any one of claims 1 to 3, characterized in that, the outputting of the first prompt information comprises: outputting the first prompt information after establishing a communication connection between the electronic device and the head-mounted device, the communication connection being used to transmit the feedback audio signal and the position information; or outputting the first prompt information after starting a first application, the first application being used to obtain an HRTF; or outputting the first prompt information in response to a first voice instruction, the first voice instruction being used to trigger the obtaining of the HRTF; or outputting the first prompt information when it is detected that a wearing state of the head-mounted device is wearing.

5. The method according to any one of claims 1 to 4, characterized in that, before the outputting of the first prompt information, the method further comprises: outputting second prompt information, the second prompt information being used to instruct the target individual to wear the head-mounted device. the outputting of the first prompt information comprises: outputting the first prompt information when it is detected that a wearing state of the head-mounted device is wearing.

6. The method according to any one of claims 1 to 5, wherein, the method further comprises: ending the playing of the test audio signal after a set time period; or ending the playing of the test audio signal based on a first user operation; or ending the playing of the test audio signal in response to a second voice instruction. the method further comprises: outputting third prompt information, the third prompt information being used to instruct the target individual to end the first operation, and / or being used to indicate that an HRTF has been generated.

7. The method according to any one of claims 1 to 6, wherein the outputting of the first prompt information comprises one or more of the following: playing voice prompt information; or 8. The method according to any one of claims 1 to 7, wherein, displaying text prompt information; or displaying an image or an animation, the image or the animation being used to show a manner in which relative motion between the electronic device and the head-mounted device is caused. the method further comprises: determining spatial positions of the electronic device relative to the head-mounted device according to position information of the electronic device and position information from the head-mounted device; 9. The method according to any one of claims 1 to 8, wherein, ​ ​ The position information of the head-mounted device is detected by a spatial posture sensor in the head-mounted device, and the position information of the electronic device is detected by a spatial posture sensor in the electronic device.

10. The method of any one of claims 1-9, wherein, The feedback audio signal includes a first feedback audio signal collected by a first microphone of the head-mounted device and a second feedback audio signal collected by a second microphone of the head-mounted device. The HRTF corresponding to the target individual is determined according to the test audio signal, the feedback audio signal, and the position information. The HRTF corresponding to the target individual is determined according to the test audio signal, the first feedback audio signal, the second feedback audio signal, and the position information.

11. The method of claim 10, wherein, The spatial position includes a first spatial position. The HRTF corresponding to the target individual is determined according to the test audio signal, the first feedback audio signal, the second feedback audio signal, and the spatial position. A first HRTF corresponding to the first spatial position is determined according to the test audio signal played by the electronic device and the first feedback audio signal collected by the first microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position. A second HRTF corresponding to the first spatial position is determined according to the test audio signal played by the electronic device and the second feedback audio signal collected by the second microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position. The first HRTF is corrected based on the second HRTF to obtain the HRTF corresponding to the first spatial position.

12. The method of any one of claims 1-11, wherein, After determining the HRTF corresponding to the target individual, one or more of the following is performed: The HRTF is corrected according to a characteristic parameter of the head-mounted device; or The HRTF is corrected according to a human body feature of the target individual; or The HRTF is corrected according to an HRTF database.

13. The method of claim 12, wherein, The HRTF is corrected according to a human body feature of the target individual, including: An image of the target individual object collected by the electronic device is obtained. The image is identified to obtain auricle structure features and / or head structure features of the target individual. The HRTF is corrected according to the auricle structure features and / or head structure features of the target individual.

14. The method of claim 12, wherein, The spatial position of the electronic device relative to the head-mounted device includes a first spatial position, and the HRTF determined according to the test audio signal, the feedback audio signal, and the position information includes a first HRTF corresponding to the first spatial position. The HRTF is corrected according to an HRTF database, including: A second HRTF matching the human body feature information and the first spatial position is obtained from the HRTF database according to the human body feature information of the target individual and the first spatial position; The first HRTF is corrected according to the second HRTF.

15. The method of any one of claims 1-14, wherein, The HRTF determined according to the test audio signal, the feedback audio signal and the position information comprises HRTFs corresponding to at least two spatial positions in a first spatial range; The method further comprises: According to the human body characteristics of the target individual, obtaining, from an HRTF database, HRTFs corresponding to at least one spatial position in a second spatial range which matches the human body characteristics; Determining the HRTFs corresponding to the at least two spatial positions in the first spatial range and the HRTFs corresponding to the at least one spatial position in the second spatial range as the HRTF corresponding to the target individual.

16. A head-related transfer function (HRTF) generation method, comprising: Applied to a head-mounted device, the method comprises: Collecting a feedback audio signal, the feedback audio signal being a test audio signal played by an electronic device and received by the head-mounted device at a plurality of different spatial positions, the spatial positions being spatial positions of the electronic device relative to the head-mounted device; Receiving position information from the electronic device, the position information being used to determine the spatial positions; Determining, according to the test audio signal, the feedback audio signal and the position information, an HRTF corresponding to the target individual.

17. The method of claim 16, wherein, Further comprising: Determining, according to the position information of the head-mounted device and the position information from the electronic device, the spatial positions of the electronic device relative to the head-mounted device; The position information of the head-mounted device is detected by a spatial attitude sensor in the head-mounted device, and the position information of the electronic device is detected by a spatial attitude sensor in the electronic device.

18. The method of any one of claims 16-17, wherein, The feedback audio signal comprises a first feedback audio signal collected by a first microphone of the head-mounted device and a second feedback audio signal collected by a second microphone of the head-mounted device; The determining, according to the test audio signal, the feedback audio signal and the position information, of the HRTF corresponding to the target individual comprises: Determining, according to the test audio signal, the first feedback audio signal and the second feedback audio signal and the position information, the HRTF corresponding to the target individual.

19. The method of claim 18, wherein, The spatial positions comprise a first spatial position; The determining, according to the test audio signal, the first feedback audio signal and the second feedback audio signal and the spatial positions, of the HRTF corresponding to the target individual comprises: According to the test audio signal, the first feedback audio signal and the second feedback audio signal and the spatial positions, determining a first HRTF corresponding to the first spatial position when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; According to the test audio signal, the first feedback audio signal and the second feedback audio signal and the spatial positions, determining a second HRTF corresponding to the first spatial position when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; Based on the second HRTF, correcting the first HRTF to obtain an HRTF corresponding to the first spatial position.

20. The method of any one of claims 16-19, wherein, After the determining of the HRTF corresponding to the target individual, the method further comprises one or more of the following: The HRTF is corrected according to a characteristic parameter of the head-mounted device; or The HRTF is corrected according to a human body feature of the target individual; or The HRTF is corrected according to an HRTF database.

21. The method of claim 20, wherein, The spatial position of the electronic device relative to the head-mounted device includes a first spatial position, and the HRTF determined according to the test audio signal, the feedback audio signal, and the position information includes a first HRTF corresponding to the first spatial position; The HRTF is corrected according to an HRTF database, including: According to the human body feature information of the target individual and the first spatial position, a second HRTF matching the human body feature information and the first spatial position is obtained from the HRTF database; The first HRTF is corrected according to the second HRTF.

22. The method of any one of claims 16-21, wherein, The HRTF determined according to the test audio signal, the feedback audio signal, and the position information includes HRTFs corresponding to at least two spatial positions in a first spatial range; The method further includes: According to the human body feature of the target individual, an HRTF corresponding to at least one spatial position in a second spatial range matching the human body feature is obtained from an HRTF database; The HRTFs corresponding to the at least two spatial positions in the first spatial range and the HRTF corresponding to the at least one spatial position in the second spatial range are determined as the HRTF corresponding to the target individual.

23. An apparatus, comprising: The apparatus includes a unit or module for performing the method of any one of claims 1-15, or a unit or module for performing the method of any one of claims 16-22.

24. An apparatus comprising: The apparatus includes: One or more processors configured to perform the method of any one of claims 1-15, or the method of any one of claims 16-22.

25. A readable storage medium characterized by, The readable storage medium stores a program or instruction, which, when executed on the apparatus, causes the apparatus to perform the method of any one of claims 1-15, or the method of any one of claims 16-22.

26. A chip system, characterized by The apparatus includes a processor for supporting a computer device to implement the method of any one of claims 1-15, or the method of any one of claims 16-22.

27. A program product, characterized by The program product includes a program; when the program is executed on a computer, the computer performs the method of any one of claims 1-15, or the method of any one of claims 16-22.

Citation Information

Patent Citations

  • Compensating for effects of headset on head related transfer functions

    CN113366863A

  • Head correlation function HRTF determination method, electronic equipment and storage medium

    CN114710739A

  • Head-related transfer function recording using positional tracking

    US9648438B1