Head related transfer function generation method and device
By using human-computer interaction to capture and process audio signals through relative movement between electronic devices and head-mounted devices, personalized HRTFs are generated, solving the problem of convenient HRTF generation and improving the accuracy of 3D sound effects and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies struggle to easily generate personalized head-related transfer functions (HRTFs), especially for achieving realistic 3D sound effects on small mobile devices such as wireless headphones or smart glasses.
By using human-computer interaction and the relative motion between electronic devices and head-mounted devices, test audio signals and feedback audio signals are collected and processed, and combined with location information, personalized HRTFs are generated.
It improves the accuracy of HRTF, enhances the user's audio listening experience, and is suitable for a variety of portable electronic devices and head-mounted devices.
Smart Images

Figure CN121751074A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio processing, and in particular to a head related transfer function generation method and device. BACKGROUND
[0002] In order to meet the needs of immersive sound experience, traditional stereo sound effects can present three-dimensional sound effects based on multi-channel technology. Although the traditional stereo sound effects have good three-dimensional sound presentation effects, they cannot be applied to small mobile devices such as wireless earphones or smart glasses due to the need to configure multiple loudspeakers.
[0003] A person can distinguish the direction of a sound source in space through both ears, and the principle lies in that the human brain can distinguish the direction of the sound source through the subtle differences in sound between both ears. The head related transfer function (HRTF) can describe the changes that occur when sound is reflected and diffracted by human body parts such as pinna, head, and torso during the process of sound being transmitted into the human ear from a specific direction. Therefore, the HRTF can be used on devices such as earphones and smart glasses to simulate the sound information heard by both ears of a person, so as to present a realistic three-dimensional sound effect.
[0004] The HRTF is closely related to an individual, and the HRTFs of different individuals have great differences. How to conveniently generate personalized HRTF is a problem to be solved at present. SUMMARY
[0005] Embodiments of the present application provide a head related transfer function generation method and device to conveniently generate personalized head related transfer functions based on human-computer interaction.
[0006] In a first aspect, a HRTF generation method is provided, which can be applied to an electronic device, such as a convenient electronic device such as a mobile phone, a tablet computer, a notebook computer, etc.; it can also be a wearable device such as a watch, a bracelet, etc.; or it can also be a vehicle-mounted device, etc. The method comprises the following steps: outputting first prompt information, the first prompt information being used to instruct a target individual wearing a head-mounted device to perform a first operation, the first operation causing a relative motion between the electronic device and the head-mounted device; playing a test audio signal during the first operation, and receiving a feedback audio signal and position information from the head-mounted device; wherein the feedback audio signal is the test audio signal received by the head-mounted device at a plurality of different spatial positions, the spatial position being a spatial position of the electronic device relative to the head-mounted device, and the position information being used to determine the spatial position; determining the HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information.
[0007] The implementation manner can generate the HRTF based on human-computer interaction by using the electronic device and the head-mounted device when the target individual wears the head-mounted device, so that the personalized HRTF can be conveniently generated.
[0008] In a possible implementation manner, the first operation is that the target individual holds the electronic device and rotates the head-mounted device around the head while keeping the head still, or the first operation is that the target individual rotates the head while keeping the electronic device held still.
[0009] In a possible implementation manner, the first prompt information further indicates one or more of the following: a distance between the electronic device and the head-mounted device, a rotation direction, a rotation angle, a rotation range, and a duration.
[0010] In a possible implementation manner, the outputting of the first prompt information comprises: outputting the first prompt information after establishing a communication connection between the electronic device and the head-mounted device, the communication connection being used to transmit the feedback audio signal and the position information; or outputting the first prompt information after starting a first application program, the first application program being used to obtain the HRTF; or outputting the first prompt information in response to a first voice instruction, the first voice instruction being used to trigger the obtaining of the HRTF; or outputting the first prompt information when it is detected that the wearing state of the head-mounted device is wearing.
[0011] In a possible implementation manner, before the outputting of the first prompt information, the method further comprises: outputting second prompt information, the second prompt information being used to instruct the target individual to wear the head-mounted device; and the outputting of the first prompt information comprises: outputting the first prompt information when it is detected that the wearing state of the head-mounted device is wearing.
[0012] In a possible implementation manner, the method further comprises: ending the playing of the test audio signal after a set time period; or ending the playing of the audio signal based on a first user operation; or ending the playing of the test audio signal in response to a second voice instruction.
[0013] In a possible implementation manner, the method further comprises: outputting third prompt information, the third prompt information being used to instruct the target individual to end the first operation and / or being used to indicate that the HRTF has been generated.
[0014] In a possible implementation manner, the outputting of the first prompt information comprises one or more of the following: playing voice prompt information; or displaying text prompt information; or displaying an image or an animation, the image or the animation being used to show a manner in which relative motion is generated between the electronic device and the head-mounted device.
[0015] In a possible implementation, the method further includes: determining a spatial position of the electronic device relative to the head-mounted device according to the position information of the electronic device and the position information of the head-mounted device, wherein the position information of the head-mounted device is detected by a spatial posture sensor in the head-mounted device, and the position information of the electronic device is detected by a spatial posture sensor in the electronic device.
[0016] In a possible implementation, the feedback audio signal includes a first feedback audio signal collected by a first microphone of the head-mounted device and a second feedback audio signal collected by a second microphone of the head-mounted device, and the determining the HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information includes: determining the HRTF corresponding to the target individual according to the test audio signal, the first feedback audio signal, the second feedback audio signal, and the position information.
[0017] In the implementation, the feedback audio signal of the test audio signal can be collected by using multiple microphones in the head-mounted device, and the HRTF is generated accordingly, so that the accuracy of the HRTF can be improved, and the audio listening experience of the user can be improved.
[0018] In a possible implementation, the spatial position includes a first spatial position, and the determining the HRTF corresponding to the target individual according to the test audio signal, the first feedback audio signal, the second feedback audio signal, and the spatial position includes: determining a first HRTF corresponding to the first spatial position according to the test audio signal played by the electronic device and the first feedback audio signal collected by the first microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; determining a second HRTF corresponding to the first spatial position according to the test audio signal played by the electronic device and the second feedback audio signal collected by the second microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; and correcting the first HRTF based on the second HRTF to obtain the HRTF corresponding to the first spatial position.
[0019] In a possible implementation, after the HRTF corresponding to the target individual is determined, one or more of the following is further included: correcting the HRTF according to a characteristic parameter of the head-mounted device; or correcting the HRTF according to a human body feature of the target individual; or correcting the HRTF according to an HRTF database.
[0020] In the above implementation manner, the calculated HRTF can be further corrected according to the characteristic parameter of the head-mounted device or the human body feature of the target individual or the HRTF database, so that the accuracy of the HRTF can be improved.
[0021] In a possible implementation manner, the correcting the HRTF according to the human body feature of the target individual comprises: acquiring an image of the target individual collected by the electronic device; identifying the image to obtain an auricle structure feature and / or a head structure feature of the target individual; and correcting the HRTF according to the auricle structure feature and / or the head structure feature of the target individual.
[0022] In a possible implementation manner, the spatial position of the electronic device relative to the head-mounted device comprises a first spatial position, the HRTF determined according to the test audio signal, the feedback audio signal and the position information comprises a first HRTF corresponding to the first spatial position, and the correcting the HRTF according to the HRTF database comprises: acquiring, from the HRTF database, a second HRTF matched with the human body feature information and the first spatial position according to the human body feature information of the target individual and the first spatial position; and correcting the first HRTF according to the second HRTF.
[0023] In a possible implementation manner, the HRTF determined according to the test audio signal, the feedback audio signal and the position information comprises HRTFs corresponding to at least two spatial positions in a first spatial range, and the method further comprises: acquiring, from the HRTF database, HRTFs corresponding to at least one spatial position in a second spatial range matched with a human body feature of the target individual according to the human body feature of the target individual; and determining the HRTFs corresponding to the at least two spatial positions in the first spatial range and the HRTFs corresponding to the at least one spatial position in the second spatial range as HRTFs corresponding to the target individual.
[0024] In a second aspect, a method for generating HRTF is provided. The method can be applied to a head-mounted device, such as a headphone, or smart glasses, or a virtual reality (VR) headset or VR glasses, or an augmented reality (AR) headset or AR glasses, or a mixed reality (MR) headset or MR glasses, etc. The method can include the following steps: collecting a feedback audio signal, the feedback audio signal being a test audio signal played by an electronic device and received by the head-mounted device at a plurality of different spatial positions, the spatial positions being spatial positions of the electronic device relative to the head-mounted device; receiving position information from the electronic device, the position information being used to determine the spatial positions; and determining the HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information.
[0025] In a possible implementation, the method further includes: determining the spatial positions of the electronic device relative to the head-mounted device according to position information of the head-mounted device and position information from the electronic device; wherein the position information of the head-mounted device is detected by a spatial attitude sensor in the head-mounted device, and the position information of the electronic device is detected by a spatial attitude sensor in the electronic device.
[0026] In a possible implementation, the feedback audio signal includes a first feedback audio signal collected by a first microphone of the head-mounted device and a second feedback audio signal collected by a second microphone of the head-mounted device; and the determining the HRTF corresponding to the target individual according to the test audio signal, the feedback audio signal, and the position information includes: determining the HRTF corresponding to the target individual according to the test audio signal, the first feedback audio signal, the second feedback audio signal, and the position information.
[0027] In a possible implementation, the spatial position includes a first spatial position; and the determining the HRTF corresponding to the target individual according to the test audio signal, the first feedback audio signal and the second feedback audio signal, and the spatial position includes: determining a first HRTF corresponding to the first spatial position according to the test audio signal played by the electronic device and the first feedback audio signal collected by the first microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; determining a second HRTF corresponding to the first spatial position according to the test audio signal played by the electronic device and the second feedback audio signal collected by the second microphone when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; and correcting the first HRTF based on the second HRTF to obtain the HRTF corresponding to the first spatial position.
[0028] In a possible implementation, after the HRTF corresponding to the target individual is determined, the method further includes one or more of the following: correcting the HRTF according to a characteristic parameter of the head-mounted device; or correcting the HRTF according to a human body feature of the target individual; or correcting the HRTF according to an HRTF database.
[0029] In a possible implementation, the spatial position of the electronic device relative to the head-mounted device includes a first spatial position, and the HRTF determined according to the test audio signal, the feedback audio signal and the position information includes a first HRTF corresponding to the first spatial position; and the correcting the HRTF according to the HRTF database includes: obtaining, from the HRTF database, a second HRTF matching the human body feature information of the target individual and the first spatial position according to the human body feature information of the target individual and the first spatial position; and correcting the first HRTF according to the second HRTF.
[0030] In a possible implementation, the HRTF determined according to the test audio signal, the feedback audio signal and the position information includes HRTFs corresponding to at least two spatial positions in a first spatial range; and the method further includes: obtaining, from an HRTF database, HRTFs corresponding to at least one spatial position in a second spatial range matching a human body feature of the target individual according to the human body feature of the target individual; and determining the HRTFs corresponding to the at least two spatial positions in the first spatial range and the HRTFs corresponding to the at least one spatial position in the second spatial range as the HRTF corresponding to the target individual.
[0031] In a third aspect, a method for generating HRTF is provided. The method can be applied to an electronic device. The method comprises: outputting first prompt information, the first prompt information being used to instruct a target individual wearing a head-mounted device to perform a first operation, the first operation causing relative movement between the electronic device and the head-mounted device; playing a test audio signal during the first operation and receiving a feedback audio signal and position information from the head-mounted device, wherein the feedback audio signal is the test audio signal received by the head-mounted device at a plurality of different spatial positions, the spatial positions being spatial positions of the electronic device relative to the head-mounted device, and the position information being used to determine the spatial positions; sending the test audio signal, the feedback audio signal, and information of the spatial positions to a server, and receiving an HRTF from the server, the HRTF being determined based on the test audio signal, the feedback audio signal, and the information of the spatial positions.
[0032] In a possible implementation, the first operation is that the target individual holds the electronic device and rotates the electronic device around the head while keeping the head still, or the first operation is that the target individual rotates the head while keeping the electronic device held in the hand still.
[0033] In a possible implementation, the first prompt information is further used to instruct one or more of the following: a distance between the electronic device and the head-mounted device, a rotation direction, a rotation angle, a rotation range, and a duration.
[0034] In a possible implementation, the outputting of the first prompt information comprises: outputting the first prompt information after establishing a communication connection between the electronic device and the head-mounted device, the communication connection being used to transmit the feedback audio signal and the position information; or outputting the first prompt information after starting a first application, the first application being used to obtain the HRTF; or outputting the first prompt information in response to a first voice instruction, the first voice instruction being used to trigger the obtaining of the HRTF; or outputting the first prompt information when it is detected that the head-mounted device is worn.
[0035] In a possible implementation, before the outputting of the first prompt information, the method further comprises: outputting second prompt information, the second prompt information being used to instruct the target individual to wear the head-mounted device; and the outputting of the first prompt information comprises: outputting the first prompt information when it is detected that the head-mounted device is worn.
[0036] In a possible implementation, the method further includes: ending playing the test audio signal after a set time duration; or ending playing the audio signal based on a first user operation; or ending playing the test audio signal in response to a second voice instruction.
[0037] In a possible implementation, the method further includes: outputting third prompt information, the third prompt information being used to instruct the target individual to end the first operation and / or being used to indicate that the HRTF has been generated.
[0038] In a possible implementation, the outputting of the first prompt information includes one or more of the following: playing voice prompt information; or displaying text prompt information; or displaying an image or an animation, the image or the animation being used to show a manner in which relative motion is generated between the electronic device and the head-mounted device.
[0039] In a possible implementation, the method further includes: determining a spatial position of the electronic device relative to the head-mounted device according to position information of the electronic device and position information from the head-mounted device, wherein the position information of the head-mounted device is detected by a spatial attitude sensor in the head-mounted device, and the position information of the electronic device is detected by a spatial attitude sensor in the electronic device.
[0040] In a fourth aspect, a system is provided, the system including an electronic device and a head-mounted device, and the electronic device can implement the method in any one of the first aspect.
[0041] In a fifth aspect, a system is provided, the system including an electronic device and a head-mounted device, and the head-mounted device can implement the method in any one of the second aspect.
[0042] In a sixth aspect, a system is provided, the system including an electronic device, a head-mounted device, and a server, and the electronic device can implement the method in any one of the third aspect.
[0043] In a seventh aspect, an apparatus is provided, the apparatus including units or modules for performing the method in any one of the first aspect, or performing the method in any one of the second aspect, or performing the method in any one of the third aspect.
[0044] In an eighth aspect, an apparatus is provided, the apparatus including one or more processors configured to perform the method in any one of the first aspect, or perform the method in any one of the second aspect, or perform the method in any one of the third aspect.
[0045] In a ninth aspect, a readable storage medium is provided, which stores a program or instructions, when the program or instructions are run on an apparatus, cause the apparatus to perform the method in any one of the first aspect, or perform the method in any one of the second aspect, or perform the method in any one of the third aspect.
[0046] In a tenth aspect, a chip system is provided, which comprises a processor for supporting a computer device to implement the method in any one of the first aspect, or implement the method in any one of the second aspect, or implement the method in any one of the third aspect.
[0047] In an eleventh aspect, a program product is provided, which comprises a program; when the program is run on a computer, cause the computer to perform the method in any one of the first aspect, or perform the method in any one of the second aspect, or perform the method in any one of the third aspect. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 A hardware structure schematic diagram of an electronic device provided in an embodiment of the present application;
[0049] Figure 2 A software structure schematic diagram of an electronic device provided in an embodiment of the present application;
[0050] Figure 3 A structure of an earphone and a schematic diagram of a wearing state of the earphone in an embodiment of the present application;
[0051] Figure 4 A schematic diagram of an application scenario in an embodiment of the present application;
[0052] Figure 5 An example interface for prompting a user to perform a first operation through an interface in an embodiment of the present application;
[0053] Figure 6 A schematic diagram of human-computer interaction provided in an embodiment of the present application;
[0054] Figure 7 A schematic diagram of another human-computer interaction provided in an embodiment of the present application;
[0055] Figure 8 A schematic diagram of a head-centered spatial coordinate system in an embodiment of the present application;
[0056] Figure 9 A flowchart of a HRTF generation method provided in an embodiment of the present application;
[0057] Figure 10 A flowchart of another HRTF generation method provided in an embodiment of the present application;
[0058] Figure 11 A flowchart of another HRTF generation method provided by an embodiment of the present application is shown in FIG. 6.
[0059] Figure 12 A schematic diagram of another application scenario in an embodiment of the present application is shown in FIG. 7.
[0060] Figure 13 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 8. DETAILED DESCRIPTION
[0061] In the following, some terms in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.
[0062] At least one of the embodiments of the present application includes one or more, and the plurality refers to two or more. In addition, it should be understood that in the description of the present application, the terms "first", "second", etc. are used only for the purpose of distinguishing the description, and cannot be understood as indicating or implying relative importance, nor can it be understood as indicating or implying an order. For example, the first operation and the second operation do not represent the importance or order of the two, but only distinguish the description. In the embodiments of the present application, "and / or" is only used to describe the association relationship, which means that there are three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper generally represents an "or" relationship between the front and rear associated objects.
[0063] In the description of the embodiments of the present application, it should be noted that, unless otherwise explicitly specified and limited, the terms "mounting", "connecting" should be understood in a broad sense, for example, "connecting" can be detachable connection, or can be non-detachable connection; can be direct connection, or can be indirect connection through intermediate medium. The orientation terms mentioned in the embodiments of the present application, such as "up", "down", "left", "right", "inner", "outer" and the like, are only the direction of the drawing, therefore, the orientation terms used are for better and clearer illustration and understanding of the embodiments of the present application, and are not intended to indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore, it cannot be understood as a limitation on the embodiments of the present application.
[0064] Reference within this specification to "one embodiment" or "an embodiment" or "a specific embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" or "in some embodiments" or "in other embodiments" or "in still other embodiments" or other similar phrases in various places in this specification are not necessarily all referring to the same embodiment, but can refer to different embodiments. The term "comprising" or "comprises" or "including" or "includes" or "containing" or "contains" or "characterized by" or "complemented by" means "including, but not limited to" unless otherwise noted.
[0065] Currently, the method for obtaining HRTF needs to use special measuring equipment and experimental environment to measure the target individual (i.e., the user), which is complex.
[0066] In view of this, the embodiments of the present application provide a method for generating HRTF and related devices and systems that can implement the method, to generate personalized HRTF based on human-computer interaction. In the method, the target individual wears a head-mounted device, the electronic device plays a test audio signal, the head-mounted device collects a feedback audio signal of the test audio signal, and based on the test audio signal and the corresponding feedback audio signal, the personalized HRTF corresponding to the target individual can be obtained.
[0067] The head-mounted device described above can be worn on the head of the user, has an audio playing function and an audio signal collecting function. Optionally, it also has a function of determining HRTF according to the test audio signal and the feedback audio signal. In some embodiments, the head-mounted device can be earphones, or smart glasses, or a virtual reality (VR) helmet or VR glasses, or an augmented reality (AR) helmet or AR glasses, or a mixed reality (MR) helmet or MR glasses, etc. In general, the type of the head-mounted device is not limited in the present application.
[0068] The electronic device described above has a function of playing a test audio signal, and optionally, also has a function of determining HRTF according to the test audio signal and the feedback audio signal collected from the head-mounted device. In some embodiments, the electronic device can be a mobile phone, a tablet computer, a notebook computer, etc. portable electronic device; it can also be a watch, a bracelet, etc. wearable device; or it can also be a vehicle-mounted device, etc. In general, the type of the electronic device is not limited in the present application.
[0069] Figure 1 The hardware structure schematic diagram of the electronic device provided by the embodiments of the present application is shown. As shown in FIG. 1, the electronic device includes a processor 101, a memory 102, a power supply 103, a bus 104, a communication interface 105, a display 106, a sensor 107, and an input device 108. Figure 1As shown, the electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, a space attitude sensor, etc. The space attitude sensor can be, for example, an inertial measurement unit (IMU).
[0070] An IMU is a device for detecting the spatial attitude of a carrier, for example, can detect the three-axis angular velocity and acceleration of the carrier, etc. Spatial pose sensor An IMU is a device for detecting the spatial attitude of a carrier, for example, can detect the three-axis angular velocity and acceleration of the carrier, etc.
[0071] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated in one or more processors. The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching instructions and executing instructions. The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can directly call from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system. The execution of the audio playing method in the embodiments of the present application can be controlled by the processor 110 or other components to complete, for example, calling the processing program of the embodiments of the present application stored in the internal memory 121, or calling the processing program of the embodiments of the present application stored in the third party device through the external memory interface 120, so that the electronic device 100 can automatically switch the audio playing mode, solving the problem of low efficiency caused by manually switching the audio playing mode in the prior art. In addition, in the embodiments of the present application, the electronic device 100 can adaptively adjust the audio playing state or the audio playing volume when it is determined that the user has the intention to switch the audio playing mode (for example, the electronic device 100 is close to or away from the user's head), for example, when switching the audio playing mode (for example, from the first audio playing mode to the second audio playing mode), pausing the playing of the audio, or gradually reducing the volume of the audio played in the first audio playing mode and gradually reducing the volume of the audio played in the second audio playing mode, so that the user experience is not affected when the audio playing mode is switched due to the sudden change of the audio playing volume, and the user's needs are adapted.
[0072] The internal memory 121 can be used to store computer executable program codes including instructions. The processor 110 performs various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system and software codes of at least one application program (e.g., an iQiyi application, a WeChat application, etc.), etc. The data storage area can store data (e.g., images, videos, etc.) generated during use of the electronic device 100, etc. In addition, the internal memory 121 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one of a magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0073] The external memory interface 120 can be used to connect an external memory card such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to perform data storage functions. For example, files such as pictures and videos can be saved in the external memory card.
[0074] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface 130, etc.
[0075] The I2C interface is a bidirectional synchronous serial bus including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can include multiple sets of I2C buses. The processor 110 can be coupled to a touch sensor 180K, a charger, a flash, a camera 193, etc. via different I2C bus interfaces, respectively.
[0076] I2S interface can be used for audio communication. In some embodiments, processor 110 can include multiple sets of I2S bus. Processor 110 can be coupled with audio module 170 through I2S bus to enable communication between processor 110 and audio module 170. In some embodiments, audio module 170 can deliver audio signals to wireless communication module 160 through I2S interface to enable the function of answering phone through Bluetooth earphone.
[0077] PCM interface can also be used for audio communication, which samples, quantizes and encodes analog signals. In some embodiments, audio module 170 can be coupled with wireless communication module 160 through PCM bus interface. In some embodiments, audio module 170 can also deliver audio signals to wireless communication module 160 through PCM interface to enable the function of playing music through Bluetooth earphone. Both I2S interface and PCM interface can be used for audio communication.
[0078] UART interface is a universal serial data bus used for asynchronous communication. The bus can be a bidirectional communication bus. It converts data to be transmitted between serial communication and parallel communication. In some embodiments, UART interface is usually used to connect processor 110 and wireless communication module 160. For example, processor 110 communicates with Bluetooth module in wireless communication module 160 through UART interface to enable Bluetooth function. In some embodiments, audio module 170 can deliver audio signals to wireless communication module 160 through UART interface to enable the function of playing music through Bluetooth earphone.
[0079] MIPI interface can be used to connect processor 110 and peripheral devices such as display screen 194 and camera 193. MIPI interface includes camera serial interface (CSI), display serial interface (DSI) and the like. In some embodiments, processor 110 and camera 193 communicate through CSI interface to enable the shooting function of electronic device 100. Processor 110 and display screen 194 communicate through DSI interface to enable the display function of electronic device 100.
[0080] GPIO interface can be configured by software. GPIO interface can be configured as control signal or data signal. In some embodiments, GPIO interface can be used to connect processor 110 and camera 193, display screen 194, wireless communication module 160, audio module 170, sensor module 180 and the like. GPIO interface can also be configured as I2C interface, I2S interface, UART interface, MIPI interface and the like.
[0081] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection modes or a combination of multiple interface connection modes in the above embodiments.
[0082] The electronic device 100 can realize the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc. Among them, the ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing, and converts it into an image visible to the naked eye. The ISP can also optimize the algorithm of the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be arranged in the camera 193.
[0083] The electronic device 100 can realize the audio function through the audio module 170, the speaker 170A, the earpiece 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, music playing, recording, etc.
[0084] The audio module 170 is used to convert a digital audio signal into an analog audio signal output, and is also used to convert an analog audio signal into a digital audio signal. The audio module 170 can also be used to encode and decode an audio signal. In some embodiments, the audio module 170 can be arranged in the processor 110, or part of the function modules of the audio module 170 can be arranged in the processor 110.
[0085] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into an acoustic signal. The electronic device 100 can listen to music or listen to an external call through one or more speakers 170A in an external call scene.
[0086] In some embodiments, the speaker 170A and / or the earpiece 170B can include a single channel or multiple channels. In some embodiments, the multiple channels are used to provide the effect of stereo sound. In some other embodiments, the multiple channels can be combined. Taking a dual-channel as an example, the left channel and the right channel can play the same sound, which can increase the volume size.
[0087] The microphone 170C, also referred to as a "microphone" or a "microphone", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, a user can speak into the microphone 170C by placing the mouth close to the microphone 170C, and input a sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones, which can realize a noise reduction function in addition to collecting a sound signal. In other embodiments, the electronic device 100 can be provided with three, four or more microphones, which can realize a sound signal collection function, a noise reduction function, and can also identify a sound source to realize a directional recording function, etc.
[0088] The earphone interface 170D is used to connect a wired earphone. The earphone interface can be a USB interface, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0089] It can be understood that, Figure 1 The components shown do not constitute a specific limitation on the electronic device. The electronic device in the embodiments of the present application can include more or fewer components than those shown in the figure. In addition, Figure 1 the combination / connection relationship between the components in Figure 1 may also be adjusted and modified.
[0090] Figure 2 A software structure schematic diagram of an electronic device provided by an embodiment of the present application is shown. As Figure 2 shown, the software structure of the electronic device can be a layered architecture, for example, the software can be divided into several layers, each layer has a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the operating system is divided into four layers, from top to bottom, the application layer, the application framework layer (framework, FWK), the runtime and the system library, and the kernel layer.
[0091] The application layer can include a series of application packages. As Figure 2 shown, the application layer can include a camera, a setting, a skin module, a user interface (UI), a third-party application, etc. Among them, the third-party application can include a gallery, a calendar, a call, a map, a navigation, a wireless local area network (WLAN), Bluetooth, music, video, short message, etc.
[0092] The application framework layer provides an application programming interface (API) and a programming framework for applications of the application layer. The application framework layer can include some pre-defined functions. As shown in Figure 2 the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, and a notification manager.
[0093] The runtime includes a core library and a virtual machine. The runtime is responsible for scheduling and management of the operating system.
[0094] The core library contains two parts: one part is the function function that the java language needs to call, and the other part is the core library of the operating system. The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java file of the application layer and the application framework layer into a binary file. The virtual machine is used to perform object lifecycle management, stack management, thread management, security and exception management, and garbage collection functions.
[0095] The system library can include a plurality of functional modules. For example: a surface manager, media libraries, a three-dimensional graphics processing library (such as OpenGL ES), a 2D graphics engine (such as SGL), etc.
[0096] The surface manager is used to manage the display subsystem and provides 2D and 3D layer fusion for multiple applications.
[0097] The media library supports a variety of commonly used audio, video format playback and recording, and static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H. 264, MP3, AAC, AMR, JPG, PNG, etc.
[0098] The three-dimensional graphics processing library is used to realize three-dimensional graphics drawing, image rendering, synthesis, and layer processing, etc.
[0099] The 2D graphics engine is a drawing engine for 2D drawing.
[0100] The kernel layer is a layer between hardware and software. The kernel layer at least contains display driver, camera driver, audio driver, sensor driver.
[0101] The hardware layer can include various sensors, such as acceleration sensor, gravity sensor, touch sensor, etc.
[0102] Based on the above Figure 2In a possible implementation of the present application, the application program can be installed in the electronic device, and the application program can be used to implement the HRTF generation method provided by the embodiments of the present application. For example, the first application program can be included in the application program layer of the electronic device, and the first application program can be used to implement the HRTF generation method provided by the embodiments of the present application. Accordingly, the related layer in the operating system can provide an interface for the first application program, so that the first application program can call the system function through the interface to generate the HRTF.
[0103] Based on the above Figure 2 In another possible implementation of the present application, the operating system can be enhanced to have the related function of the HRTF generation method provided by the embodiments of the present application. For example, a service module or a function module can be added in one or more of the application program framework layer, the system library layer and the kernel layer, and the service module or the function module can be used to implement the related function of the HRTF generation method provided by the embodiments of the present application.
[0104] Taking the headset as an example of the head-mounted device, Figure 3 The structure of the headset and the schematic diagram of the headset wearing state are exemplarily shown. As shown in Figure 3 As shown, the processor 31, the loudspeaker 32 and the first microphone 33 are arranged in the left ear headset 30.
[0105] The loudspeaker 32 can play the audio signal processed by the HRTF under the control of the processor 31, and the first microphone 33 can collect the test audio signal emitted by the electronic device.
[0106] As shown in Figure 3 When the user wears the headset 30, the audio signal played by the electronic device is reflected and / or diffracted by the human body parts such as the pinna, the head and the torso of the user, and the first microphone 33 in the headset 30 can collect the audio signal after the reflection and / or diffraction, which is referred to as the feedback audio signal in the embodiments of the present application.
[0107] Optionally, the second microphone 34 can also be arranged in the headset 30, and the second microphone 34 can be arranged on the outside of the headset. When the user wears the headset, the second microphone 34 is closer to the mouth of the user than the first microphone 33, so as to collect the user's speech signal.
[0108] It should be understood that the right ear headset and the left ear headset have similar structures, which will not be repeated.
[0109] It should be understood that the right earphone and the left earphone have similar structures, and the earphone described below can be interpreted as the left earphone, the right earphone, or as a pair of earphones.
[0110] In addition, other devices or modules can be installed inside the headphones, such as communication modules, battery modules, etc.
[0111] It should be understood that although the above description is based on in-ear headphones, this application can also be applied to semi-in-ear headphones, open-back headphones, or similar head-mounted devices.
[0112] See Figure 4 This is a schematic diagram illustrating an application scenario in an embodiment of this application. The user's head 301 is fitted with earphones 302, each with an independent microphone (also called a transducer). The user holds a mobile phone 303, the distance between the phone 303 and the user's head 301 being approximately an arm's length, for example, about 40 to 50 centimeters. A first communication connection is established between the mobile phone 303 and the earphones 302. This first communication connection can be a wired connection or a wireless connection. The wireless connection can be a short-range communication connection, such as a Bluetooth connection, a Wi-Fi connection, or a Nearlink connection, etc., and this application does not impose any limitations on this.
[0113] The HRTF generation process can include a data acquisition phase and a calculation phase. In the data acquisition phase, the earphone 302 worn on the head 301 can acquire the feedback audio signal of the test audio signal emitted by the mobile phone 303. In the calculation phase, the HRTF can be calculated based on the acquired data.
[0114] like Figure 4 As shown in (a), during the data acquisition phase, the user's head 301 remains stationary, while the user slowly moves the mobile phone 302 around the head 301. The direction of movement can be roughly horizontal. To obtain more data from various locations to improve the accuracy of HRTF, the direction of movement can include both horizontal and vertical directions. The range of movement in the horizontal direction can be the maximum range of rotation that the user can achieve, and can even include the area behind the user's head. For example, with the head 301 as the origin of the coordinate system and the direction directly in front of the head 301 as the positive direction of the coordinate axis, the user's range of movement in the horizontal direction can be required to be no less than [-60 degrees, 60 degrees]. The range of rotation in the vertical direction can be the maximum range of rotation that the user can achieve. For example, with the user's head 301 as the origin of the coordinate system and the direction directly in front of the head 301 as the positive direction of the coordinate axis, the user's range of rotation in the vertical direction can be required to be no less than [-60 degrees, 60 degrees].
[0115] During the data acquisition phase, the speaker of mobile phone 303 plays a test audio signal, and the microphone of earphone 302 collects the feedback signal of the test audio signal. After the earphone 302 performs analog-to-digital conversion on the feedback signal, it sends the digital information of the feedback signal to mobile phone 303 through the first communication connection with the mobile phone.
[0116] During the calculation phase, the mobile phone 303 generates the HRTF corresponding to the earphone 302 based on the test audio signal and the feedback audio signal from the earphone 302.
[0117] The calculation phase can be performed during the data acquisition phase, or it can be performed after the data acquisition phase.
[0118] The specific implementation process of the above data acquisition and calculation stages can be found below. Figure 9 The process is shown below.
[0119] HRTF is a set of filters that uses technologies such as interaural time delay (HDITD), interaural amplitude difference (IAD), and auricular frequency vibration to generate stereo sound effects. When sound is transmitted to the auricle, ear canal, and eardrum in the ear, the listener will have the feeling of surround sound. Through the digital signal processor (DSP) in the headphones, HRTF can process the sound source in real time.
[0120] After generating the HRTF, when the user listens to audio played by the mobile phone 303 using the headset 302, the mobile phone 303 can process the audio signal using the HRTF corresponding to the headset 302 stored in the mobile phone 303, and play the processed audio signal, so that the audio signal received by the headset 302 can present a realistic three-dimensional sound effect to the user. In another possible approach, the mobile phone 302 sends the generated HRTF to the headset 302 through the first communication connection. When the speaker of the headset 302 plays the audio signal, it can use the HRTF to process the received audio signal, thereby simulating the sound information heard by human binaural hearing and presenting a realistic three-dimensional sound effect to the user.
[0121] exist Figure 4 In another possible application scenario shown in (b) above, with Figure 4The difference between the application scenario shown in (a) in the above and the application scenario shown in (a) in the above is that: in the data collection stage, the user holds the mobile phone 302 and keeps it still, the user's head 301 slowly rotates, for example, it can rotate left and right or up and down, or rotate at will. The rotation range can be the maximum rotation range that the user can reach. For example, taking the user's head 301 as the coordinate system origin and the front of the head 301 as the positive direction of the coordinate axis, the user's head rotation range can be required to be not less than [-60 degrees, 60 degrees]. The specific implementation of the data calculation stage in this scenario is the same as the implementation of the above data calculation stage.
[0122] In another possible application scenario, the difference between the above application scenario and the above application scenario is that: in the calculation stage, the microphone (also known as the microphone) of the earphone 302 collects the feedback audio signal of the test audio signal, and then calculates the HRTF according to the test audio signal and the feedback audio signal. Optionally, the frequency response range of the test audio signal can be pre-set, or the mobile phone 303 can send the digital information of the test audio signal to the earphone 302 through the first communication connection. After generating the HRTF, when the user uses the earphone 302 to listen to the audio played by the mobile phone 303, the earphone 302 can process the received audio signal by using the generated HRTF, so as to simulate the sound information heard by the human ears and present the user with realistic three-dimensional sound effects.
[0123] In another possible application scenario, the difference between the above application scenario and the above application scenario is that: in the calculation stage, the mobile phone 303 sends the feedback audio signal from the earphone 302 to the cloud platform (or cloud server), and the cloud platform generates the HRTF according to the test audio signal and the feedback audio signal, and sends the HRTF to the mobile phone 303. Optionally, the mobile phone 303 can send the HRTF to the earphone 302.
[0124] Optionally, the frequency response range of the test audio signal can be pre-set, or the mobile phone 303 can send the test audio signal to the cloud platform.
[0125] Taking the scenario shown in the above Figure 4 For example, in order to obtain the HRTF at multiple spatial positions based on the movement of the mobile phone 303 relative to the head 301 (or the earphone 302), the electronic device can output first prompt information, and the first prompt information prompts the user to perform a first operation, and the relative movement between the mobile phone 303 and the earphone 302 can be generated by performing the first operation.
[0126] For example, the first prompt information can prompt the user to perform the first operation as shown in (a) in the above: the user holds the mobile phone 303 and rotates it around the head 301 while keeping the head still (i.e., keeping the earphone 302 still). Figure 4 For example, the first prompt information can prompt the user to perform the first operation as shown in (a) in the above: the user holds the mobile phone 303 and rotates it around the head 301 while keeping the head still (i.e., keeping the earphone 302 still).
[0127] For another example, the first prompt information can prompt the user to perform the first operation as shown in (b) of FIG. 3: the user rotates the head 301 while holding the handset 303 still. Figure 4
[0128] In a possible implementation, the first prompt information is further used to indicate one or more of the following:
[0129] - the distance between the handset 303 and the earphone 302 (or the head 301). For example, the first prompt information can prompt the distance between the handset 303 and the head 301 to be one arm or so.
[0130] - the rotation direction. For example, in the scenario shown in (a) of FIG. 3, the first prompt information can prompt the user to rotate the handset around the head to the left and / or to the right, or to rotate the handset to the left first and then to the right. Figure 4 Figure 4 - the rotation angle or rotation range. For example, in the scenario shown in (a) of FIG. 3, the first prompt information can prompt the user to rotate the handset around the head to the left by 60 degrees and to the right by 60 degrees. For another example, in the scenario shown in (b) of FIG. 3, the first prompt information can prompt the user to hold the handset still, and rotate the head to the left by 45 degrees and to the right by 45 degrees.
[0131] Figure 4 Figure 4
[0132] - the duration. The first prompt information can prompt the user for the duration of performing the first operation.
[0133] In the embodiments of this application, the electronic device can adopt a single prompt mode or a combination of multiple prompt modes to instruct the user to perform the first operation. For example, the electronic device can prompt the user to perform the first operation by using the following prompt modes:
[0134] - a voice prompt mode: the electronic device can play voice prompt information to instruct the user to perform the first operation.
[0135] - an interface prompt mode: the electronic device can display the first prompt information on an interface to instruct the user to perform the first operation. Optionally, the first prompt information can include one or more of text, image, and animation. The image and animation can show the relative motion between the electronic device and the head-mounted device (or the head), for example, the relative motion shown in (a) or (b) of FIG. 3. Figure 4
[0136] Exemplarily, Figure 5 Several examples of prompting a user to perform a first operation through an interface are shown.
[0137] As shown in (a) of FIG. 5, the interface 510 includes a text prompt area 511 and a first control 512, and optionally, a second control 513. The text prompt area 511 displays information prompting the user to perform a first operation, for example, the following content can be prompted: “Please hold the phone parallel to the line of sight, and keep a one-arm distance from the head. After clicking the first control, keep the phone still, and slowly turn the head to the left and then to the right.” The first control 512 is used to trigger an HRTF generation process, for example, “Start test” can be displayed on the first control 512. After the first control 512 is triggered, the phone enters a data collection stage, and the phone can send an instruction to the earphone through a first communication connection, so that the earphone also enters the data collection stage. The second control 513 is used to end the HRTF generation process, for example, “End test” can be displayed on the second control 513. After the second control 513 is triggered, the phone ends the data collection stage and the calculation stage, and the phone can send an instruction to the earphone through the first communication connection, so that the earphone also ends the data collection stage and the calculation stage. Figure 5 Optionally, the interface 510 can also include a third control 514, which is used to cancel the operation of obtaining the HRTF, or in other words, to close the interface 510. Considering that the same earphone can be used by different users, if the user has generated an HRTF before using the earphone this time, and the HRTF is the personalized HRTF of the user, the third control 514 can be triggered to cancel the operation of generating the personalized HRTF this time, and directly enter the use stage of the earphone.
[0138] As shown in (b) of FIG. 5, the interface 520 includes a text prompt area 521, an animation area 522, and a first control 523, and optionally, a second control 524. The text prompt area 521 displays information prompting the user to perform a first operation, and the animation area 522 displays an animation example of the first operation. For example, the animation area 522 displays an animation example of the first operation as shown in (b) of FIG. 5, and the text prompt area 521 can prompt the following content: “After clicking the first control, keep the phone still, and slowly turn the head to the left and then to the right, as shown in the following figure.”
[0139] Figure 5 The first control 523 is used to trigger an HRTF generation process, for example, “Start test” can be displayed on the first control 523. The second control 524 is used to end the HRTF generation process, for example, “End test” can be displayed on the second control 524. The functions of the first control 523 and the second control 524 can be referred to the above embodiments. Figure 4
[0140] Optionally, the interface 510 or the interface 520 can further include a third control 525 for canceling the operation of acquiring the HRTF, or in other words, for closing the interface 520.
[0141] In another possible implementation, the interface 510 and the interface 520 can not include the first control and the second control. After the interface 510 or the interface 520 is opened, a time counting is started, and after a first time length, the data collection stage is automatically entered, and after a second time length after entering the data collection stage, the data collection stage and the calculation stage are automatically ended.
[0142] It should be understood that, Figure 5 The present application is not limited to the above examples of interfaces.
[0143] In a possible implementation, before outputting the first prompt information, the electronic device can first output second prompt information in a voice manner and / or an interface manner, the second prompt information being used to instruct the user to wear the head-mounted device. When the electronic device detects that the wearing state of the head-mounted device is “wearing”, the first prompt information is outputted.
[0144] Taking the application scenario of the above mobile phone and earphone as an example, the triggering manner of the data collection stage can include the following:
[0145] Triggering manner 1:
[0146] After the mobile phone and the earphone establish the first communication connection, the mobile phone and the earphone enter the data collection stage.
[0147] For example, after the mobile phone and the earphone establish the first communication connection, the mobile phone and the earphone can automatically enter the data collection stage.
[0148] For another example, after the mobile phone and the earphone establish the first communication connection, the mobile phone can output the first prompt information, and enter the data collection stage in response to a user operation (for example, an operation of triggering the first control).
[0149] For example, after the distance between the mobile phone and the earphone satisfies the distance specified by the first communication protocol, the mobile phone and the earphone can automatically establish the first communication connection, or the mobile phone can establish the first communication connection with the earphone based on a user setting operation.
[0150] Triggering manner 2:
[0151] The user starts a first application program (which is used to acquire the HRTF) on the mobile phone, and the screen of the mobile phone displays a user interface of the first application program, and the user interface can display the above-mentioned first prompt information. For example, after the first application program is started, the above-mentioned interface 510 or interface 520 can be displayed.
[0152] In another possible implementation, when the first application is started, the mobile phone can output the first prompt information in a voice playing manner to trigger mode 3:
[0153] The user can issue a first voice instruction, which is used to trigger data collection, or in other words, which is used to trigger the acquisition of the HRTF. The mobile phone can enter the data collection stage in response to the first voice instruction and send an instruction to the earphone through the first communication connection, so that the earphone also enters the data collection stage.
[0154] Optionally, after the mobile phone responds to the first voice instruction, it can first prompt the user to perform the first operation described above in a voice broadcast manner or by displaying the first prompt information on the screen.
[0155] Mode 4:
[0156] The mobile phone enters the data collection stage when it detects that the wearing state of the earphone is wearing.
[0157] For example, when the mobile phone detects that the wearing state of the earphone is wearing, it can first output the first prompt information, then enter the data collection stage in response to a user operation (for example, an operation of triggering a first control), or enter the data collection stage after a set time period and send an instruction to the earphone through the first communication connection, so that the earphone also enters the data collection stage.
[0158] It can be understood that the above exemplary lists several possible triggering modes of the data collection stage, and the present application does not limit this.
[0159] Taking the application scenario of the above mobile phone and earphone as an example, the ending mode of the data collection stage can include the following:
[0160] Ending mode 1:
[0161] After the mobile phone enters the data collection stage, it automatically ends the data collection stage after a set time period, for example, the mobile phone ends playing the test audio signal. For example, the mobile phone can be provided with a timer, which is started when the mobile phone enters the data collection stage, and when the timer times out, the data collection stage ends. Optionally, the countdown time of the timer can be displayed on the screen of the mobile phone, or the progress of the data collection stage can be displayed through a progress bar.
[0162] Optionally, the earphone enters the data collection stage, and automatically ends the data collection stage after the set time period.
[0163] Optionally, when the mobile phone automatically ends the data collection stage, it can send an instruction to the earphone through the first communication connection, which is used to instruct the earphone to end the data collection stage.
[0164] Ending mode 2:
[0165] The first user operation ends the collection phase, for example, the phone ends playing the test audio signal.
[0166] For example, after the phone enters the data collection phase, the phone displays a user interface on the screen, the user interface displays a "end test" control, when the control is triggered, the phone ends the data collection phase, and sends an instruction to the earphone through the first communication connection, to instruct the earphone to end the data collection phase.
[0167] For another example, the phone displays a user interface on the screen after entering the data collection phase, the user interface displays a "end test" control, when the control is triggered, the phone ends the data collection phase, and sends an instruction to the earphone through the first communication connection, to instruct the earphone to end the data collection phase. Figure 5 For example, after the user presses the first control of "start test", the first control is kept pressed during the data collection phase, when the user lifts the first control, the phone ends the data collection phase, and sends an instruction to the earphone through the first communication connection, to instruct the earphone to end the data collection phase.
[0168] End mode 3:
[0169] The user can issue a second voice instruction, the second voice instruction is used to end the data collection. The phone can end the data collection phase in response to the second voice instruction, and send an instruction to the earphone through the first communication connection, so that the earphone also enters the data collection phase.
[0170] In a possible implementation, after the electronic device obtains the HRTF, or the head-mounted device notifies the electronic device that the HRTF is obtained, the electronic device can output third prompt information, the third prompt information is used to instruct the user to end the first operation, and / or is used to indicate that the HRTF has been generated. Optionally, the electronic device can output the third prompt information in the form of playing a voice and / or in the form of an interface prompt.
[0171] It should be understood that the above-mentioned triggering mode and stopping mode of data collection can be combined with each other. The embodiments of the present application do not limit the combination of the triggering mode and the stopping mode, for example, the triggering mode 1 and the ending mode 1 can be combined; for another example, the triggering mode 2 and the ending mode 2 can be combined; for another example, the triggering mode 3 and the ending mode 3 can be combined.
[0172] In the following, several user interface change processes involved in HRTF detection based on human-computer interaction are given in combination with Figure 6 and Figure 7 .
[0173] Referring to Figure 6Fig. 6 is a schematic diagram of an embodiment of the present application, which illustrates a process of obtaining HRTF. As shown in Fig. 6, a user selects the earphone 302 from the list of available devices in the communication connection setting interface 610 of the mobile phone 303. In response to the user operation, the mobile phone 303 establishes a first communication connection with the earphone 302. After the first communication connection is established, the mobile phone displays a prompt window 620, which displays second prompt information to prompt the user to wear the earphone. The user can wear the earphone according to the second prompt information. When the mobile phone detects that the earphone is worn, the mobile phone displays the interface 520. The user can click the first control 523 according to the prompt information in the interface 520, and perform a first operation according to the prompt information after clicking the first control 523. In response to the event that the first control 523 is triggered, the mobile phone plays a test audio signal, receives a feedback audio signal and position information from the earphone, and determines the HRTF of the user according to the test audio signal, the feedback audio signal and the position information.
[0174] Optionally, during the process of obtaining the HRTF, the mobile phone can display the motion of the head of the target individual and / or the motion of the hand holding the mobile phone in the user interface in real time based on the detected position information and the position information from the earphone, so as to facilitate the user to timely adjust the motion of the head or the motion of the hand.
[0175] Optionally, during the process of obtaining the HRTF, the mobile phone can output new prompt information to instruct the user to perform a corresponding operation according to the new prompt information. For example, if the mobile phone determines that the rotation angle of the mobile phone relative to the head is less than the rotation angle indicated by the first prompt information according to the position information from the earphone, the mobile phone can output new prompt information to instruct the user to rotate a larger angle. Optionally, the new prompt information can be prompted by playing a voice and / or by interface display.
[0176] Optionally, after the mobile phone obtains the HRTF, a window 630 can be displayed, which displays third prompt information that can be used to indicate that the HRTF has been generated and the user can use the earphone to listen to audio. The display window 630 can be automatically closed after a set time period, or the display window 630 is closed based on a user operation, for example, the display window 630 includes a control for "confirm" or "close", and when the control is triggered, the display window 630 is closed.
[0177] It should be understood that the "window" described above can be a floating window, which is replaced by "interface", and the present application does not limit this.
[0178] Referring to Figure 7Another schematic diagram of human-computer interaction provided by an embodiment of the present application is shown. If the HRTF has been saved in the mobile phone, the mobile phone can directly use the HRTF to process the played audio signal. If a new HRTF is needed, the related function of obtaining the HRTF can be called through voice, or the related function of obtaining the HRTF can be started through the setting interface, or the related function of obtaining the HRTF can be started through other manners. Figure 7 The application scenario of starting the related function of obtaining the HRTF through the setting interface is taken as an example for description.
[0179] As shown in Figure 7 , the mobile phone and the earphone have established a first communication connection, and the mobile phone has detected that the wearing state of the earphone is "wearing". The display interface 520 is displayed when the user opens the setting interface 710 on the mobile phone and triggers the "obtain HRTF" control in the setting interface. The subsequent operation can refer to the related content in Figure 6 , which is not repeated here.
[0180] The application scenario shown in Figure 4 , the relative spatial position change between the earphone 302 (or the head 301 of the user) and the mobile phone 303 occurs, for example, relative rotation or relative movement. The relative spatial position change can be described based on a spatial coordinate system with the head 301 as the center. Figure 8 A schematic diagram of a spatial coordinate system with the head 301 as the center is shown. The black dot in the coordinate system represents the spatial position of the mobile phone relative to the head, which can be represented by the distance from the head, the pitch angle relative to the head, the horizontal angle relative to the head, and the like. Taking Figure 8 , the spatial position A in the spatial coordinate system, as an example, the distance between the spatial position A and the head center point O is r, the pitch angle relative to the head is φ, and the horizontal angle relative to the head is θ.
[0181] In a possible implementation manner, the spatial attitude sensor (for example, IMU) is arranged in the mobile phone 303 and the earphone 302 respectively.
[0182] In the scenario of calculating the HRTF by the mobile phone 303, the earphone 302 can send the detection data of the spatial attitude sensor in the earphone 302 to the mobile phone 303 through the first communication connection. The mobile phone 303 determines the spatial position of the mobile phone 303 relative to the head 301 based on the detection data of the spatial attitude sensor in the mobile phone 303 and the detection data of the spatial attitude sensor in the earphone 302, and can calculate the HRTF based on the spatial position, or optimize the calculated HRTF based on the spatial position.
[0183] In the scenario where the earphone 302 calculates the HRTF, the mobile phone 303 can send the detection data of the spatial posture sensor in the mobile phone to the earphone 302 through the first communication connection, and the earphone 302 determines the spatial position of the mobile phone 303 relative to the head 301 according to the detection data of the spatial posture sensor in the earphone and the detection data of the spatial posture sensor in the mobile phone 303, and can calculate the HRTF according to the spatial position, or optimize the calculated HRTF according to the spatial position.
[0184] It can be understood that in some other application scenarios, the earphone can be replaced by other head-mounted devices, for example, it can be replaced by smart glasses, or a VR helmet or VR glasses, or an AR helmet or AR glasses, or an MR helmet or MR glasses, etc. Those skilled in the art can understand that in the case of replacing the mobile phone with other head-mounted devices, the prompt information of the electronic device (such as the mobile phone) can be adjusted accordingly. For example, in the case of replacing the earphone with smart glasses, the user interface displayed by the mobile phone can display prompt information prompting the user to correctly wear the glasses, and when it is detected that the wearing state of the glasses is "correct", the mobile phone displays prompt information prompting the user to rotate the head or hold the mobile phone around the head.
[0185] It can also be understood that in some other application scenarios, the mobile phone can be replaced by other electronic devices, for example, it can be replaced by a tablet computer, or a smart watch, or a smart bracelet, or other wearable devices.
[0186] Based on the above application scenarios, Figure 9 A flowchart of a HRTF generation method provided by an embodiment of the present application is shown. The electronic device in the flowchart can be a mobile phone or the like in the above application scenarios, and the head-mounted device in the flowchart can be an earphone or the like in the above application scenarios. Before executing the following flow, the head-mounted device has been worn on the head of a target individual (i.e. a user), and the electronic device and the head-mounted device enter the data collection phase in the manner described in the above application scenarios. In the data collection phase, the electronic device and the head-mounted device move relative to each other in the manner described in the above application scenarios.
[0187] As shown in the figure, the method can include the following steps: Figure 9
[0188] Step 901: The electronic device plays a test audio signal.
[0189] In the embodiment of the present application, the test audio signal can also be referred to as a probe audio signal or a reference audio signal, etc., and the present application does not limit this.
[0190] Optionally, the test audio signal can be an audio signal with a frequency response range. For example, the frequency response range can include the frequency range of sounds that can be heard by human beings, specifically, 20 Hz to 20 KHz; or a middle frequency band in the frequency range, for example, 50 MHz to 150 MHz.
[0191] Optionally, the test audio signal can be an audio signal with a frequency gradually increasing in order from low frequency to high frequency in a set frequency response range, or an audio signal with a frequency gradually decreasing in order from low frequency to high frequency in the frequency response range, or an audio signal with a frequency constantly cycling from high to low, or a piece of music melody with a frequency range in the set frequency response range. The application does not limit the test audio signal.
[0192] In this step, in the data collection phase, the electronic device plays the test audio signal through the loudspeaker thereof.
[0193] Optionally, in the data collection phase, the electronic device can continuously play the test audio signal.
[0194] Step 902: The head-mounted device worn on the head of the target individual collects a feedback audio signal of the test audio signal.
[0195] The head-mounted device can collect the feedback audio signal of the test audio signal generated at the ear of the target individual. In other words, the audio signal formed after the test audio signal played by the electronic device is reflected and / or diffracted by the ear and / or head of the target individual can be collected by the microphone of the head-mounted device.
[0196] In one possible implementation, taking earphones as an example, the microphone in the left earphone and the microphone in the right earphone respectively collect the feedback audio signal of the test audio signal. Since the earphones are worn on the head of the target individual, the microphone in the left earphone and the microphone in the right earphone are located at the entrance of the ear canal of the target individual, and thus the feedback audio signal collected by the microphone in the left earphone and the microphone in the right earphone can be a head related impulse response (HRIR), that is, the feedback audio signal is an audio signal generated after the test audio signal is reflected and / or diffracted by the ear, head and other human body parts of the target individual.
[0197] Step 903: The head-mounted device sends the collected feedback audio signal and the position information of the head-mounted device to the electronic device.
[0198] Optionally, the head-mounted device can collect the feedback audio signal according to a preset sampling frequency to obtain a time domain sequence of the feedback audio signal. Embodiments of the present application take the feedback audio signal corresponding to N sampling points in the time domain sequence of the feedback audio signal collected in the data collection stage as an example for description, where N is an integer greater than 1. Optionally, the N sampling points can be equal time intervals or can not be equal time intervals.
[0199] Optionally, the head-mounted device can send the feedback audio signal to the electronic device after the end of the data collection stage, or can send the feedback audio signal to the electronic device immediately after collecting the feedback audio signal, which is not limited in the present application.
[0200] In specific implementation, the head-mounted device can perform analog-to-digital conversion on the collected feedback audio signal of the analog signal type to obtain digital information of the feedback audio signal, and send the digital information of the feedback audio signal to the electronic device through the first communication connection between the head-mounted device and the electronic device.
[0201] Optionally, the position information of the head-mounted device is detected by a spatial attitude sensor in the head-mounted device. Optionally, the position information further includes a time stamp.
[0202] Optionally, the head-mounted device can send the position information to the electronic device after the end of the data collection stage, or can send the position information to the electronic device immediately after obtaining the position information, which is not limited in the present application.
[0203] Step 904: The electronic device determines the HRTF corresponding to the spatial position of the electronic device relative to the head-mounted device according to the spatial position of the electronic device relative to the head-mounted device, and the test audio signal and the feedback audio signal corresponding to the spatial position.
[0204] In this step, taking the earphone as an example, the electronic device can determine the HRTF of the left earphone at the spatial position of the electronic device relative to the left earphone according to the spatial position of the electronic device relative to the left earphone, and the test audio signal and the feedback audio signal corresponding to the spatial position. Similarly, the electronic device determines the HRTF of the right earphone at the spatial position of the electronic device relative to the right earphone according to the spatial position of the electronic device relative to the right earphone, and the test audio signal and the feedback audio signal corresponding to the spatial position.
[0205] For example, taking the spatial position of the electronic device relative to the head-mounted device as the first spatial position, the test audio signal corresponding to the first spatial position refers to the test audio signal played by the electronic device when the spatial position of the electronic device relative to the head-mounted device is the first spatial position; and the feedback audio signal corresponding to the first spatial position refers to the feedback audio signal collected by the head-mounted device when the spatial position of the electronic device relative to the head-mounted device is the first spatial position.
[0206] In the data collection phase, the electronic device and the head-mounted device move relatively, and in the time-domain sequence of the feedback audio signals collected by the head-mounted device, the feedback audio signal corresponding to each sampling point is related to the spatial position of the electronic device relative to the head-mounted device at the corresponding time, or in other words, the feedback audio signal corresponding to each sampling point is associated with the spatial position of the electronic device relative to the head-mounted device at the time. The electronic device can determine, for each sampling point, the HRTF at the spatial position corresponding to the time according to the test audio signal and the feedback audio signal at the time.
[0207] In a possible implementation, the electronic device can determine, based on a clock of the electronic device, a timestamp of the test audio signal corresponding to each sampling point in the time-domain sequence of the test spectral signal played by the electronic device, and the feedback audio signal received by the electronic device from the head-mounted device can include a timestamp of the feedback audio signal corresponding to each sampling point in the time-domain sequence of the feedback audio signal. In addition, the position information detected by the spatial posture sensor of the electronic device includes a timestamp, and the position information received by the electronic device from the head-mounted device includes a timestamp, so that the electronic device can determine, according to the timestamp, the test audio signal, the feedback audio signal, and the spatial position corresponding to each sampling point, and thus can calculate the HRTF at the spatial position according to the test audio signal, the feedback audio signal, and the spatial position corresponding to each sampling point.
[0208] In a possible implementation, the spatial position of the electronic device relative to the head-mounted device can be calculated according to the position information detected by the spatial posture sensor of the electronic device and the position information received from the head-mounted device. Optionally, the spatial position of the electronic device relative to the head-mounted device can be calculated in a coordinate system with the head of the target individual as the center.
[0209] In a possible implementation, the electronic device can perform convolution operation on the test audio signal and the feedback audio signal, to obtain the HRTF.
[0210] For example, for a first spatial position on the relative motion path of the electronic device and the head-mounted device, the electronic device can perform Fourier transform on the time-domain signal of the test audio signal and the time-domain signal of the feedback audio signal of the left earphone at the first spatial position, respectively, to obtain corresponding frequency-domain signals, and then multiply the two frequency-domain signals, and the result can be used as the HRTF of the left earphone at the spatial position. Similarly, for the first spatial position, the electronic device can perform Fourier transform on the time-domain signal of the test audio signal and the time-domain signal of the feedback audio signal of the right earphone at the first spatial position, respectively, to obtain corresponding frequency-domain signals, and then multiply the two frequency-domain signals, and the result can be used as the HRTF of the right earphone at the spatial position.
[0211] Considering the limited spatial range of relative motion between the electronic device and the head-mounted device, in some embodiments, after obtaining the HRTF corresponding to multiple spatial positions based on the test audio signal and the feedback audio signal, an interpolation algorithm can be used to calculate the HRTF corresponding to even more spatial positions to optimize the HRTF. This improves the user's listening experience when playing audio based on the optimized HRTF. For example, the spatial range of relative motion between the electronic device and the head-mounted device is a 45-degree horizontal and vertical space centered on the target individual's head. After interpolation, an HRTF of a 90-degree horizontal and vertical space in front of the target individual, or an HRTF of an even larger spatial range, can be obtained.
[0212] Considering that head-mounted devices may have multiple microphones, for example, Figure 3 Taking the illustrated headset as an example, a first microphone (e.g., a feedback microphone) and a second microphone (e.g., a call microphone) are respectively installed at different positions in the left and right earpieces. Therefore, in some embodiments, the microphones at different positions in the head-mounted device can separately collect test audio signals, and the head-mounted device can feed back the feedback audio signals collected by the microphones at different positions to the electronic device. For a certain spatial position in the relative movement of the electronic device and the head-mounted device, the electronic device can determine the HRTF corresponding to that spatial position based on the test audio signal and the feedback audio signals collected by the different microphones.
[0213] For example, an electronic device can calculate a first HRTF (also called a preliminary HRTF) based on a test audio signal and a feedback audio signal collected by a first microphone (e.g., a feedback microphone) in the left earphone. Then, it can calculate a second HRTF based on the test audio signal and a feedback audio signal collected by a second microphone (e.g., a call microphone) in the left earphone. The first HRTF is then corrected using the second HRTF to obtain the HRTF corresponding to the left earphone. The accuracy of the HRTF obtained in this way is typically higher than the accuracy of the uncorrected HRTF.
[0214] Similarly, the HRTF corresponding to the right earphone can also be calculated in the same way as above.
[0215] It should be understood that although the above description uses the example of a first microphone and a second microphone in the right earphone or the left earphone, in actual applications, if there are more microphones in the earphone, the HRTF can also be calculated using the above method.
[0216] The test audio signal is collected and fed back by the microphones at different positions in the head-mounted device to serve as the basis for calculating the HRTF. Compared with collecting and feeding back the test audio signal by using a single microphone, the accuracy of the HRTF can be improved, and thus the user's audio listening experience can be improved when the head-mounted device plays audio based on the HRTF.
[0217] In view of the fact that the characteristics (or geometric characteristics) of the head-mounted device can affect the HRTF, in some embodiments, after the HRTF is calculated, the calculated HRTF can be corrected based on the characteristic parameters of the head-mounted device to improve the accuracy of the HRTF.
[0218] Optionally, the characteristic parameters of the head-mounted device can include one or more of the following:
[0219] The size of the head-mounted device;
[0220] The shape or structure of the head-mounted device;
[0221] The material of the head-mounted device;
[0222] The position of the microphone in the head-mounted device;
[0223] The acoustic response characteristic parameters of the head-mounted device, for example, the acoustic characteristic parameters can include frequency response, phase, etc. The acoustic characteristic parameters can be obtained from the relevant manual of the head-mounted device, or by detecting the head-mounted device by professional equipment. In this way, the HRTF detected by the microphone can be compensated according to the frequency response and other characteristics of the microphone in the head-mounted device.
[0224] In order to improve the accuracy of the HRTF, in some embodiments, after the HRTF corresponding to the target individual is calculated, the calculated HRTF can be corrected based on the human body characteristics (such as ear structure characteristics and / or head structure characteristics) of the target individual, thereby improving the accuracy of the HRTF.
[0225] Optionally, the image of the target individual can be collected by using the front camera on the electronic device, and the ear structure characteristics (such as the shape of the pinna, the distance between the two ears, etc.) and / or the head structure characteristics of the target individual can be obtained by analyzing or recognizing the image.
[0226] Optionally, in the case where the user authorizes the use of the front camera in the data collection stage, the electronic device can automatically start the front camera and use the front camera to collect the image of the target individual. Optionally, in another scenario, the electronic device can obtain the user's use permission of the front camera through voice or interface, and start the front camera to collect the image of the target individual after obtaining the authorization.
[0227] Optionally, the inter-aural distance affects the HRTF, and in a possible implementation manner in which the HRTF is corrected based on the ear features and / or structural features of the target individual, the difference between the HRTFs of the two ears can be corrected according to the inter-aural distance, for example, the inter-aural time difference (ITD) and / or the inter-aural level difference (ILD).
[0228] To improve the accuracy of the HRTF, in some embodiments, after the HRTF is calculated according to the test audio signal and the feedback audio signal, the HRTF can be corrected based on an HRTF database.
[0229] The HRTF database can be obtained by performing acoustic-related tests on users with different human features. For example, the HRTF database can include HRTF models corresponding to a plurality of human features. Each HRTF model includes HRTFs corresponding to a plurality of spatial positions, for example, HRTFs corresponding to spatial positions in the spatial range shown in FIG. 6. Figure 8 Each human feature can correspond to a set of human feature information, for example, including pinna shape, inter-aural distance, head circumference, head shape, etc.
[0230] Optionally, the HRTF database can be obtained in various ways, for example, it can be downloaded from the network side or pre-set in the electronic device, which is not limited in the present application.
[0231] Taking the first HRTF corresponding to the first spatial position calculated according to the test audio signal and the feedback audio signal as an example, a possible implementation manner of correcting the HRTF based on the HRTF database is as follows: the electronic device can query the HRTF database according to the human feature information of the target individual, obtain a matched HRTF model, and obtain a second HRTF corresponding to the first spatial position in the HRTF model; if the error between the first HRTF and the second HRTF is within a set range, the first HRTF does not need to be corrected; if the error between the first HRTF and the second HRTF exceeds the set range, the first HRTF can be corrected according to the second HRTF, so that the error between the corrected first HRTF and the second HRTF is within the set range, for example, the first HRTF is corrected to the second HRTF.
[0232] Optionally, the human features of the target individual can be obtained in the following ways:
[0233] Manner 1: providing a user interface for a user to input his / her own anthropometric information based on the user interface, or to select anthropometric information matching the user from the anthropometric information options in the user interface.
[0234] Manner 2: using a camera of the electronic device to take a photo of the user, and obtaining the anthropometric information of the user based on image analysis.
[0235] For example, the electronic device can automatically collect the image of the target individual in the data collection stage.
[0236] Manner 3: prompting the user to input anthropometric information in a voice manner, and obtaining the anthropometric information input by the user in the voice manner.
[0237] With the above implementation manners, the HRTF obtained by the embodiments of the present application can present similar stereo effect as the HRTF obtained in the laboratory environment in the medium and high frequency bands.
[0238] With the above implementation manners, the accuracy of the HRTF can be improved. For example, the HRTF obtained by the embodiments of the present application can have similar stereo effect as the HRTF obtained in the laboratory environment in the medium and high frequency bands.
[0239] It should be understood that the above only exemplarily shows one possible implementation manner of correcting the calculated HRTF according to the HRTF queried from the HRTF database, and the implementation manner of how to correct the HRTF is not limited in the present application.
[0240] In one possible implementation manner, after the HRTF corresponding to the target individual is calculated, the HRTF of a larger spatial range can be obtained based on the HRTF database.
[0241] For example, the spatial range of the relative motion between the electronic device and the head-mounted device is a first spatial range (for example, a spatial range of 45 degrees left and right and 30 degrees up and down in front of the target individual) centered on the head of the target individual, and based on the test audio signal and the feedback audio signal, the HRTF corresponding to a plurality of spatial positions in the first spatial range can be obtained. Based on the HRTF database described above, the electronic device can query the HRTF database according to the human feature information of the target individual to obtain an HRTF model matched with the human feature information; and then obtain the HRTF corresponding to at least one spatial position in a second spatial range centered on the head of the target individual from the HRTF model. The second spatial range does not overlap with the first spatial range. For example, the electronic device can obtain the HRTF corresponding to a plurality of spatial positions other than the first spatial range in the spatial range centered on the head of the target individual from the HRTF model, so that the HRTF corresponding to the target individual includes the HRTF in all spatial ranges within a certain distance and centered on the head of the target individual.
[0242] In some embodiments, in the data collection stage, the target individual can be instructed to perform relative motion in the maximum spatial range as much as possible, such as being instructed to turn the mobile phone to the back of the head for testing, or being instructed to turn the head at the maximum angle for testing. In this way, the HRTF obtained by using the embodiments of the present application can have similar stereo sound effects as the HRTF obtained in a laboratory environment at a large angle, that is, the accuracy of the HRTF can be improved.
[0243] According to simulation experiments, based on the left ear HRTF and the right ear HRTF obtained by the process shown in Figure 9 the HRTF obtained in the laboratory environment, the HRTF in the low and medium frequency band of 2KHz can be obtained according to the above process; the HRTF in the frequency band of 2KHz to 6KHz has a high correlation with the HRTF obtained in the laboratory environment, and the obtained HRTF can be corrected by combining related modeling, simulation and fitting, etc. on the basis of the above process. In the frequency band above 6KHz, the accuracy of the HRTF can be improved by the above-mentioned correction method of the calculated HRTF.
[0244] In the above embodiments of the present application, the HRTF of the head-mounted device can be obtained by using the electronic device and the head-mounted device to perform simple testing, without the need for testing in a laboratory environment to obtain the HRTF, thereby simplifying the process of obtaining the HRTF and improving the convenience of obtaining the HRTF. In addition, the HRTF obtained by using the method provided in the embodiments of the present application does not require the user to provide user privacy information such as human feature parameters, and therefore the safety of the user privacy information can be protected.
[0245] Based on the above application scenarios, Figure 10 A flowchart of a method for generating HRTF is shown. The electronic device in the flowchart can be a mobile phone or other device in the above application scenarios. The head-mounted device in the flowchart can be a headset or other device in the above application scenarios. Before executing the following flowchart, the head-mounted device has been worn on the head of the target individual (i.e., the user), and the electronic device and the head-mounted device start entering the data collection phase in the manner described in the above application scenarios. In the data collection phase, the electronic device and the head-mounted device move relative to each other in the manner described in the above application scenarios.
[0246] As shown in Figure 10 , the method can include the following steps:
[0247] Step 1001: The electronic device plays a test audio signal.
[0248] The specific implementation of this step can refer to step 901 in Figure 9 .
[0249] Step 1002: The head-mounted device worn on the head of the target individual collects a feedback audio signal of the test audio signal.
[0250] The specific implementation of this step can refer to step 902 in Figure 9 .
[0251] Step 1003: The head-mounted device determines the HRTF corresponding to the spatial position of the electronic device relative to the head-mounted device, according to the test audio signal and the feedback audio signal corresponding to the spatial position.
[0252] The implementation of the head-mounted device determining the HRTF can refer to the implementation of the electronic device determining the HRTF in Figure 9 .
[0253] Optionally, the test audio signal can be pre-set in the head-mounted device, or the electronic device can send the digital information of the test audio signal to the head-mounted device through the first communication connection between the electronic device and the head-mounted device.
[0254] Based on the above application scenarios, Figure 11 A flowchart of a method for generating HRTF is shown. The electronic device in the flowchart can be a mobile phone or other device in the above application scenarios. The head-mounted device in the flowchart can be a headset or other device in the above application scenarios. Before executing the following flowchart, the head-mounted device has been worn on the head of the target individual (i.e., the user), and the electronic device and the head-mounted device start entering the data collection phase in the manner described in the above application scenarios. In the data collection phase, the electronic device and the head-mounted device move relative to each other in the manner described in the above application scenarios.
[0255] As shown in Figure 11 , the method can include the following steps:
[0256] Step 1101: The electronic device plays a test audio signal.
[0257] The specific implementation of this step can refer to step 901 in Figure 9 .
[0258] Step 1102: The head-mounted device worn on the head of the target individual collects a feedback audio signal of the test audio signal.
[0259] The specific implementation of this step can refer to step 902 in Figure 9 .
[0260] Step 1103: The head-mounted device sends the collected feedback audio signal and the position information of the head-mounted device to the electronic device.
[0261] The specific implementation of this step can refer to step 903 in Figure 9 .
[0262] Step 1104: The electronic device sends the test audio signal, the feedback audio signal, and the information of the spatial position of the electronic device relative to the head-mounted device to the server.
[0263] Step 1105: The server determines the HRTF corresponding to the spatial position of the head-mounted device relative to the electronic device, and the test audio signal and the feedback audio signal corresponding to the spatial position.
[0264] The implementation of the server determining the HRTF can refer to the implementation of the electronic device determining the HRTF in Figure 9 .
[0265] Step 1106: The server sends the HRTF to the electronic device.
[0266] The embodiments of the present application also provide a method for generating HRTF and related devices and systems that can implement the method, to generate personalized HRTF based on human-computer interaction mode. In the method, the electronic device instructs the head of the target individual to move relative to the electronic device based on the human-computer interaction mode, collects multi-directional images of the head and / or ear of the target individual by the electronic device, models the head based on the images, and obtains the personalized HRTF corresponding to the target individual based on the head modeling. The structure and type of the electronic device can refer to the foregoing embodiments.
[0267] See Figure 12This is a schematic diagram of an application scenario in an embodiment of this application. The user holds a mobile phone 1202, and the distance between the mobile phone 1202 and the user's head 1201 is approximately the length of an arm's length.
[0268] The HRTF generation process can include a data acquisition stage and a calculation stage. In the data acquisition stage, the mobile phone 1202 (e.g., the front-facing camera in the mobile phone) can acquire multi-angle images of the target individual's head. In the calculation stage, a three-dimensional model of the target individual's head can be generated based on the acquired multi-angle images of the target individual's head, and the personalized HRTF of the target individual can be determined based on the three-dimensional model.
[0269] like Figure 12 As shown in (a) above, during the data acquisition phase, the user's head 1201 remains stationary, while the user holds a mobile phone 1202 and moves it slowly around the head 1201. Figure 12 As shown in (b), during the data acquisition phase, the user's head 1201 rotates, while the user holds the mobile phone 1202 stationary. The direction of movement can be referenced... Figure 4 The relevant content in [the document / document].
[0270] During the computation phase, the mobile phone 1202 generates a head model of the target individual based on the acquired images, and then generates an HRTF based on the model.
[0271] In one possible implementation, during the data acquisition phase, as the phone 1202 rotates around the head, the spatial attitude sensor in the phone 1202 can detect the phone's position information (which reflects the relative positional relationship between the phone and the head), and model the head based on the acquired images and the detected position information. Since the spatial attitude sensor's detection data reflects the relative positional relationship between the phone and the head, that is, it can reflect information such as the depth of field of the head object in the acquired images, combining the position information and the images for analysis and modeling can improve modeling accuracy, thereby improving the accuracy of HRTF (Head-to-Face Transformation).
[0272] In another possible implementation, the target individual's head 1201 can wear headphones; other processing can be referenced. Figure 12 (a) shows the scenario. In this scenario, the earphone may partially obstruct the ear in the image captured by the mobile phone 1202. The structure of the entire auricle can be reconstructed based on the algorithm using the unobstructed part of the auricle.
[0273] In another possible implementation, the head 1201 of the target individual can wear a headset, and a spatial posture sensor in the headset can send the detected data to the mobile phone 1202. The mobile phone 1202 can determine the relative position of the mobile phone and the headset (or the head of the target individual) according to the position information detected by the spatial posture sensor in the mobile phone and the position information received from the headset, combine the relative position and the image together for analysis and modeling, improve the modeling accuracy, and further improve the accuracy of the HRTF.
[0274] In the above process, the mobile phone 1202 can output prompt information to prompt the user to perform an operation, or to guide the user to complete the HRTF generation process. The prompting manner of the mobile phone 1202 can refer to the foregoing embodiments.
[0275] It should be understood that the electronic device can also be replaced with another device, and the type of the another device can refer to the foregoing embodiments. When the mobile phone is replaced with another electronic device, a person skilled in the art can make necessary adjustments or improvements, for example, when the electronic device is a large-screen device (for example, a smart television), the electronic device can guide the user to rotate the head to complete the HRTF generation process based on the human-computer interaction manner.
[0276] It should be understood that the headset can be replaced with another device, and the specific implementation can refer to the foregoing embodiments. When the headset is replaced with another electronic device, a person skilled in the art can make necessary adjustments or improvements.
[0277] In a possible implementation, after the HRTF is obtained based on the head three-dimensional model in the foregoing manner, a larger range of HRTF can be obtained based on the HRTF database, or a part of the HRTF (for example, the HRTF that is greatly different from the HRTF in the HRTF database) can be corrected based on the HRTF database or the HRTF based on the characteristics (or the geometric characteristics) of the headset. The specific implementation can refer to the foregoing embodiments.
[0278] In the foregoing embodiments of the present application, the HRTF of the headset can be obtained by using the electronic device to perform a simple test, without the need to perform a test in a laboratory environment to obtain the HRTF, thereby simplifying the HRTF acquisition process and improving the convenience of HRTF acquisition.
[0279] Based on the foregoing embodiments and the same concept, the present application further provides an electronic device for implementing the method performed by the electronic device provided in the embodiments of the present application.
[0280] As Figure 13As shown, the electronic device 1300 can include a memory 1301, one or more processors 1302, and one or more computer programs (not shown in the figure). The above-mentioned devices can be coupled through one or more communication buses 1303. Optionally, when the electronic device 1300 is used to implement the method performed by the electronic device provided in the embodiments of the present application, the electronic device 1300 can further include a display screen 1304.
[0281] In the memory 1301, one or more computer programs (codes) are stored, and the one or more computer programs include computer instructions; the one or more processors 1302 invoke the computer instructions stored in the memory 1301, so that the electronic device 1300 performs the method provided in the embodiments of the present application. The display screen 1304 is used to display images, videos, application interfaces, and other related user interfaces.
[0282] In a specific implementation, the memory 1301 can include a high-speed random access memory, and can also include a non-volatile memory, for example, one or more disk storage devices, flash devices, or other non-volatile solid-state storage devices. The memory 1301 can store an operating system (hereinafter referred to as a system), for example, an ANDROID, IOS, WINDOWS, or LINUX embedded operating system. The memory 1301 can be used to store the implementation program of the embodiments of the present application. The memory 1301 can also store a network communication program, which can be used to communicate with one or more additional devices, one or more user devices, and one or more network devices. The one or more processors 1302 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.
[0283] It should be noted that, Figure 13 The above is only one implementation of the electronic device 1300 provided in the embodiments of the present application, and in actual applications, the electronic device 1300 can further include more or fewer components, which are not limited here.
[0284] Based on the above embodiments and the same concept, the embodiments of the present application further provide a computer-readable storage medium, which stores a computer program, when the computer program runs on a computer, the computer program causes the computer to perform the method performed by the electronic device in the method provided in the above embodiments.
[0285] Based on the above embodiments and the same concept, the embodiments of the present application further provide a computer program product, which comprises computer programs or instructions, and when the computer programs or instructions are run on a computer, the computer programs or instructions make the computer execute the method performed by the electronic device in the method provided by the above embodiments.
[0286] Those skilled in the art will understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer usable program code.
[0287] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0288] These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction apparatus, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0289] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 an apparatus that performs the functions specified in one or more flows and / or blocks.
[0290] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A method for generating Head-Related Transfer Function (HRTF), characterized in that, Applied to electronic devices, the method includes: Output a first prompt message, which is used to instruct a target individual wearing a head-mounted device to perform a first operation, the first operation causing relative motion between the electronic device and the head-mounted device; During the first operation, a test audio signal is played, and feedback audio signals and position information are received from the head-mounted device; wherein, the feedback audio signal is a test audio signal received by the head-mounted device at multiple different spatial locations, the spatial location is the spatial location of the electronic device relative to the head-mounted device, and the position information is used to determine the spatial location; Based on the test audio signal, the feedback audio signal, and the location information, the HRTF corresponding to the target individual is determined.
2. The method as described in claim 1, characterized in that, The first operation is that the target individual holds the electronic device and rotates it around their head while keeping their head still; or The first operation is for the target individual to turn their head while keeping the electronic device in their hand.
3. The method as described in claim 2, characterized in that, The first prompt information is also used to indicate one or more of the following: the distance between the electronic device and the head-mounted device, the direction of rotation, the angle of rotation, the range of rotation, and the duration.
4. The method according to any one of claims 1-3, characterized in that, The output of the first prompt information includes: After establishing a communication connection between the electronic device and the head-mounted device, the first prompt message is output, and the communication connection is used to transmit the feedback audio signal and location information; or After the first application is started, the first prompt message is output, and the first application is used to obtain HRTF; or In response to the first voice command, the first prompt message is output, wherein the first voice command is used to trigger the acquisition of HRTF; or When the head-mounted device is detected to be in a wearing state, the first prompt message is output.
5. The method according to any one of claims 1-4, characterized in that, Before outputting the first prompt message, it also includes: Output a second prompt message, which is used to instruct the target individual to wear the head-mounted device; The output of the first prompt information includes: When the head-mounted device is detected to be in a wearing state, the first prompt message is output.
6. The method according to any one of claims 1-5, characterized in that, Also includes: The test audio signal will stop playing after the set duration. or The audio signal playback ended based on the first user's operation; or In response to the second voice command, the playback of the test audio signal is stopped.
7. The method according to any one of claims 1-6, characterized in that, Also includes: Output a third prompt message, which is used to instruct the target individual to end the first operation and / or to indicate that an HRTF has been generated.
8. The method according to any one of claims 1-7, characterized in that, The first prompt message output includes one or more of the following: Play voice prompts; or Display a text prompt message; or Display images or animations to demonstrate the relative motion between the electronic device and the head-mounted device.
9. The method according to any one of claims 1-8, characterized in that, Also includes: The spatial position of the electronic device relative to the head-mounted device is determined based on the location information of the electronic device and the location information from the head-mounted device. The position information of the head-mounted device is detected by the spatial attitude sensor in the head-mounted device, and the position information of the electronic device is detected by the spatial attitude sensor in the electronic device.
10. The method according to any one of claims 1-9, characterized in that, The feedback audio signal includes a first feedback audio signal collected by the first microphone of the head-mounted device and a second feedback audio signal collected by the second microphone of the head-mounted device; The step of determining the HRTF corresponding to the target individual based on the test audio signal, the feedback audio signal, and the location information includes: The HRTF corresponding to the target individual is determined based on the test audio signal, the first feedback audio signal, the second feedback audio signal, and the location information.
11. The method as described in claim 10, characterized in that, The spatial location includes a first spatial location; The step of determining the HRTF corresponding to the target individual based on the test audio signal, the first feedback audio signal, the second feedback audio signal, and the spatial location includes: Based on the test audio signal played by the electronic device and the first feedback audio signal collected by the first microphone when the electronic device is in a first spatial position relative to the head-mounted device, the first HRTF corresponding to the first spatial position is determined. Based on the test audio signal played by the electronic device and the second feedback audio signal collected by the second microphone when the electronic device is in a first spatial position relative to the head-mounted device, the second HRTF corresponding to the first spatial position is determined. The first HRTF is corrected based on the second HRTF to obtain the HRTF corresponding to the first spatial location.
12. The method according to any one of claims 1-11, characterized in that, After determining the HRTF corresponding to the target individual, the process also includes one or more of the following: The HRTF is corrected based on the characteristic parameters of the head-mounted device; or The HRTF is modified based on the human characteristics of the target individual; or The HRTF is corrected based on the HRTF database.
13. The method as described in claim 12, characterized in that, The step of modifying the HRTF based on the human characteristics of the target individual includes: Acquire an image of the target individual object captured by the electronic device; The image is identified to obtain the auricular structural features and / or head structural features of the target individual; The HRTF is modified based on the auricular and / or head structural features of the target individual.
14. The method as described in claim 12, characterized in that, The spatial position of the electronic device relative to the head-mounted device includes a first spatial position, and the HRTF determined based on the test audio signal, the feedback audio signal, and the position information includes the first HRTF corresponding to the first spatial position; The step of correcting the HRTF based on the HRTF database includes: Based on the target individual's human body feature information and the first spatial location, a second HRTF that matches the human body feature information and the first spatial location is obtained from the HRTF database; The first HRTF is modified according to the second HRTF.
15. The method according to any one of claims 1-14, characterized in that, The HRTF determined based on the test audio signal, the feedback audio signal, and the location information includes HRTFs corresponding to at least two spatial locations within the first spatial range; The method further includes: Based on the human characteristics of the target individual, obtain the HRTF corresponding to at least one spatial location within a second spatial range that matches the human characteristics from the HRTF database; The HRTF corresponding to at least two spatial locations within the first spatial range and the HRTF corresponding to at least one spatial location within the second spatial range are determined as the HRTF corresponding to the target individual.
16. A method for generating a Head-Related Transfer Function (HRTF), characterized in that, Applied to a head-mounted device, the method includes: Collect feedback audio signals, which are test audio signals played by electronic devices and received by the head-mounted device at multiple different spatial locations, where the spatial location is the spatial position of the electronic device relative to the head-mounted device; Receive location information from the electronic device, the location information being used to determine the spatial location; Based on the test audio signal, the feedback audio signal, and the location information, the HRTF corresponding to the target individual is determined.
17. The method as described in claim 16, characterized in that, Also includes: Based on the location information of the head-mounted device and the location information from the electronic device, the spatial position of the electronic device relative to the head-mounted device is determined; The position information of the head-mounted device is detected by the spatial attitude sensor in the head-mounted device, and the position information of the electronic device is detected by the spatial attitude sensor in the electronic device.
18. The method according to any one of claims 16-17, characterized in that, The feedback audio signal includes a first feedback audio signal collected by the first microphone of the head-mounted device and a second feedback audio signal collected by the second microphone of the head-mounted device; The step of determining the HRTF corresponding to the target individual based on the test audio signal, the feedback audio signal, and the location information includes: The HRTF corresponding to the target individual is determined based on the test audio signal, the first feedback audio signal, the second feedback audio signal, and the location information.
19. The method as described in claim 18, characterized in that, The spatial location includes a first spatial location; The step of determining the HRTF corresponding to the target individual based on the test audio signal, the first feedback audio signal, the second feedback audio signal, and the spatial location includes: Based on the test audio signal played by the electronic device and the first feedback audio signal collected by the first microphone when the electronic device is in a first spatial position relative to the head-mounted device, the first HRTF corresponding to the first spatial position is determined. Based on the test audio signal played by the electronic device and the second feedback audio signal collected by the second microphone when the electronic device is in a first spatial position relative to the head-mounted device, the second HRTF corresponding to the first spatial position is determined. The first HRTF is corrected based on the second HRTF to obtain the HRTF corresponding to the first spatial location.
20. The method according to any one of claims 16-19, characterized in that, After determining the HRTF corresponding to the target individual, the process also includes one or more of the following: The HRTF is corrected based on the characteristic parameters of the head-mounted device; or The HRTF is modified based on the human characteristics of the target individual; or The HRTF is corrected based on the HRTF database.
21. The method as described in claim 20, characterized in that, The spatial position of the electronic device relative to the head-mounted device includes a first spatial position, and the HRTF determined based on the test audio signal, the feedback audio signal, and the position information includes the first HRTF corresponding to the first spatial position; The step of correcting the HRTF based on the HRTF database includes: Based on the target individual's human body feature information and the first spatial location, a second HRTF that matches the human body feature information and the first spatial location is obtained from the HRTF database; The first HRTF is modified according to the second HRTF.
22. The method according to any one of claims 16-21, characterized in that, The HRTF determined based on the test audio signal, the feedback audio signal, and the location information includes HRTFs corresponding to at least two spatial locations within the first spatial range; The method further includes: Based on the human characteristics of the target individual, obtain the HRTF corresponding to at least one spatial location within a second spatial range that matches the human characteristics from the HRTF database; The HRTF corresponding to at least two spatial locations within the first spatial range and the HRTF corresponding to at least one spatial location within the second spatial range are determined as the HRTF corresponding to the target individual.
23. An apparatus, characterized in that, It includes units or modules for performing the method as described in any one of claims 1-15, or units or modules for performing the method as described in any one of claims 16-22.
24. An apparatus, characterized in that, include: One or more processors are configured to perform the method as described in any one of claims 1-15, or to perform the method as described in any one of claims 16-22.
25. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed on the device, cause the device to perform the method as described in any one of claims 1-15, or the method as described in any one of claims 16-22.
26. A chip system, characterized in that, Includes a processor for supporting a computer device in implementing the method as claimed in any one of claims 1-15, or in implementing the method as claimed in any one of claims 16-22.
27. A program product, characterized in that, The program product includes a program; when the program is run on a computer, it causes the computer to perform the method as described in any one of claims 1-15, or to perform the method as described in any one of claims 16-22.