Audio device, control method thereof, and storage medium
By monitoring the head rotation angle to adjust the virtual speaker position in the audio device and using the HRTF in the personalized acoustic transmission data, the problem of poor sound quality of the audio device when the head rotates is solved, and accurate sound source perception and sound quality improvement are achieved when the head rotates.
Patent Information
- Application Number
- CN202511100959.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-07
AI Technical Summary
When the head moves, the sound image of existing audio equipment becomes blurred, resulting in a poor sense of sound source positioning and poor sound quality.
By monitoring the target user's head rotation angle, adjusting the virtual speaker position in the audio device, and using the head-related transfer function HRTF in the personalized acoustic transmission data for audio rendering, the audio playback at the virtual playback position is simulated.
Even when the head turns, the target user can accurately perceive the direction of the sound source, improving the sound quality performance of the audio device.
Smart Images

Figure CN120602885A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of audio technology, and in particular to an audio device, a control method thereof, and a storage medium. Background Art
[0002] In recent years, binaural audio systems using head tracking technology have become increasingly popular. Head tracking technology allows listeners to perceive sound coming from multiple directions, creating an immersive listening experience. However, in practice, it has been found that when the head turns, the sound image becomes blurred, resulting in a poor sense of sound source localization, which in turn leads to poor sound quality in audio devices. Therefore, the problem of poor sound quality in audio devices caused by head rotation currently exists.
[0003] The above content is only used to assist in understanding the technical solutions of the embodiments of the present application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to provide an audio device and a control method and a storage medium thereof, aiming to solve the problem of poor sound quality performance of the audio device caused by head rotation.
[0005] To achieve the above objectives, an embodiment of the present application provides a method for controlling an audio device, the method comprising: Monitoring the head rotation angle of the target user and obtaining personalized acoustic transmission data of the target user, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) trained based on the personal auditory physiological data of the target user; Adjusting the position of a virtual speaker in the audio device according to the head rotation angle to obtain a virtual playback position of the virtual speaker; A target head-related transfer function HRTF corresponding to the virtual playback posture is obtained from the personalized acoustic transmission data, and the audio played by the audio device is rendered according to the target head-related transfer function HRTF to simulate playing the audio at the virtual playback posture.
[0006] In one embodiment, the step of obtaining the personalized acoustic transmission data of the target user includes: The preset model is trained by using the acquired group auditory training data set to obtain a trained general model, wherein the group auditory training data set includes a plurality of human auditory physiological data and their corresponding HRTF labels; Acquiring personal auditory physiological data of the target user; A universal decoder is obtained from the universal model, and personalized training is performed on the universal decoder using the personal auditory physiological data to obtain a personalized decoder. Personalized acoustic transmission data corresponding to the personal auditory physiological data is generated based on the personalized decoder.
[0007] In one embodiment, the preset model includes a preset encoder and a preset decoder; and the step of training the preset model using the acquired group auditory training dataset to obtain a trained universal model includes: Acquiring target human auditory physiological data from the group auditory training data set, inputting the target human auditory physiological data into the preset model, encoding the target human auditory physiological data through a preset encoder in the preset model, inputting the encoded result into the preset decoder, and outputting the HRTF training result through the preset decoder; Determine the training loss between the HRTF training result and the target HRTF label of the target human auditory physiological data. If the training loss is greater than a preset loss threshold, re-acquire new target human auditory physiological data from the group auditory training dataset, and return to the step of inputting the target human auditory physiological data into the preset model until the training loss is less than or equal to the preset loss threshold, thereby obtaining a trained general model.
[0008] In one embodiment, the step of obtaining the target user's personal auditory physiological data includes: Obtaining head structural parameters of the target user, including auricle height, auricle width, ear canal inclination, head width, and shoulder and neck parameters; A head geometric model of the target user is constructed using the head structural parameters, and personal auditory physiological data is extracted from the head geometric model.
[0009] In one embodiment, the step of obtaining the target user's personal auditory physiological data includes: Obtaining three-dimensional head scan data of the user, and determining personal auditory physiological data from the three-dimensional head scan data; and / or, Obtaining head measurement parameters of the user, obtaining target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters from a preset auditory physiological database, and obtaining personal auditory physiological data consisting of the plurality of target sub-auditory physiological data, wherein the head measurement parameters include a plurality of sub-measurement parameters, namely, an outer ear structure parameter, a head-neck connection parameter, and a head body parameter, and the preset auditory physiological database includes the preset sub-auditory physiological data corresponding to the plurality of preset sub-measurement parameters.
[0010] In one embodiment, the head rotation angle includes a horizontal rotation angle, the virtual speaker includes a first speaker and a second speaker, and the virtual playback posture includes a first playback posture and a second playback posture; The step of adjusting the posture of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playback posture of the virtual speaker includes: Determining a target horizontal angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal angle; According to the target horizontal angle, the angles of the first speaker and the second speaker in the horizontal direction are adjusted respectively to obtain a first playback position of the first speaker and a second playback position of the second speaker, so that the angle between the first speaker and the second speaker in the horizontal direction is the target horizontal angle.
[0011] In one embodiment, the head rotation angle includes a horizontal rotation angle and a vertical pitch angle, and the virtual playback posture includes a first playback posture and a second playback posture; The step of adjusting the posture of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playback posture of the virtual speaker includes: Determining a target vertical angle corresponding to the vertical pitch angle in a mapping relationship between a preset pitch angle and a preset vertical angle; Determining a target horizontal angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal angle; Adjusting the vertical angles of a first speaker and a second speaker in the virtual speakers according to the target vertical angle, and adjusting the horizontal angles of the first speaker and the second speaker according to the target horizontal angle to obtain a first playback position of the first speaker and a second playback position of the second speaker; The angle between the first playback posture and the second playback posture in the horizontal direction is the target horizontal angle, and the angle between the first playback posture and the second playback posture in the vertical direction is the target vertical angle.
[0012] In one embodiment, the method for controlling the audio device further includes: receiving a posture adjustment instruction for the virtual speaker, and determining a position adjustment parameter and an angle adjustment parameter of the virtual speaker from the posture adjustment instruction; Adjusting the position of the virtual speaker according to the position adjustment parameter and the angle adjustment parameter to obtain an adjusted position of the virtual speaker in the audio device; Based on the adjusted posture, the audio played by the audio device is rendered.
[0013] In one embodiment, the step of rendering the audio played by the audio device based on the adjusted posture includes: Rendering the audio played by the audio device according to the adjusted position and personalized acoustic transmission data of the virtual speaker; or Obtain the initial posture of the virtual speaker before receiving the posture adjustment indication, and the initial rendering parameters of the audio device for the audio; determine the posture deviation between the initial posture and the adjusted posture, and adjust the initial rendering parameters to obtain target rendering parameters based on the sound source compensation coefficient corresponding to the posture deviation; render the audio played by the audio device using the target rendering parameters.
[0014] In addition, to achieve the above-mentioned purpose, this embodiment further provides a control device for an audio device, the device comprising: A monitoring module, configured to monitor the target user's head rotation angle and obtain personalized acoustic transmission data of the target user, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) trained based on the target user's personal auditory physiological data; An adjustment module, configured to adjust the position of the virtual speaker in the audio device according to the head rotation angle to obtain a virtual playback position of the virtual speaker; The rendering module is used to obtain the target head-related transfer function HRTF corresponding to the virtual playback posture in the personalized acoustic transmission data, and render the audio played by the audio device based on the target head-related transfer function HRTF to simulate the audio playing at the virtual playback posture.
[0015] In addition, to achieve the above-mentioned purpose, an embodiment of the present application also provides an audio device, which includes: a memory, a processor, and a program of the control method of the audio device stored in the memory and runnable on the processor. When the program of the control method of the audio device is executed by the processor, the steps of the control method of the audio device as described above can be implemented.
[0016] In addition, to achieve the above-mentioned purpose, an embodiment of the present application also provides a computer-readable storage medium, on which a program for implementing a method for controlling an audio device is stored. When the program for implementing a method for controlling an audio device is executed by a processor, the steps of the method for controlling an audio device as described above are implemented.
[0017] In addition, to achieve the above-mentioned purpose, an embodiment of the present application further provides a computer program product, including a computer program, which implements the steps of the above-mentioned audio device control method when executed by a processor.
[0018] One or more technical solutions proposed in the embodiments of the present application have at least the following technical effects: the present application can monitor the head rotation angle of the target user, and adjust the position of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playback position of the virtual speaker. Since the virtual speaker is used to simulate the spatial position of the real sound source, the angle of the virtual speaker will also affect the brain's judgment of the direction of the sound source. Therefore, the present application adjusts the position of the virtual speaker by the head rotation angle to obtain the virtual playback position of the virtual speaker, thereby compensating for the head rotation angle of the target user by adjusting the position of the virtual speaker, so that the target user can still accurately perceive the direction of the sound source.
[0019] Specifically, the target head-related transfer function HRTF (Head-Related Transfer Function) corresponding to the virtual playback posture is obtained in the personalized acoustic transmission data, and the audio played by the audio device is rendered according to the target head-related transfer function HRTF to simulate playing the audio at the virtual playback posture. Since the personalized acoustic transmission data includes at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user, each head-related transfer function HRTF in the personalized acoustic transmission data can reflect the personal auditory physiological characteristics of the target user. Therefore, the present application obtains the target head-related transfer function HRTF through the virtual playback posture, and renders the audio played by the audio device through the target head-related transfer function HRTF, so that the rendered audio is more in line with the auditory physiological characteristics of the target user and can better match the head rotation angle of the target user, so that even if the target user's head rotates, the target user can accurately perceive the direction of the sound source, thereby improving the sound quality performance of the audio device. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the embodiments of the present application, and together with the specification are used to explain the principles of the embodiments of the present application.
[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 This is a flow chart of an embodiment of a method for controlling an audio device according to an embodiment of the present application; Figure 2This is a flow chart of obtaining personal auditory physiological data in the method for controlling an audio device according to an embodiment of the present application; Figure 3 A schematic diagram of a scenario for generating personalized acoustic transmission data in a method for controlling an audio device according to an embodiment of the present application; Figure 4 A schematic diagram of a curve showing how the included angle corresponding to the virtual speaker changes with the head rotation angle in the control method of the audio device according to an embodiment of the present application; Figure 5 This is a schematic diagram of the angles corresponding to the virtual speakers in the method for controlling the audio device according to an embodiment of the present application; Figure 6 This is a module diagram of a control device for an audio device according to an embodiment of the present application; Figure 7 Schematic diagram of the device structure of the hardware operating environment involved in the audio device control method in the embodiment of the present application.
[0023] The purpose, features and advantages of the embodiments of the present application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0024] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the embodiments of the present application and are not intended to limit the embodiments of the present application.
[0025] In order to better understand the technical solutions of the embodiments of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0026] Binaural audio systems using head tracking are becoming increasingly popular. Head tracking allows listeners to perceive sounds as coming from multiple directions, creating an immersive and realistic listening experience and providing a sense of space and direction. With the rise of virtual and augmented reality (VR and AR), binaural audio systems have become even more important. Binaural audio uses head-related transfer functions (HRTFs), which are crucial for localizing sounds and spatial perception. In VR and AR, binaural audio systems can create realistic soundscapes that match the user's visual environment, providing a more immersive and engaging experience than using visual cues alone. The demand for binaural 3D spatialization is also growing in fields ranging from gaming and entertainment to medicine and education and training.
[0027] The binaural synthesis process uses head-related transfer functions (HRTFs) as input to assign directions of arrival to sounds originating from virtual sound sources within the created virtual environment. This is typically achieved using a universal HRTF. While universal HRTFs are widely used in spatial audio to create more accurate VR / AR experiences, the auditory physiology of different users varies greatly. Using universal HRTFs for audio rendering can make it difficult to provide a good sound quality experience for every user.
[0028] During the application process, it was found that when the head turns, the sound image becomes blurred, resulting in a poor sense of sound source positioning, which in turn leads to poor sound quality of the audio device. Therefore, there is currently a problem of poor sound quality of audio devices caused by head rotation.
[0029] Therefore, this embodiment provides a control method for an audio device, which can provide each user with a good sound quality experience in a targeted manner, and can provide the user with a good sound quality experience even if the user's head turns. In this embodiment, the head rotation angle of the target user can be monitored, and according to the head rotation angle, the position of the virtual speaker in the audio device is adjusted to obtain the virtual playback position of the virtual speaker. Since the virtual speaker is used to simulate the spatial position of the real sound source, the angle of the virtual speaker will also affect the brain's judgment of the direction of the sound source. Therefore, this application adjusts the position of the virtual speaker by the head rotation angle to obtain the virtual playback position of the virtual speaker, thereby compensating for the target user's head rotation angle by adjusting the position of the virtual speaker, so that the target user can still accurately perceive the direction of the sound source.
[0030] Specifically, the target head-related transfer function HRTF corresponding to the virtual playback posture is obtained in the personalized acoustic transmission data, and the audio played by the audio device is rendered based on the target head-related transfer function HRTF to simulate the audio playing at the virtual playback posture. Since the personalized acoustic transmission data includes at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user, each head-related transfer function HRTF in the personalized acoustic transmission data can reflect the personal auditory physiological characteristics of the target user. Therefore, the present application obtains the target head-related transfer function HRTF through the virtual playback posture, and renders the audio played by the audio device through the target head-related transfer function HRTF, so that the rendered audio is more in line with the auditory physiological characteristics of the target user, and can better match the head rotation angle of the target user, so that even when the target user's head rotates, the target user can accurately perceive the direction of the sound source, thereby improving the sound quality performance of the audio device.
[0031] Based on this, the embodiment of the present application provides a method for controlling an audio device, referring to Figure 1 , Figure 1 This is a flow chart of a first embodiment of a method for controlling an audio device according to an embodiment of the present application. The method for controlling an audio device includes steps S10 to S30: Step S10: monitoring the target user's head rotation angle and obtaining the target user's personalized acoustic transmission data, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) trained based on the target user's personal auditory physiological data; It should be noted that audio devices can be either near-ear or far-ear audio devices. Near-ear open-type devices include, but are not limited to, AR devices, VR devices, smart audio glasses, neck-mounted speakers, open-type headphones, mobile phones, tablets, etc., while far-ear open-type devices include, but are not limited to, audio devices such as speakers and televisions. The head rotation angle is the rotation angle of the target user relative to a preset zero-position direction. The preset zero-position direction can be the direction in which the target user's head is facing straight ahead. In other embodiments, the preset zero-position direction can also be determined based on actual conditions. For example, the preset zero-position direction corresponding to the audio device can be determined based on different audio devices. This embodiment does not specifically limit this.
[0032] The personalized acoustic transmission data may include at least one head-related transfer function HRTF. The head-related transfer function HRTF is a frequency domain filter function that describes the reflection, diffraction, and scattering effects of sound waves on physiological structures such as the head, auricle, and shoulders during the process of sound propagating from the free field to the two ears. The head-related transfer function HRTF can be used to simulate the changes in sound waves propagating from the sound source position to the human ear, so that the virtual sound source can have real spatial properties. For the same target user, there may be multiple head-related transfer functions HRTF. Different head-related transfer functions HRTF simulate different sound source positions. The commonality of different head-related transfer function HRTFs for the same target user is that they all conform to the personal auditory physiological characteristics of the target user, but the simulated sound source positions may be different.
[0033] Exemplarily, the head rotation angle of the target user is monitored, the personal auditory physiological data of the target user and the personalized decoder are obtained, and personalized acoustic transmission data of the personal auditory physiological data is generated by the personalized decoder.
[0034] In a feasible embodiment, step S10 further includes steps S11 to S13: Step S11: training a preset model using the acquired group auditory training dataset to obtain a trained general model, wherein the group auditory training dataset includes a plurality of human auditory physiological data and their corresponding HRTF labels; It should be noted that the group auditory training dataset includes multiple human auditory physiological data. Specifically, the group auditory training dataset is obtained by training a deep learning model using a large amount of existing HRTF data (i.e., open-source HRTF data). This existing HRTF data can include HRTF data from a large number of different individuals, covering a wide range of ages, ear shapes, and head shapes. In other words, the multiple human auditory physiological data originate from different individuals. The preset model can include a preset encoder and a preset decoder, and each piece of human auditory physiological data corresponds to an HRTF label. The group auditory training dataset can be predetermined, and each piece of human auditory physiological data can reflect the auditory physiological characteristics of the human body. For example, human auditory physiological data can include data such as the ear canal resonance frequency, the angle between the tragus and the cymba concha, and the complexity of the auricular folds, though this embodiment does not specifically limit this. The human auditory physiological data can be acquired in the same manner as the target user's personal auditory physiological data. The human auditory physiological data can be processed by the preset encoder. The universal model includes a universal decoder and a universal encoder. After the preset model is trained, a universal model consisting of the universal decoder and universal encoder can be obtained.
[0035] Exemplarily, a group auditory training data set is obtained, and a preset model is trained using the group auditory training data set to obtain a general model consisting of a general decoder and a general encoder. The general model can also be used to generate a general HRTF, and the general HRTF can also simulate the changes in sound waves transmitted from the sound source to the human ear, but the general HRTF is difficult to adapt to the individual differences in user hearing. Therefore, when rendering audio through a general HRTF, there will still be a problem of poor sound quality performance of the audio device. Therefore, after determining the general model, this embodiment will fine-tune the general decoder in the general model to obtain a personalized decoder, so that the personalized acoustic transmission data of the target user can be generated through the personalized decoder later.
[0036] In a feasible embodiment, step S11 further includes steps S111 to S112: Step S111, obtaining target human auditory physiological data from a group auditory training dataset, inputting the target human auditory physiological data into a preset model, encoding the target human auditory physiological data through a preset encoder in the preset model, inputting the encoded result into a preset decoder, and outputting the HRTF training result through the preset decoder; Step S112, determine the training loss between the HRTF training result and the target HRTF label of the target human auditory physiological data. When the training loss is greater than the preset loss threshold, obtain new target human auditory physiological data from the group auditory training data set, and return to the step of inputting the target human auditory physiological data into the preset model until the training loss is less than or equal to the preset loss threshold, thereby obtaining a trained general model.
[0037] It should be noted that the preset encoder can encode the target human auditory physiological data, and can encode the high-dimensional target human auditory physiological data into a low-dimensional feature vector related to hearing to downsample and encode the target human auditory physiological data. Specifically, the preset encoder can encode the features that affect HRTF from the target human auditory physiological data, thereby reducing the complexity of the model and improving the accuracy of training. The result of the encoding of the target human auditory physiological data by the preset encoder can be input into the preset decoder, and the preset decoder decodes the encoded result to generate an HRTF training result. The HRTF training result is the HRTF corresponding to the target human auditory physiological data output by the preset decoder in the preset model.
[0038] The training loss can represent the difference between the target HRTF label and the HRTF training result. The greater the training loss, the greater the difference between the target HRTF label and the HRTF training result. The preset loss threshold can be determined based on actual conditions, and this embodiment does not specifically limit this. When the training loss is greater than the preset loss threshold, it means that the difference between the target HRTF label and the HRTF training result is large, so it is necessary to obtain new target human auditory physiological data for training until the training loss is less than or equal to the preset loss threshold, thereby obtaining a trained general model.
[0039] Exemplarily, the target human auditory physiological data is input into a preset model, the target human auditory physiological data is encoded by a preset encoder in the preset model, the encoded result is input into a preset decoder, and the HRTF training result is output by the preset decoder; the training loss between the HRTF training result and the target HRTF label of the target human auditory physiological data is calculated, and when the training loss is greater than a preset loss threshold, new target human auditory physiological data is obtained from the group auditory training data set, and the step of inputting the target human auditory physiological data into the preset model is returned until the training loss is less than or equal to the preset loss threshold, thereby obtaining a trained general model.
[0040] This embodiment trains a preset model through a group auditory training data set to obtain a universal model, which facilitates the subsequent use of a universal decoder in the universal model to determine personalized acoustic transmission data, so as to provide each user with a good sound quality experience.
[0041] Step S12, obtaining the target user's personal auditory physiological data; It should be noted that personal auditory physiological data can reflect the target user's personal auditory physiological characteristics. Personal auditory physiological data can be obtained through three-dimensional head scanning, from head measurement parameters of the target user, or through geometric modeling of the user's head. This embodiment does not specifically limit this.
[0042] In a feasible embodiment, step S12 further includes steps S121 and S122: Step S121, obtaining the target user's head structural parameters, which include auricle height, auricle width, ear canal inclination, head width, and shoulder and neck parameters; Step S122: constructing a head geometric model of the target user using the head structure parameters, and extracting personal auditory physiological data from the head geometric model.
[0043] It's important to note that head width is the lateral distance between the left and right tragus, which determines the interaural time difference (ITD). Neck and shoulder parameters, including shoulder width and neck height, affect the reflection and scattering of low-frequency sound waves. Auricle height is the vertical distance from the tragus to the highest point of the helix. Auricle width can be defined as the maximum horizontal distance from the front edge of the helix to the back edge of the antihelix. The ear canal inclination is the angle between the ear canal axis and the sagittal plane (the plane of bilateral symmetry) of the head, which affects the direction of sound wave incidence.
[0044] A head geometric model of the target user can be constructed using the head structural parameters. 3D modeling software can be called to construct a head geometric model based on the head structural parameters. Features directly related to auditory function can be directly extracted from the head geometric model to obtain personal auditory physiological data. For example, personal auditory physiological data may include the ear canal resonance frequency, the angle between the tragus and the cymba concha, and the ITD baseline value. This embodiment does not specifically limit this, and the specific value can be determined based on actual conditions. In other embodiments, the head structural parameters may also include parameters such as mandibular width and ear canal length, which are not specifically limited in this embodiment.
[0045] For example, the head structure parameters of the target user are obtained. The head structure parameters can be obtained by taking photos. Modeling software is called to build a head geometry model of the target user based on the head structure parameters, and personal auditory physiological data is extracted from the head geometry model.
[0046] This embodiment constructs a head geometry model of the target user through head structure parameters, thereby facilitating the extraction of personal auditory physiological data from the head geometry model, and further facilitating the subsequent determination of personalized acoustic transmission data, so as to subsequently improve the sound quality performance of the audio device.
[0047] In a feasible embodiment, step S12 further includes step S123 and / or step S124: Step S123, obtaining the user's three-dimensional head scan data, and determining the individual's auditory physiological data from the three-dimensional head scan data; Step S124: Obtain the user's head measurement parameters, and obtain target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters from a preset auditory physiological database to obtain personal auditory physiological data consisting of multiple target sub-auditory physiological data. The head measurement parameters include multiple sub-measurement parameters, namely, external ear structure parameters, head-neck connection parameters, and head body parameters. The preset auditory physiological database includes preset sub-auditory physiological data corresponding to the multiple preset sub-measurement parameters.
[0048] It should be noted that the three-dimensional head scan data can be obtained by performing 3D structured light or laser scanning on the target user, and personal auditory physiological data related to hearing can be extracted from the three-dimensional head scan data.
[0049] Cephalometry parameters can be obtained by directly measuring the target user's head and the neck and shoulders connected to the head. For example, they can be obtained through multi-view photogrammetry. Cephalometry parameters can include multiple sub-parameters, which can be external ear structure parameters, head-neck connection parameters, and head body parameters. Each sub-parameter can also include multiple sub-parameters. The external ear structure parameters can include multiple sub-parameters such as the auricle contour, auricle folds, auricle size, ear canal length and diameter, and concha depth. The head-neck connection parameters can include multiple sub-parameters such as shoulder width and neck length, and the head body parameters can include multiple sub-parameters such as head diameter and interaural distance.
[0050] The preset auditory physiological database may be predetermined based on actual conditions and may include preset sub-auditory physiological data corresponding to a plurality of preset sub-measurement parameters. Target sub-auditory physiological data for each sub-measurement parameter may be searched in the preset auditory physiological database.
[0051] For example, 3D structured light and / or laser scanning may be performed on the target user to obtain three-dimensional head scanning data, and personal auditory physiological data related to hearing may be acquired from the three-dimensional head scanning data.
[0052] The user's head measurement parameters can also be obtained. Target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters can be obtained from a preset auditory physiological database, thereby obtaining personal auditory physiological data consisting of multiple target sub-auditory physiological data. Specifically, each sub-measurement parameter includes multiple sub-parameters. The preset sub-auditory physiological data corresponding to each sub-parameter can be searched in the preset auditory physiological database to obtain personal auditory physiological data consisting of multiple preset sub-auditory physiological data. The preset sub-measurement parameters also include multiple preset sub-parameters, and the preset sub-measurement parameters can correspond to multiple preset sub-auditory physiological data. This is not specifically limited in this embodiment.
[0053] This embodiment can obtain personal auditory physiological data in a variety of ways, thereby improving the flexibility of data acquisition. For example, you can also refer to Figure 2 , Figure 2 Various methods of obtaining personal auditory physiological data are given. Personal auditory physiological data can be extracted from a head geometric model, from three-dimensional head scanning data, or by searching a preset auditory physiological database for target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters to obtain personal auditory physiological data composed of multiple target sub-auditory physiological data.
[0054] Step S13: obtaining a universal decoder from the universal model, performing personalized training on the universal decoder using the personal auditory physiological data to obtain a personalized decoder, and generating personalized acoustic transmission data corresponding to the personal auditory physiological data based on the personalized decoder.
[0055] It should be noted that personal auditory physiological data can be used to perform unsupervised training on the universal decoder to perform personalized training on the universal decoder. Since the universal decoder has learned the general mapping rules between the physiological characteristics of the group and HRTF, the personal auditory physiological data can be used to perform unsupervised learning on the universal decoder, and then the universal decoder can be adjusted in a targeted manner. For example, the network parameters of the universal decoder can be adjusted to obtain a personalized decoder, where the network parameters of the universal decoder refer to the weights and biases of the neurons in the universal decoder.
[0056] Exemplarily, a universal decoder is obtained from a universal model, and the universal decoder is unsupervisedly trained using personal auditory physiological data to obtain a personalized decoder. The personal auditory physiological data can be input into a universal encoder or a preset encoder to encode the personal auditory physiological data using the universal encoder or the preset encoder. The encoded personal auditory physiological data is input into a personalized decoder, and the personalized decoder outputs personalized acoustic transmission data of the personal auditory physiological data. The personalized acoustic transmission data can include left-ear acoustic transmission data and right-ear acoustic transmission data. The left-ear acoustic transmission data and the right-ear acoustic transmission data can each include their corresponding head-related transfer functions (HRTFs). The left-ear acoustic transmission data and the right-ear acoustic transmission data can be symmetrical, or can be trained asymmetrically to adapt to the personalized characteristics of the left and right ears of the target user.
[0057] Specifically, in other embodiments, the left-ear acoustic transmission data and the right-ear acoustic transmission data may also be asymmetric. For example, the left-ear personal physiological data and the right-ear personal physiological data may be obtained from the personal auditory physiological data. The left-ear personal physiological data may be input into a universal encoder or a preset encoder to encode the left-ear personal physiological data. The encoded left-ear personal physiological data may be input into a personalized decoder, which outputs the left-ear acoustic transmission data of the left-ear personal physiological data. The right-ear personal physiological data may be input into a universal encoder or a preset encoder to encode the right-ear personal physiological data. The encoded right-ear personal physiological data may be input into a personalized decoder, which outputs the right-ear acoustic transmission data of the right-ear personal physiological data. This facilitates improving the accuracy of the personalized acoustic transmission data and can also accommodate differences between the left and right ears of the same user.
[0058] This embodiment first trains a universal decoder, thereby facilitating targeted adjustment of the network parameters of the universal decoder based on personal auditory physiological data to obtain a personalized decoder, and then facilitates generation of personalized acoustic transmission data in combination with the personalized decoder to improve the accuracy of audio rendering, so that the rendered audio meets the auditory physiological characteristics of the target user and provides the user with a good sound quality experience.
[0059] To better understand this embodiment, please refer to Figure 3 , Figure 3A schematic diagram illustrates the personalization phase and universal model training phase of this embodiment. During the universal model training phase, a group auditory training dataset can be input into a preset encoder. The output of the preset encoder is then input into a preset decoder, which then outputs HRTF training results. The preset encoder and decoder can be trained until a trained universal decoder is obtained. The universal decoder can be personalized using individual auditory physiological data to obtain a personalized decoder.
[0060] In the personalization stage, acoustic measurements can be performed to obtain personal auditory physiological data. The personal auditory physiological data can be used to perform personalized training on the universal decoder to obtain a personalized decoder. The personal auditory physiological data is input into the personalized decoder, and the personalized decoder can output personalized acoustic transmission data. Figure 3 The process of personalized training of the general decoder through personal auditory physiological data is not shown in FIG.
[0061] Step S20, adjusting the position of the virtual speaker in the audio device according to the head rotation angle to obtain a virtual playback position of the virtual speaker; It should be noted that when the target user rotates their head, the position of their ear may change. When the ear position changes, the sound waves transmitted from the original sound source to the ear also change, which may affect the target user's judgment of the sound source location and lead to poor localization. Therefore, this embodiment adjusts the position of the virtual speaker based on the head rotation angle. This adjustment can compensate for the head rotation angle by adjusting the position of the virtual speaker, so that the user can still experience good sound quality.
[0062] The virtual speakers may include a first speaker and a second speaker. The first speaker may be a speaker near the left ear, and the second speaker may be a speaker near the right ear. The virtual playback position is the position and angle at which the virtual speakers ultimately play audio. Both the first speaker and the second speaker are virtual speakers.
[0063] Exemplarily, in this embodiment, the angle of the virtual speaker in the audio device can be adjusted according to the head rotation angle to obtain the virtual playback posture of the virtual speaker, so as to compensate for the head rotation angle through the virtual playback posture. Specifically, the angles of the first speaker and the second speaker in the audio device can be adjusted to obtain the first playback posture of the first speaker and the second playback posture of the second speaker, wherein the virtual playback posture may include the first playback posture and the second playback posture.
[0064] Step S30: obtaining a target head-related transfer function HRTF corresponding to the virtual playback posture in the personalized acoustic transmission data, and rendering the audio played by the audio device according to the target head-related transfer function HRTF to simulate playing the audio at the virtual playback posture.
[0065] It should be noted that the virtual playback posture may include a first playback posture of a first speaker and a second playback posture of a second speaker. The first speaker may be a left ear speaker, and the second speaker may be a right ear speaker. This is not specifically limited in this embodiment. The first playback posture includes the position and angle of the left ear speaker, and the second playback posture includes the position and angle of the right ear speaker. The target head-related transfer function HRTF may include a first target head-related transfer function HRTF for the left ear and a second target head-related transfer function HRTF for the right ear.
[0066] Exemplarily, the audio played by the audio device may be rendered using a first target head-related transfer function HRTF and a second target head-related transfer function HRTF to simulate playing the audio at a first playback posture and a second playback posture.
[0067] This embodiment can monitor the head rotation angle of the target user and adjust the position of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playback position of the virtual speaker. Since the virtual speaker is used to simulate the spatial position of the real sound source, the angle of the virtual speaker will also affect the brain's judgment of the direction of the sound source. Therefore, this embodiment adjusts the position of the virtual speaker by the head rotation angle to obtain the virtual playback position of the virtual speaker, thereby compensating for the head rotation angle of the target user by adjusting the position of the virtual speaker, so that the target user can still accurately perceive the direction of the sound source.
[0068] Specifically, the target head-related transfer function HRTF corresponding to the virtual playback posture is obtained in the personalized acoustic transmission data, and the audio played by the audio device is rendered based on the target head-related transfer function HRTF to simulate the audio playing at the virtual playback posture. Since the personalized acoustic transmission data includes at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user, each head-related transfer function HRTF in the personalized acoustic transmission data can reflect the personal auditory physiological characteristics of the target user. Therefore, this embodiment obtains the target head-related transfer function HRTF through the virtual playback posture, and renders the audio played by the audio device through the target head-related transfer function HRTF, so that the rendered audio is more in line with the auditory physiological characteristics of the target user, and can better match the head rotation angle of the target user, so that even when the target user's head rotates, the target user can accurately perceive the direction of the sound source, thereby improving the sound quality performance of the audio device.
[0069] In a feasible embodiment, the head rotation angle includes a horizontal rotation angle, the virtual speaker includes a first speaker and a second speaker, and the virtual playback posture includes a first playback posture and a second playback posture; step S20 further includes steps S21 and S22: Step S21, determining a target horizontal angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal angle; Step S22: Adjust the horizontal angles of the first speaker and the second speaker respectively according to the target horizontal angle to obtain a first playback position of the first speaker and a second playback position of the second speaker, so that the horizontal angle between the first speaker and the second speaker is the target horizontal angle.
[0070] It should be noted that the mapping relationship between the preset horizontal angle and the preset horizontal angle can be predetermined, and this embodiment does not specifically limit this. The absolute value of the preset horizontal angle is negatively correlated with the preset horizontal angle. For example, you can refer to Figure 4 , Figure 4 The mapping relationship between the preset horizontal transmission angle and the preset horizontal angle is shown in Figure 4 This is one example of the change in the angle, and the specific value of the angle is not limited in this embodiment. Figure 4 The horizontal axis is the head rotation angle, and the vertical axis is the angle corresponding to the virtual speaker. Figure 4 The given data point in represents the extreme value of the angle change. Figure 4 The maximum value of the included angle is 120 degrees, and the minimum value is 60 degrees. When the head rotates at 0 degrees, the included angle reaches its maximum value of 120 degrees, and when the head rotates at 90 degrees, the included angle reaches its minimum value of 60 degrees. When the head rotates from 0 to 90 degrees, the included angle changes from 120 to 60 degrees. When the head rotates from 90 to 180 degrees, the included angle changes from 60 to 120 degrees, and so on. The included angle corresponding to the virtual speaker is the angle between the first and second speakers.
[0071] For two-channel audio, when the head does not rotate (the head rotation angle is 0 degrees), the angle between the first speaker and the second speaker is 120 degrees, the angle of the first speaker can be negative 60 degrees, and the angle of the second speaker can be positive 60 degrees. When the head rotates, when rotating to the extreme left (negative 90 degrees) or extreme right (positive 90 degrees), the angle between the first speaker and the second speaker can be 60 degrees, the angle of the first speaker can be negative 30 degrees, and the angle of the second speaker can be positive 30 degrees. When the head rotation angle changes from 0 degrees to 90 degrees, or from 0 degrees to -90 degrees, the horizontal angle of the first speaker can change from negative 60 degrees to negative 30 degrees, and the horizontal angle of the second speaker can change from positive 60 degrees to positive 30 degrees.
[0072] In this embodiment, between -90 degrees and 90 degrees, -90 degrees can refer to the position where the head is turned to the extreme left, and 90 degrees can refer to the position where the head is turned to the extreme right. The higher the absolute value of the head rotation angle, the smaller the angle between the first speaker and the second speaker. The lower the absolute value of the head rotation angle, the larger the angle between the first speaker and the second speaker. The angles of the first speaker and the second speaker are symmetrical. When the absolute value of the head rotation angle is higher, it means that the amplitude of the head rotation is large. If the angle of the virtual speaker is not adjusted, the sound field constructed by the first speaker and the second speaker in the virtual speaker is large, and the sound image becomes virtual and the sense of positioning becomes poor. Therefore, this embodiment reduces the angle between the first speaker and the second speaker to reduce the sound field constructed by the first speaker and the second speaker, so as to solve the problem of virtual sound image and poor sense of positioning. Therefore, in this embodiment, when the absolute value of the head rotation angle increases, the angle between the first speaker and the second speaker is reduced. When the absolute value of the head rotation angle decreases, the angle between the first speaker and the second speaker is increased. The angle here can be a horizontal angle. For example, you can also refer to Figure 5 , Figure 5 In the figure, a refers to a schematic diagram of the angle between the first speaker and the second speaker when the head rotation angle is 0 degrees, and the angle of J1 is 120 degrees. In the figure, b refers to a schematic diagram of the angle between the first speaker and the second speaker when the head rotation angle is rotated to the extreme left, and the angle of J2 is 60 degrees. In the figure, c refers to a schematic diagram of the angle between the first speaker and the second speaker when the head rotation angle is rotated to the extreme right, and the angle of J2 is 60 degrees. Figure 5 The direction indicated by the middle arrow is the direction of head rotation. For example, the arrow in a indicates the direction is straight ahead, and the head rotation angle is 0 degrees. The arrow in b indicates the direction is far left, and the head rotation angle is 90 degrees. The arrow in c indicates the direction is far right, and the head rotation angle is 90 degrees.
[0073] Exemplarily, the target horizontal angle corresponding to the horizontal rotation angle is determined in the mapping relationship between the preset horizontal angle and the preset horizontal angle, and the target horizontal angle is evenly divided to obtain the horizontal symmetry adjustment angle; based on the horizontal symmetry adjustment angle, the angles of the first speaker and the second speaker in the horizontal direction are adjusted respectively to obtain the first playback position of the first speaker and the second playback position of the second speaker, so that the angle between the first speaker and the second speaker in the horizontal direction is the target horizontal angle. The horizontal symmetry adjustment angle is half of the target horizontal angle. The horizontal symmetry adjustment angle is the absolute value of the angle of the first speaker relative to the preset central axis, and is also the absolute value of the angle of the second speaker relative to the preset central axis. The preset central axis is the central axis with the head facing straight ahead and horizontal to the ground. In this embodiment, after the angles of the first speaker and the second speaker are adjusted, the orientation of the sound range between the first speaker and the second speaker remains unchanged, but the size of the sound range changes. The sound range refers to the angle area between the first speaker and the second speaker.
[0074] In other embodiments, after determining the target horizontal angle, the target adjustment ratio is determined based on the head rotation angle. Based on the target adjustment ratio, the target horizontal angle is divided to obtain a first adjustment angle and a second adjustment angle. Then, the orientation of the head is obtained. When the head is oriented to the right, the largest of the first adjustment angle and the second adjustment angle is determined to be the left adjustment angle, and the smallest is the right adjustment angle. When the head is oriented to the left, the largest of the first adjustment angle and the second adjustment angle is determined to be the right adjustment angle, and the smallest is the left adjustment angle. The angle of the first speaker is adjusted based on the left adjustment angle, and the angle of the second speaker is adjusted based on the right adjustment angle, so that the horizontal angle between the first speaker and the second speaker is the target horizontal angle. Thus, when the head is turned to the extreme left, to improve the sense of positioning, the angle between the second speaker and the preset central axis is greater than the angle between the first speaker and the preset central axis, thereby causing the sound field between the first and second speakers to shift to the right, so that the user can accurately locate the direction of the sound source. When the head is turned to the extreme right, to improve the sense of positioning, the angle between the second speaker and the preset central axis is less than the angle between the first speaker and the preset central axis, thereby causing the sound field between the first and second speakers to shift to the left, so that the user can accurately locate the direction of the sound source. The target adjustment ratio includes a target first adjustment ratio and a target second adjustment ratio, the sum of which is 1. The greater the head rotation angle, the greater the difference between the first and second adjustment ratios. When the head rotation angle is 0, the first and second adjustment ratios can be the same. In this embodiment, a mapping relationship between rotation angle intervals and preset adjustment ratios can be predetermined. The target rotation angle interval within which the head rotation angle falls can be found in this mapping relationship, and the preset adjustment ratio corresponding to the target rotation angle interval can be used as the target adjustment ratio. The sum of the first and second adjustment ratios in the preset adjustment ratios is 1.
[0075] This embodiment determines the target horizontal angle by the head rotation angle, and then facilitates adjusting the angles of the first speaker and the second speaker according to the target horizontal angle, so as to compensate for the problems of virtual sound image and poor positioning sense caused by head rotation by adjusting the angle of the virtual speaker, so as to improve the sound quality performance of the audio device.
[0076] In another feasible embodiment, the head rotation angle includes a horizontal rotation angle and a vertical pitch angle, and step S20 further includes steps A10 to A30: Step A10, determining a target vertical angle corresponding to the vertical pitch angle in a mapping relationship between a preset pitch angle and a preset vertical angle; Step A20, determining a target horizontal angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal angle; Step A30: adjusting the vertical angles of the first and second virtual speakers based on the target vertical angle, and adjusting the horizontal angles of the first and second virtual speakers based on the target horizontal angle, to obtain a first playback position of the first speaker and a second playback position of the second speaker. The angle between the first playback posture and the second playback posture in the horizontal direction is the target horizontal angle, and the angle between the first playback posture and the second playback posture in the vertical direction is the target vertical angle.
[0077] It should be noted that the mapping relationship between the preset pitch angle and the preset vertical angle can be predetermined, and this embodiment does not specifically limit this. The absolute value of the preset pitch angle is negatively correlated with the preset vertical angle. The target vertical angle is the preset pitch angle corresponding to the vertical pitch angle.
[0078] For example, the target vertical angle can be divided equally to obtain a vertical adjustment angle, and the vertical angles of the first speaker and the second speaker in the virtual speaker can be adjusted according to the vertical adjustment angle. The target horizontal angle can also be divided equally to obtain a horizontal adjustment angle, and the horizontal angles of the first speaker and the second speaker can be adjusted according to the horizontal adjustment angle to obtain the first playback position of the first speaker and the second playback position of the second speaker. The horizontal angle between the first playback position and the second playback position is then made the target horizontal angle, and the vertical angle is made the target vertical angle. In this embodiment, the vertical angle adjustment of the first speaker and the second speaker can be symmetrical or asymmetrical, and this embodiment does not specifically limit this.
[0079] This embodiment determines the target vertical and horizontal angles by head rotation angle, thereby facilitating adjustment of the vertical and horizontal angles of the first and second speakers. This allows for more accurate adjustment of the virtual speaker angles, thereby improving the sound quality of the audio device. In this embodiment, it is possible to adjust only the horizontal angle of the first and second speakers, or to adjust both the horizontal and vertical angles, thereby increasing the flexibility of virtual speaker angle adjustment.
[0080] In a feasible embodiment, the method for controlling an audio device further includes steps X10 to X30: Step X10, receiving a posture adjustment instruction for the virtual speaker, and determining a position adjustment parameter and an angle adjustment parameter of the virtual speaker from the posture adjustment instruction; Step X20, adjusting the position of the virtual speaker according to the position adjustment parameter and the angle adjustment parameter to obtain an adjusted position of the virtual speaker in the audio device; It should be noted that the posture adjustment indication is used to indicate the adjustment of the position and angle of the virtual speaker. The posture adjustment indication can be initiated by the user. For example, when the audio device is an AR or VR device, the audio device can display the posture of the virtual speaker on the display screen. The target user can adjust the position and / or angle of the virtual speaker on the display screen to achieve a personalized playback experience. For example, if the target user wants to listen to audio from position A, the user can adjust the posture of the virtual speaker to position A on the display screen. The posture adjustment indication can indicate the posture adjustment of the first speaker and / or the second speaker.
[0081] The position adjustment parameters may include a first adjustment position of the first speaker and a second adjustment position of the second speaker. The angle adjustment parameters may include a first adjustment angle of the first speaker and a second adjustment angle of the second speaker. The first adjustment position may be a target position of the first speaker, and the first adjustment angle is the target angle of the first speaker; the second adjustment position may be a target position of the second speaker, and the second adjustment angle is the target angle of the second speaker. The adjusted posture includes a first target posture of the first speaker and a second target posture of the second speaker.
[0082] Exemplarily, the current position of the first speaker and the current position of the second speaker are displayed on a display screen corresponding to the audio device, and a user's position adjustment instructions for the first speaker and the second speaker are received. The first adjustment position, the first adjustment angle, the second adjustment position and the second adjustment angle are determined from the position adjustment instructions. According to the first adjustment position and the first adjustment angle, the position of the first speaker is adjusted to determine the first target position. According to the second adjustment position and the second adjustment angle, the position of the second speaker is adjusted to determine the second target position.
[0083] Step X30: Render the audio played by the audio device based on the adjusted posture.
[0084] It should be noted that rendering the audio played by the audio device based on the adjusted posture can provide users with a personalized listening experience and increase the fun of listening. In this embodiment, after rendering the audio played by the audio device based on the adjusted posture, the virtual playback posture of the virtual speaker can still be determined based on the head rotation angle, and the target head-related transfer function HRTF corresponding to the virtual playback posture can be obtained from the personalized acoustic transmission data. The audio played by the audio device is rendered based on the target head-related transfer function HRTF.
[0085] For example, the audio played by the audio device can be re-rendered according to the adjusted posture.
[0086] In a feasible embodiment, step X30 further includes step X31 or step X32: Step X31: Render the audio played by the audio device based on the adjusted position of the virtual speaker and the personalized acoustic transmission data; It should be noted that the posture adjustment includes a first target posture and a second target posture. For example, a first head-related transfer function (HRTF) corresponding to the first target posture can be obtained from the personalized acoustic transmission data, and a second head-related transfer function (HRTF) corresponding to the second target posture can be obtained from the personalized acoustic transmission data; and audio played by the audio device can be rendered based on the first and second head-related transfer functions (HRTF).
[0087] Step X32, obtain the initial posture of the virtual speaker before receiving the posture adjustment indication, and the initial rendering parameters of the audio device for the audio; determine the posture deviation between the initial posture and the adjusted posture, and adjust the initial rendering parameters to obtain the target rendering parameters based on the sound source compensation coefficient corresponding to the posture deviation; render the audio played by the audio device through the target rendering parameters.
[0088] It should be noted that the initial posture is the posture of the virtual speaker before receiving the posture adjustment indication. The initial posture may include the first initial posture of the first speaker and the second initial posture of the second speaker. If possible, the first initial posture may be the first playback posture in the virtual playback posture, and the second initial posture may be the second playback posture in the virtual playback posture. This embodiment does not make any specific limitations on this.
[0089] The posture deviation is the deviation between the initial posture and the adjusted posture, and the posture deviation may include a first position deviation and a first angle deviation of the first speaker, and a second position deviation and a first angle deviation of the second speaker.
[0090] The sound source compensation coefficient may include a first compensation coefficient of the first speaker and a compensation coefficient of the second speaker, the first compensation coefficient may include a first position compensation coefficient and a first angle compensation coefficient, and the second compensation coefficient may include a second position compensation coefficient and a second angle compensation coefficient. The sound source compensation coefficient can be used to compensate for parameters of audio rendering errors caused by posture deviations of virtual speakers. The sound source compensation coefficient can be determined based on a preset posture compensation mapping relationship, which is not specifically limited in this embodiment. The posture compensation mapping relationship includes a mapping relationship between multiple preset posture deviations and their corresponding preset compensation coefficients.
[0091] The initial rendering parameters are the parameters used by the audio device to render audio at the initial virtual speaker position. The target rendering parameters are the target values used by the audio device to render audio after receiving a position adjustment instruction. For example, the initial rendering parameters may include initial binaural delay time, filter coefficients, and cutoff frequency, but this embodiment does not specifically limit these parameters.
[0092] Exemplarily, the initial posture of the virtual speaker before receiving the posture adjustment indication and the initial rendering parameters of the audio device for the audio are obtained; the posture deviation between the initial posture and the adjusted posture is calculated, the sound source compensation coefficient corresponding to the posture deviation is found in the preset posture compensation mapping relationship, and the initial rendering parameters are adjusted to obtain the target rendering parameters; through the target rendering parameters, the audio played by the audio device is re-rendered to simulate the audio being played from the adjusted posture of the virtual speaker, so that the target user can perceive that the audio is played from the adjusted posture, thereby facilitating a personalized listening experience.
[0093] The present application also provides a control device for an audio device. Figure 6 , the control device of the audio device includes: A monitoring module 10 is configured to monitor the target user's head rotation angle and obtain the target user's personalized acoustic transmission data, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) trained based on the target user's personal auditory physiological data; An adjustment module 20 is configured to adjust the position of the virtual speaker in the audio device according to the head rotation angle to obtain a virtual playback position of the virtual speaker; The rendering module 30 is used to obtain the target head-related transfer function HRTF corresponding to the virtual playback posture in the personalized acoustic transmission data, and render the audio played by the audio device according to the target head-related transfer function HRTF to simulate the audio playing at the virtual playback posture.
[0094] The control device for an audio device provided in an embodiment of the present application utilizes the control method for an audio device in the aforementioned embodiment, and is intended to address the issue of poor audio quality in an audio device caused by head rotation. Compared to the prior art, the control method for an audio device provided in an embodiment of the present application achieves the same beneficial effects as the control method for an audio device provided in the aforementioned embodiment, and the other technical features of the control device for an audio device are the same as those disclosed in the aforementioned embodiment, and are not further described here.
[0095] The present application provides an audio device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the control method of the audio device in the above-mentioned embodiment 1.
[0096] Reference below Figure 7 , which shows a schematic diagram of the structure of an audio device suitable for implementing the embodiments of the present application. The audio device in the embodiments of the present application can be a near-ear audio device or a far-ear audio device. Near-ear open-type devices include but are not limited to AR devices, VR devices, smart audio glasses, neck-hanging speakers, open-type headphones, mobile phones, tablets, etc. Far-ear open-type devices include but are not limited to audio devices such as speakers and televisions. Figure 7 The audio device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0097] like Figure 7 As shown, the audio device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the audio device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speakers, or vibrator; a storage device 1003 including, for example, a magnetic tape or hard disk; and a communication device 1009. The communication device 1009 may allow the audio device to communicate with other devices wirelessly or wired to exchange data. Although the figures show an audio device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.
[0098] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0099] The audio device provided in this application, utilizing the audio device control method of the aforementioned embodiment, can resolve the issue of poor audio quality caused by head rotation. Compared to the prior art, the beneficial effects of the audio device provided in this application are the same as those of the audio device control method provided in the aforementioned embodiment. Other technical features of this audio device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.
[0100] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0101] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0102] This embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, and the computer-readable program instructions are used to execute the control method of the audio device in the above-mentioned embodiment 1.
[0103] The computer-readable storage medium provided in the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, equipment, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable EPROM (Electrical Programmable Read Only Memory) or flash memory, optical fiber, a portable compact disc CD-ROM (compact disc read-only memory), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution device, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0104] The computer-readable storage medium may be included in the audio device, or may exist independently without being incorporated into the audio device.
[0105] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the audio device, the audio device is enabled to: monitor the head rotation angle of the target user and obtain the personalized acoustic transmission data of the target user, wherein the personalized acoustic transmission data includes at least one head-related transfer function HRTF obtained by training based on the personal auditory physiological data of the target user; adjust the posture of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playback posture of the virtual speaker; obtain the target head-related transfer function HRTF corresponding to the virtual playback posture in the personalized acoustic transmission data, and render the audio played by the audio device according to the target head-related transfer function HRTF to simulate playing the audio at the virtual playback posture. The present application solves the problem of poor sound quality performance of audio devices caused by head rotation.
[0106] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a LAN (local area network) or WAN (wide area network), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based device that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0108] The modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0109] The computer-readable storage medium provided in the embodiments of the present application stores computer-readable program instructions for executing the aforementioned audio device control method, aiming to address the issue of poor audio quality in audio devices caused by head rotation. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in the embodiments of the present application are the same as those of the audio device control method provided in the aforementioned embodiments, and are not further elaborated here.
[0110] An embodiment of the present application further provides a computer program product, including a computer program, which implements the steps of the above-mentioned method for controlling an audio device when the computer program is executed by a processor.
[0111] The computer program product provided in the embodiments of this application is intended to address the issue of poor audio quality in audio devices caused by head rotation. Compared to the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as those of the audio device control method provided in the above-mentioned embodiments, and are not further elaborated here.
[0112] The above are only preferred embodiments of the embodiments of the present application, and do not limit the patent scope of the embodiments of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of the embodiments of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the embodiments of the present application.
Claims
1. A method for controlling an audio device, characterized in that: The method comprises: Monitoring the head rotation angle of the target user and obtaining personalized acoustic transmission data of the target user, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) trained based on the personal auditory physiological data of the target user; Adjusting the position of a virtual speaker in the audio device according to the head rotation angle to obtain a virtual playback position of the virtual speaker; A target head-related transfer function HRTF corresponding to the virtual playback posture is obtained from the personalized acoustic transmission data, and the audio played by the audio device is rendered according to the target head-related transfer function HRTF to simulate playing the audio at the virtual playback posture.
2. The method for controlling an audio device according to claim 1, wherein: The step of obtaining the personalized acoustic transmission data of the target user includes: The preset model is trained by using the acquired group auditory training data set to obtain a trained general model, wherein the group auditory training data set includes a plurality of human auditory physiological data and their corresponding HRTF labels; Acquiring personal auditory physiological data of the target user; A universal decoder is obtained from the universal model, and personalized training is performed on the universal decoder using the personal auditory physiological data to obtain a personalized decoder. Personalized acoustic transmission data corresponding to the personal auditory physiological data is generated based on the personalized decoder.
3. The method for controlling an audio device according to claim 2, wherein: The preset model includes a preset encoder and a preset decoder; the step of training the preset model using the acquired group auditory training data set to obtain a trained universal model includes: Acquiring target human auditory physiological data from the group auditory training data set, inputting the target human auditory physiological data into the preset model, encoding the target human auditory physiological data through a preset encoder in the preset model, inputting the encoded result into the preset decoder, and outputting the HRTF training result through the preset decoder; Determine the training loss between the HRTF training result and the target HRTF label of the target human auditory physiological data. If the training loss is greater than a preset loss threshold, re-acquire new target human auditory physiological data from the group auditory training dataset, and return to the step of inputting the target human auditory physiological data into the preset model until the training loss is less than or equal to the preset loss threshold, thereby obtaining a trained general model.
4. The method for controlling an audio device according to claim 2, wherein: The step of obtaining the target user's personal auditory physiological data includes: Obtaining head structural parameters of the target user, including auricle height, auricle width, ear canal inclination, head width, and shoulder and neck parameters; A head geometric model of the target user is constructed using the head structural parameters, and personal auditory physiological data is extracted from the head geometric model.
5. The method for controlling an audio device according to claim 2, wherein: The step of obtaining the target user's personal auditory physiological data includes: Obtaining three-dimensional head scan data of the user, and determining personal auditory physiological data from the three-dimensional head scan data; and / or, Obtaining head measurement parameters of the user, obtaining target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters from a preset auditory physiological database, and obtaining personal auditory physiological data consisting of the plurality of target sub-auditory physiological data, wherein the head measurement parameters include a plurality of sub-measurement parameters, namely, an outer ear structure parameter, a head-neck connection parameter, and a head body parameter, and the preset auditory physiological database includes the preset sub-auditory physiological data corresponding to the plurality of preset sub-measurement parameters.
6. The method for controlling an audio device according to claim 1, wherein: The head rotation angle includes a horizontal rotation angle, the virtual speaker includes a first speaker and a second speaker, and the virtual playback posture includes a first playback posture and a second playback posture; The step of adjusting the posture of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playback posture of the virtual speaker includes: Determining a target horizontal angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal angle; According to the target horizontal angle, the angles of the first speaker and the second speaker in the horizontal direction are adjusted respectively to obtain a first playback position of the first speaker and a second playback position of the second speaker, so that the angle between the first speaker and the second speaker in the horizontal direction is the target horizontal angle.
7. The method for controlling an audio device according to claim 1, wherein: The head rotation angle includes a horizontal rotation angle and a vertical pitch angle, and the virtual playback posture includes a first playback posture and a second playback posture; The step of adjusting the posture of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playback posture of the virtual speaker includes: Determining a target vertical angle corresponding to the vertical pitch angle in a mapping relationship between a preset pitch angle and a preset vertical angle; Determining a target horizontal angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal angle; Adjusting the vertical angles of a first speaker and a second speaker in the virtual speakers according to the target vertical angle, and adjusting the horizontal angles of the first speaker and the second speaker according to the target horizontal angle to obtain a first playback position of the first speaker and a second playback position of the second speaker; The angle between the first playback posture and the second playback posture in the horizontal direction is the target horizontal angle, and the angle between the first playback posture and the second playback posture in the vertical direction is the target vertical angle.
8. The method for controlling an audio device according to claim 1, wherein: The control method of the audio device further includes: receiving a posture adjustment instruction for the virtual speaker, and determining a position adjustment parameter and an angle adjustment parameter of the virtual speaker from the posture adjustment instruction; Adjusting the position of the virtual speaker according to the position adjustment parameter and the angle adjustment parameter to obtain an adjusted position of the virtual speaker in the audio device; Based on the adjusted posture, the audio played by the audio device is rendered.
9. The method for controlling an audio device according to claim 8, wherein: The step of rendering the audio played by the audio device based on the adjusted posture includes: Rendering the audio played by the audio device according to the adjusted position and personalized acoustic transmission data of the virtual speaker; or Obtain the initial posture of the virtual speaker before receiving the posture adjustment indication, and the initial rendering parameters of the audio device for the audio; determine the posture deviation between the initial posture and the adjusted posture, and adjust the initial rendering parameters to obtain target rendering parameters based on the sound source compensation coefficient corresponding to the posture deviation; render the audio played by the audio device using the target rendering parameters.
10. An audio device, characterized in that: The audio device includes at least one processor, and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the method for controlling an audio device according to any one of claims 1 to 9.
11. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, on which is stored a program for implementing a method for controlling an audio device. The program for implementing a method for controlling an audio device is executed by a processor to implement the steps of the method for controlling an audio device according to any one of claims 1 to 9.
Citation Information
Patent Citations
Audio signal output method and device
CN110072172A
Projector sound effect implementation method and device
CN115802275A
Virtual surround sound rendering method and device, equipment and storage medium
CN116193196A
Equipment control method, electronic equipment and product
CN116367043A
Sound source direction virtualization method and device, equipment and medium
CN116567517A
Cited By
Immersive sound field rendering method and system for acoustic loudspeaker
CN121985284A