Audio device, control method thereof, and storage medium
By monitoring the head rotation angle and adjusting the virtual speaker pose, and using the Head-Related Transfer Function (HRTF) in personalized acoustic transmission data to render audio, the problem of poor sound quality caused by head rotation is solved, achieving accurate sound source perception and sound quality improvement when the head rotates.
Patent Information
- Application Number
- CN202511100959.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-07
AI Technical Summary
In existing technologies, the problem of poor sound quality performance of audio devices caused by head rotation, especially in binaural synthesized audio systems with head tracking technology, is that the sound image becomes blurred, resulting in a poor sense of sound source localization.
By monitoring the head rotation angle of the target user, the Head-Related Transfer Function (HRTF) in the personalized acoustic transmission data is obtained. The virtual speaker pose in the audio device is adjusted, and the audio is rendered according to the HRTF to simulate playing audio at the virtual playback pose.
Even when the head is turned, the target user can accurately perceive the direction of the sound source, thus improving the sound quality performance of the audio device.
Smart Images

Figure CN120602885B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of audio technology, in particular to an audio device and a control method thereof, and a storage medium. BACKGROUND
[0002] In recent years, binaural synthesis audio systems using head tracking technology have become increasingly popular. Head movement tracking technology allows listeners to perceive sound coming from multiple directions, providing an immersive listening experience. However, in the application process, it is found that when the head turns, the sound image becomes virtual, causing the sound source positioning to become poor, and thus the sound quality of the audio device is not good. Therefore, there is currently a problem of poor sound quality of the audio device caused by head rotation.
[0003] The above content is only used to assist in understanding the technical solutions of the embodiments of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY
[0004] The main purpose of the embodiments of the present application is to provide an audio device and a control method thereof, and a storage medium, which aims to solve the problem of poor sound quality of the audio device caused by head rotation.
[0005] To achieve the above-mentioned purpose, the embodiments of the present application provide a control method of an audio device, the method comprising:
[0006] monitoring the head rotation angle of a target user and obtaining individual acoustic transmission data of the target user, wherein the individual acoustic transmission data comprises at least one head-related transfer function HRTF trained based on personal auditory physiological data of the target user;
[0007] adjusting the pose of a virtual loudspeaker in the audio device according to the head rotation angle to obtain a virtual playing pose of the virtual loudspeaker;
[0008] obtaining a target head-related transfer function HRTF corresponding to the virtual playing pose in the individual acoustic transmission data, and rendering audio played by the audio device according to the target head-related transfer function HRTF to simulate playing the audio at the virtual playing pose.
[0009] In an embodiment, the step of obtaining the individual acoustic transmission data of the target user comprises:
[0010] training a preset model to obtain a trained general model by using the obtained group auditory training data set, wherein the group auditory training data set comprises multiple human auditory physiological data and respective corresponding HRTF labels;
[0011] obtaining personal auditory physiological data of the target user;
[0012] obtaining a general decoder from the general model, personalizing the general decoder by the personal auditory physiological data to obtain a personalized decoder, and generating the personalized acoustic transmission data corresponding to the personal auditory physiological data based on the personalized decoder.
[0013] In an embodiment, the preset model comprises a preset encoder and a preset decoder; and the step of training the preset model by the obtained group auditory training dataset to obtain the trained general model comprises:
[0014] obtaining target human auditory physiological data from the group auditory training dataset, inputting the target human auditory physiological data into the preset model, encoding the target human auditory physiological data by the preset encoder in the preset model, inputting the encoded result into the preset decoder, and outputting an HRTF training result by the preset decoder;
[0015] determining a training loss between the HRTF training result and a target HRTF label of the target human auditory physiological data, re-obtaining new target human auditory physiological data from the group auditory training dataset in a case where the training loss is greater than a preset loss threshold, and returning to the step of inputting the target human auditory physiological data into the preset model until the training loss is less than or equal to the preset loss threshold, to obtain the trained general model.
[0016] In an embodiment, the step of obtaining the personal auditory physiological data of the target user comprises:
[0017] obtaining head structure parameters of the target user, the head structure parameters comprising auricle height, auricle width, ear canal angle, head width, and shoulder-neck parameter;
[0018] constructing a head geometric model of the target user by the head structure parameters, and extracting personal auditory physiological data from the head geometric model.
[0019] In an embodiment, the step of obtaining the personal auditory physiological data of the target user comprises:
[0020] obtaining three-dimensional head scan data of the user, and determining personal auditory physiological data from the three-dimensional head scan data; and / or,
[0021] obtaining a head measurement parameter of the user, obtaining target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameter from a preset auditory physiological database, to obtain personal auditory physiological data composed of a plurality of target sub-auditory physiological data, wherein the plurality of sub-measurement parameters included in the head measurement parameter are respectively an external ear structure parameter, a head-neck connection parameter and a head body parameter, and the preset auditory physiological database includes preset sub-auditory physiological data corresponding to a plurality of preset sub-measurement parameters.
[0022] In an embodiment, the head rotation angle includes a horizontal rotation angle, the virtual speaker includes a first speaker and a second speaker, and the virtual playback pose includes a first playback pose and a second playback pose.
[0023] The step of adjusting the pose of the virtual speaker in the audio device according to the head rotation angle to obtain a virtual playback pose of the virtual speaker includes:
[0024] determining a target horizontal included angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal included angle;
[0025] adjusting the angles of the first speaker and the second speaker in the horizontal direction according to the target horizontal included angle to obtain a first playback pose of the first speaker and a second playback pose of the second speaker, so that the included angle of the first speaker and the second speaker in the horizontal direction is the target horizontal included angle.
[0026] In an embodiment, the head rotation angle includes a horizontal rotation angle and a vertical pitch angle, and the virtual playback pose includes a first playback pose and a second playback pose.
[0027] The step of adjusting the pose of the virtual speaker in the audio device according to the head rotation angle to obtain a virtual playback pose of the virtual speaker includes:
[0028] determining a target vertical included angle corresponding to the vertical pitch angle in a mapping relationship between a preset pitch angle and a preset vertical included angle;
[0029] determining a target horizontal included angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal included angle;
[0030] adjusting the angles of the first speaker and the second speaker in the virtual speaker in the vertical direction according to the target vertical included angle, and adjusting the angles of the first speaker and the second speaker in the horizontal direction according to the target horizontal included angle, to obtain a first playback pose of the first speaker and a second playback pose of the second speaker;
[0031] The included angle of the first playing pose and the second playing pose in the horizontal direction is a target horizontal included angle, and the included angle in the vertical direction is a target vertical included angle.
[0032] In an embodiment, the control method of the audio device further includes:
[0033] receiving a pose adjustment instruction of the virtual speaker, determining a position adjustment parameter and an angle adjustment parameter of the virtual speaker from the pose adjustment instruction;
[0034] adjusting the pose of the virtual speaker according to the position adjustment parameter and the angle adjustment parameter to obtain an adjusted pose of the virtual speaker in the audio device;
[0035] rendering the audio played by the audio device based on the adjusted pose.
[0036] In an embodiment, the step of rendering the audio played by the audio device based on the adjusted pose includes:
[0037] rendering the audio played by the audio device according to the adjusted pose of the virtual speaker and the individual acoustic transmission data; or,
[0038] obtaining an initial pose of the virtual speaker before receiving the pose adjustment instruction and an initial rendering parameter of the audio by the audio device; determining a pose deviation between the initial pose and the adjusted pose, adjusting the initial rendering parameter to obtain a target rendering parameter according to a sound source compensation coefficient corresponding to the pose deviation; and rendering the audio played by the audio device through the target rendering parameter.
[0039] In addition, to achieve the above-mentioned purpose, the embodiment further provides a control device of an audio device, the device includes:
[0040] a monitoring module configured to monitor a head rotation angle of a target user and obtain individual acoustic transmission data of the target user, wherein the individual acoustic transmission data includes at least one head-related transfer function (HRTF) trained based on personal auditory physiological data of the target user;
[0041] an adjusting module configured to adjust a pose of a virtual speaker in the audio device according to the head rotation angle to obtain a virtual playing pose of the virtual speaker;
[0042] a rendering module configured to obtain a target head-related transfer function (HRTF) corresponding to the virtual playing pose from the individual acoustic transmission data, and render the audio played by the audio device according to the target head-related transfer function (HRTF) to simulate playing the audio at the virtual playing pose.
[0043] Further, the embodiment of the present application also provides an audio device, which comprises a memory, a processor and a program of the control method of the audio device stored in the memory and executable on the processor, and the program of the control method of the audio device is executable to realize the steps of the control method of the audio device when executed by the processor.
[0044] Further, the embodiment of the present application also provides a computer readable storage medium, which stores a program of the control method of the audio device, and the program of the control method of the audio device is executable to realize the steps of the control method of the audio device when executed by the processor.
[0045] Further, the embodiment of the present application also provides a computer program product, which comprises a computer program, and the computer program is executable to realize the steps of the control method of the audio device when executed by the processor.
[0046] The one or more technical solutions provided by the embodiment of the present application have at least the following technical effects: the present application can monitor the head rotation angle of the target user, and adjust the pose of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playing pose of the virtual speaker. Since the virtual speaker is used to simulate the spatial position of the real sound source, the angle of the virtual speaker will also affect the judgment of the brain on the direction of the sound source. Therefore, the present application adjusts the pose of the virtual speaker according to the head rotation angle to obtain the virtual playing pose of the virtual speaker, so as to compensate the head rotation angle of the target user by adjusting the pose of the virtual speaker, so that the target user can still accurately perceive the direction of the sound source.
[0047] Specifically, the target head-related transfer function HRTF (Head-Related Transfer Function) corresponding to the virtual playing pose is acquired in the personalized acoustic transmission data, and the audio played by the audio equipment is rendered according to the target head-related transfer function HRTF, so as to simulate playing the audio at the virtual playing pose. Since the personalized acoustic transmission data includes at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user, each head-related transfer function HRTF in the personalized acoustic transmission data can reflect the personal auditory physiological characteristics of the target user. Therefore, the target head-related transfer function HRTF is acquired through the virtual playing pose, and the audio played by the audio equipment is rendered through the target head-related transfer function HRTF, so that the rendered audio is more consistent with the auditory physiological characteristics of the target user and can better match the head rotation angle of the target user. Even in the case that the head of the target user rotates, the target user can accurately perceive the sound source direction, thereby improving the sound quality performance of the audio equipment. BRIEF DESCRIPTION OF DRAWINGS
[0048] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate the embodiments of the present application and, together with the description, serve to explain the principles of the present application.
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, those skilled in the art can obtain other drawings according to these drawings without any creative effort.
[0050] Figure 1 A flowchart of an embodiment of the control method of the audio equipment of the present application;
[0051] Figure 2 A flowchart of acquiring personal auditory physiological data in the control method of the audio equipment of the present application;
[0052] Figure 3 A scene diagram of generating personalized acoustic transmission data in the control method of the audio equipment of the present application;
[0053] Figure 4 A curve diagram of the angle of the virtual loudspeaker corresponding to the included angle changing with the head rotation angle in the control method of the audio equipment of the present application;
[0054] Figure 5 An angle diagram of the angle of the virtual loudspeaker corresponding to the included angle in the control method of the audio equipment of the present application;
[0055] Figure 6 A module schematic diagram of a control device of an audio device according to an embodiment of the present application;
[0056] Figure 7 A device structure schematic diagram of a hardware running environment involved in a control method of an audio device according to an embodiment of the present application.
[0057] The purposes, functional features and advantages of the embodiments of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0058] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the embodiments of the present application, and are not intended to limit the embodiments of the present application.
[0059] In order to better understand the technical solutions of the embodiments of the present application, the following will be described in detail in combination with the accompanying drawings and specific embodiments.
[0060] Binaural synthesis audio systems using head tracking technology are increasingly popular. Head movement tracking technology can make the listener feel that the sound comes from multiple directions, creating an immersive and realistic listening experience, providing spatial and directional cues for the listener. With the rise of virtual and augmented reality technology (VR (Virtual Reality) and AR (Augmented Reality)), binaural audio systems have become more important, and binaural audio uses head-related transfer functions (HRTF), which are crucial for positioning sound and spatial cues in space. In virtual and augmented reality, binaural audio systems can create realistic soundscapes that match the user's visual environment, providing a more immersive and engaging experience compared to using visual cues alone. From gaming and entertainment or medicine to education and training, the demand for binaural synthesized 3D (3-Dimensional) spatialization is also increasing.
[0061] The binaural synthesis process uses head-related transfer functions (HRTF) as input to assign a direction of arrival for sound from a virtual sound source in the created virtual environment. A generic HRTF is usually used for this purpose. Although the generic HRTF has been widely used in the field of spatial audio and can create a more accurate VR / AR experience, the auditory physiological characteristics of different users differ individually. If a generic HRTF is used for audio rendering, it is difficult to provide a good sound quality experience for each user.
[0062] And in the application process, it is found that when the head turns, the sound image becomes virtual, the sound source positioning feeling becomes poor, and the sound quality performance of the audio equipment is poor. Therefore, the sound quality performance of the audio equipment caused by the head rotation is poor.
[0063] Therefore, the control method of the audio equipment provided in the embodiment can provide good sound quality experience for each user, and can provide good sound quality experience for the user even if the user's head rotates. In the embodiment, the head rotation angle of the target user can be monitored, and the virtual playback pose of the virtual loudspeaker can be obtained by adjusting the pose of the virtual loudspeaker in the audio equipment according to the head rotation angle. Since the virtual loudspeaker is used to simulate the spatial position of the real sound source, the angle of the virtual loudspeaker will also affect the judgment of the brain on the sound source direction. Therefore, the virtual playback pose of the virtual loudspeaker is obtained by adjusting the pose of the virtual loudspeaker according to the head rotation angle, so that the head rotation angle of the target user can be compensated by adjusting the pose of the virtual loudspeaker, so that the target user can still accurately perceive the sound source direction.
[0064] Specifically, the target head-related transfer function HRTF corresponding to the virtual playback pose is obtained in the individual acoustic transfer data, and the audio played by the audio equipment is rendered according to the target head-related transfer function HRTF, so as to simulate the playback of the audio at the virtual playback pose. Since the individual acoustic transfer data includes at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user, each head-related transfer function HRTF in the individual acoustic transfer data can reflect the personal auditory physiological characteristics of the target user. Therefore, the target head-related transfer function HRTF is obtained through the virtual playback pose, the audio played by the audio equipment is rendered through the target head-related transfer function HRTF, so that the rendered audio is more in line with the auditory physiological characteristics of the target user, and can better match the head rotation angle of the target user. Even in the case that the head of the target user rotates, the target user can accurately perceive the sound source direction, thereby improving the sound quality performance of the audio equipment.
[0065] Based on this, the control method of the audio equipment provided in the embodiment of the application is provided, which is described with reference to Figure 1 , Figure 1 The flowchart of the control method of the audio equipment in the first embodiment of the embodiment of the application is shown. The control method of the audio equipment includes steps S10-S30:
[0066] Step S10, monitor the head rotation angle of the target user, and obtain the individual acoustic transfer data of the target user, wherein the individual acoustic transfer data includes at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user;
[0067] It should be noted that the audio device can be a near-ear audio device or a far-ear audio device. The near-ear open device includes, but is not limited to, an AR device, a VR device, smart audio glasses, a neck-hung audio box, an open earphone, a mobile phone, a tablet, etc. The far-ear open device includes, but is not limited to, an audio device of a sound box, a television, etc. The head rotation angle is a rotation angle of the target user relative to a preset zero position direction. The preset zero position direction can be a direction in which the head of the target user faces the front. In other embodiments, the preset zero position direction can be determined based on actual conditions. For example, the preset zero position direction corresponding to the audio device can be determined based on different audio devices. The present embodiment does not make specific limitations.
[0068] The personalized acoustic transmission data can include at least one head-related transfer function (HRTF). The head-related transfer function (HRTF) is a frequency domain filter function describing the reflection, diffraction, and scattering effects of physiological structures such as the head, pinna, and shoulder on sound waves during the propagation of sound from a free field to the ears. The head-related transfer function (HRTF) can be used to simulate the changes of sound waves from the sound source position to the human ear, and can make the virtual sound source have real spatial properties. For the same target user, there can be multiple head-related transfer functions (HRTF). The head-related transfer functions (HRTF) are different, and the simulated sound source positions are different. The commonality of different head-related transfer functions (HRTF) of the same target user is that they all conform to the personal auditory physiological characteristics of the target user, and only the simulated sound source positions can be different.
[0069] For example, the head rotation angle of the target user is monitored, the personal auditory physiological data of the target user is obtained, and the personalized decoder is used to generate the personalized acoustic transmission data of the personal auditory physiological data.
[0070] In a feasible embodiment, step S10 further includes steps S11-S13:
[0071] In step S11, the preset model is trained based on the obtained group auditory training data set to obtain a trained general model. The group auditory training data set includes multiple human auditory physiological data and respective corresponding HRTF labels.
[0072] It should be noted that the population auditory training dataset includes a plurality of human auditory physiological data, specifically, a deep learning model is trained using a large amount of existing HRTF data (the existing HRTF data is open source HRTF data), that is, the population auditory training dataset is obtained. The existing HRTF data can include a large amount of HRTF data of different people, covering a large amount of HRTF data of people of different ages, ear shapes, and head shapes. That is, the plurality of human auditory physiological data is derived from different human bodies, and the preset model can include a preset encoder and a preset decoder, and each human auditory physiological data corresponds to an HRTF label; the population auditory training dataset can be predetermined, and each human auditory physiological data can reflect the auditory physiological characteristics of the human body, for example, the human auditory physiological data can include data of the ear canal resonance frequency, the angle between the tragus and the antihelix, and the complexity of the pinna fold, and the present embodiment does not make specific limitation thereon. The acquisition method of the human auditory physiological data can be the same as that of the personal auditory physiological data of the target user. The human auditory physiological data is data that can be processed by the preset encoder. The general model includes a general decoder and a general encoder. After the preset model is trained, the general model composed of the general decoder and the general encoder can be obtained.
[0073] For example, the population auditory training dataset is obtained, the preset model is trained through the population auditory training dataset, and the general model composed of the general decoder and the general encoder is obtained. The general model can also be used to generate a general HRTF, and the general HRTF can also simulate the change of sound waves when the sound source propagates to the human ear, but the general HRTF is difficult to adapt to the individual differences of the user's hearing, so there is still a problem that the sound quality of the audio device is not good when the audio is rendered through the general HRTF. Therefore, after the general model is determined, the general decoder in the general model is fine-tuned to obtain a personalized decoder, so that the personalized acoustic transmission data of the target user can be generated through the personalized decoder subsequently.
[0074] In a feasible embodiment, step S11 further includes steps S111-S112:
[0075] Step S111, obtaining target human auditory physiological data from the population auditory training dataset, inputting the target human auditory physiological data into the preset model, encoding the target human auditory physiological data through the preset encoder in the preset model, inputting the encoding result into the preset decoder, and outputting the HRTF training result through the preset decoder;
[0076] In step S112, a training loss between the HRTF training result and the target HRTF label of the target human auditory physiological data is determined, and in a case where the training loss is greater than a preset loss threshold, new target human auditory physiological data is re-acquired from the group auditory training data set, and the step of inputting the target human auditory physiological data into the preset model is returned until the training loss is less than or equal to the preset loss threshold, and a general model trained is obtained.
[0077] It should be noted that the preset encoder can encode the target human auditory physiological data, and can encode the high-dimensional target human auditory physiological data into a low-dimensional feature vector related to hearing, so as to perform down-sampling encoding on the target human auditory physiological data. Specifically, the preset encoder can encode the features affecting the HRTF from the target human auditory physiological data, thereby reducing the model complexity and improving the training accuracy. The result of encoding the target human auditory physiological data by the preset encoder can be input into the preset decoder, and the result of encoding is decoded by the preset decoder, thereby generating the HRTF training result. The HRTF training result is the HRTF corresponding to the target human auditory physiological data output by the preset decoder in the preset model.
[0078] The training loss can represent the difference between the target HRTF label and the HRTF training result. The greater the training loss, the greater the difference between the target HRTF label and the HRTF training result. The preset loss threshold can be determined based on actual conditions, which is not limited in the embodiment. In a case where the training loss is greater than the preset loss threshold, it indicates that the difference between the target HRTF label and the HRTF training result is large, so new target human auditory physiological data needs to be acquired for training until the training loss is less than or equal to the preset loss threshold, and a general model trained is obtained.
[0079] Exemplarily, the target human auditory physiological data is input into the preset model, the target human auditory physiological data is encoded by the preset encoder in the preset model, the encoded result is input into the preset decoder, and the HRTF training result is output by the preset decoder; the training loss between the HRTF training result and the target HRTF label of the target human auditory physiological data is calculated, and in a case where the training loss is greater than a preset loss threshold, new target human auditory physiological data is re-acquired from the group auditory training data set, and the step of inputting the target human auditory physiological data into the preset model is returned until the training loss is less than or equal to the preset loss threshold, and a general model trained is obtained.
[0080] The preset model is trained by the group auditory training data set to obtain a general model, and then the general decoder in the general model is used to determine the individual acoustic transmission data, so that each user can provide a good sound quality experience.
[0081] Step S12: Obtain the target user's personal auditory physiological data;
[0082] It should be noted that personal auditory physiological data can reflect the individual auditory physiological characteristics of the target user. Personal auditory physiological data can be obtained through 3D head scanning, from measurements of the target user's head, or through geometric modeling of the user's head. This embodiment does not impose specific limitations on these methods.
[0083] In a feasible embodiment, step S12 further includes steps S121 to S122:
[0084] Step S121: Obtain the head structure parameters of the target user, including auricle height, auricle width, ear canal tilt angle, head width, and shoulder and neck parameters.
[0085] Step S122: Construct a head geometric model of the target user using head structure parameters, and extract personal auditory physiological data from the head geometric model.
[0086] It should be noted that head width is the lateral distance between the left and right tragus, and head width determines the interaural time difference (ITD). Shoulder and neck parameters can include shoulder width, neck height, etc., and these parameters affect the reflection and scattering of low-frequency sound waves. Auricular height is the vertical distance from the tragus to the highest point of the helix, and auricular width can be the maximum horizontal distance from the anterior edge of the helix to the posterior edge of the antihelix. The ear canal tilt angle is the angle between the ear canal axis and the sagittal plane (plane of symmetry) of the head, which affects the direction of sound wave incidence.
[0087] A geometric model of the target user's head can be constructed using head structural parameters. 3D modeling software can be used to construct this geometric model based on these parameters. Features directly related to auditory function can be extracted directly from the head geometric model to obtain individual auditory physiological data. For example, this data may include ear canal resonant frequencies, the angle between the tragus and cymba conchae, and the ITD baseline value. This embodiment does not impose specific limitations on these parameters; they can be determined based on actual circumstances. In other embodiments, head structural parameters may also include jaw width, ear canal length, etc., which are not specifically limited in this embodiment.
[0088] For example, the head structure parameters of the target user can be obtained by taking a picture. Modeling software is then called to construct a geometric model of the target user's head based on the head structure parameters, and personal auditory physiological data can be extracted from the head geometric model.
[0089] The embodiment constructs a head geometry model of the target user through head structure parameters, and then extracts personal auditory physiological data from the head geometry model, and then determines personalized acoustic transmission data, so as to improve the sound quality performance of the audio equipment.
[0090] In an available embodiment, the step S12 further comprises a step S123 and / or a step S124.
[0091] In the step S123, three-dimensional head scanning data of the user is acquired, and personal auditory physiological data is determined from the three-dimensional head scanning data.
[0092] In the step S124, head measurement parameters of the user are acquired, target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters is acquired from a preset auditory physiological database, and personal auditory physiological data composed of a plurality of target sub-auditory physiological data is obtained, wherein the plurality of sub-measurement parameters included in the head measurement parameters are respectively external ear structure parameters, head-neck connection parameters, and head body parameters, and the preset auditory physiological database includes preset sub-auditory physiological data corresponding to a plurality of preset sub-measurement parameters.
[0093] It should be noted that the three-dimensional head scanning data can be obtained by 3D structured light and / or laser scanning on the target user, and the personal auditory physiological data related to hearing can be extracted from the three-dimensional head scanning data.
[0094] The head measurement parameters can be obtained by directly measuring the head of the target user and the shoulder and neck connected with the head, for example, can be obtained by multi-angle photography. The head measurement parameters can include a plurality of sub-measurement parameters, and the plurality of sub-measurement parameters can be respectively external ear structure parameters, head-neck connection parameters, and head body parameters. Each sub-measurement parameter can further include a plurality of sub-parameters, and the plurality of sub-parameters included in the external ear structure parameters can be respectively auricle contour, auricle wrinkle, auricle size, ear canal length and diameter, concha cavity depth, etc. The plurality of sub-parameters included in the head-neck connection parameters can be respectively shoulder width and neck length, and the plurality of sub-parameters included in the head body parameters can include head diameter and binaural distance.
[0095] The preset auditory physiological database can be determined in advance based on actual conditions, and the preset auditory physiological database can include preset sub-auditory physiological data corresponding to a plurality of preset sub-measurement parameters. The target sub-auditory physiological data of each sub-measurement parameter can be searched in the preset auditory physiological database.
[0096] For example, the target user can be 3D structured light and / or laser scanned to obtain three-dimensional head scanning data, and the personal auditory physiological data related to hearing can be acquired from the three-dimensional head scanning data.
[0097] The head measurement parameters of the user can also be acquired, and target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters can be acquired from a preset audiological database to obtain personal audiological data composed of multiple target sub-auditory physiological data. Specifically, each sub-measurement parameter further includes multiple sub-parameters, and preset sub-auditory physiological data corresponding to each sub-parameter can be searched from the preset audiological database to obtain personal audiological data composed of multiple preset sub-auditory physiological data. The preset sub-measurement parameter also includes multiple preset sub-parameters, and the preset sub-measurement parameter can correspond to multiple preset sub-auditory physiological data, which is not limited in the embodiment.
[0098] The personal audiological data can be acquired in various ways in the embodiment, thereby improving the flexibility of data acquisition. For example, the personal audiological data can also be acquired by referring to Figure 2 , Figure 2 The various ways of acquiring personal audiological data are given, and the personal audiological data can be extracted from the head geometric model, or extracted from the three-dimensional head scan data, or searched from the preset audiological database to obtain target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters, and then obtain personal audiological data composed of multiple target sub-auditory physiological data.
[0099] In step S13, a general decoder is acquired from the general model, and the general decoder is personalized trained by the personal audiological data to obtain a personalized decoder, and the personalized decoder is used to generate individual acoustic transmission data corresponding to the personal audiological data.
[0100] It should be noted that the general decoder can be unsupervised trained by the personal audiological data to personalize the general decoder. Since the general decoder has learned the general mapping rule between the physiological characteristics of the group and the HRTF, the general decoder can be unsupervised learned by the personal audiological data, and the general decoder can be adjusted accordingly, for example, the network parameters of the general decoder can be adjusted, and then the personalized decoder can be obtained, wherein the network parameters of the general decoder refer to the weights and biases of neurons in the general decoder.
[0101] Exemplarily, the general decoder is obtained from the general model, the general decoder is unsupervised trained through the personal auditory physiological data, and the personalized decoder is obtained. The personal auditory physiological data can be input into the general encoder or the preset encoder to encode the personal auditory physiological data through the general encoder or the preset encoder. The result of the personal auditory physiological data after encoding is input into the personalized decoder, and the personalized decoder outputs the individualized acoustic transmission data of the personal auditory physiological data. The individualized acoustic transmission data can include left ear acoustic transmission data and right ear acoustic transmission data. The left ear acoustic transmission data and the right ear acoustic transmission data can each include a corresponding head-related transfer function HRTF. The left ear acoustic transmission data and the right ear acoustic transmission data can be symmetrical, or can be individually trained to be asymmetrical according to the characteristics of the left and right ears of the target user.
[0102] Specifically, in other embodiments, the left ear acoustic transmission data and the right ear acoustic transmission data can also be asymmetrical. For example, left ear personal physiological data and right ear personal physiological data can be obtained from the personal auditory physiological data. The left ear personal physiological data is input into the general encoder or the preset encoder to encode the left ear personal physiological data through the general encoder or the preset encoder. The result of the left ear personal physiological data after encoding is input into the personalized decoder, and the personalized decoder outputs the left ear acoustic transmission data of the left ear personal physiological data. The right ear personal physiological data is input into the general encoder or the preset encoder to encode the right ear personal physiological data through the general encoder or the preset encoder. The result of the right ear personal physiological data after encoding is input into the personalized decoder, and the personalized decoder outputs the right ear acoustic transmission data of the right ear personal physiological data. In this way, the accuracy of the individualized acoustic transmission data can be improved, and the differences between the left and right ears of the same user can be adapted.
[0103] The embodiment trains the general decoder first, so as to facilitate the adjustment of the network parameters of the general decoder according to the personal auditory physiological data, obtain the personalized decoder, and then facilitate the generation of the individualized acoustic transmission data in combination with the personalized decoder, so as to improve the accuracy of the audio rendering, make the rendered audio conform to the auditory physiological characteristics of the target user, and provide a good sound quality experience for the user.
[0104] For better understanding of the embodiment, reference can be made to Figure 3 , Figure 3A scene diagram of the personalization stage and the training stage of the general model in this embodiment is given. In the training stage of the general model, the group auditory training data set can be input to a preset encoder, and the output result of the preset encoder is input to a preset decoder. The preset decoder can output an HRTF training result. The preset encoder and the preset decoder can be trained until the training of the general decoder is completed. The general decoder can be personalized trained by using the personal auditory physiological data to obtain a personalized decoder.
[0105] In the personalization stage, acoustic measurement can be performed to obtain personal auditory physiological data. The personal auditory physiological data can be used to personalize train the general decoder to obtain a personalized decoder. The personal auditory physiological data can be input to the personalized decoder, and the personalized decoder can output personalized acoustic transmission data. Figure 3 The process of personalizing training the general decoder by using the personal auditory physiological data is not shown in the embodiment.
[0106] In step S20, the pose of the virtual loudspeaker in the audio device is adjusted according to the head rotation angle to obtain a virtual playing pose of the virtual loudspeaker.
[0107] It should be noted that when the target user rotates the head, it indicates that the position of the human ear can change. When the position of the human ear changes, the sound waves propagating from the original sound source position to the human ear will also change, which can affect the target user's judgment of the sound source position and cause poor positioning.
[0108] The virtual loudspeaker can include a first loudspeaker and a second loudspeaker. The first loudspeaker can be a loudspeaker close to the left ear, and the second loudspeaker can be a loudspeaker close to the right ear. The virtual playing pose is the position and angle at which the virtual loudspeaker finally plays audio. Both the first loudspeaker and the second loudspeaker are virtual loudspeakers.
[0109] For example, in this embodiment, the angle of the virtual loudspeaker in the audio device can be adjusted according to the head rotation angle to obtain a virtual playing pose of the virtual loudspeaker to compensate for the head rotation angle through the virtual playing pose. Specifically, the angles of the first loudspeaker and the second loudspeaker in the audio device can be adjusted to obtain a first playing pose of the first loudspeaker and a second playing pose of the second loudspeaker. The virtual playing pose can include the first playing pose and the second playing pose.
[0110] In step S30, the target head-related transfer function HRTF corresponding to the virtual playback pose is obtained in the personalized acoustic transmission data, and the audio played by the audio equipment is rendered according to the target head-related transfer function HRTF, so as to simulate the playback of the audio at the virtual playback pose.
[0111] It should be noted that the virtual playback pose can include a first playback pose of the first loudspeaker and a second playback pose of the second loudspeaker, the first loudspeaker can be a left-ear loudspeaker, and the second loudspeaker can be a right-ear loudspeaker, which are not limited in the embodiment. The first playback pose includes the position and angle of the left-ear loudspeaker, and the second playback pose includes the position and angle of the right-ear loudspeaker. The target head-related transfer function HRTF can include a first target head-related transfer function HRTF of the left ear and a second target head-related transfer function HRTF of the right ear.
[0112] For example, the audio played by the audio equipment can be rendered by the first target head-related transfer function HRTF and the second target head-related transfer function HRTF, so as to simulate the playback of the audio at the first playback pose and the second playback pose.
[0113] The embodiment can monitor the head rotation angle of the target user, and adjust the pose of the virtual loudspeaker in the audio equipment to obtain the virtual playback pose of the virtual loudspeaker according to the head rotation angle. Since the virtual loudspeaker is used to simulate the spatial position of the real sound source, the angle of the virtual loudspeaker will also affect the judgment of the brain on the direction of the sound source. Therefore, by adjusting the pose of the virtual loudspeaker according to the head rotation angle to obtain the virtual playback pose of the virtual loudspeaker, the head rotation angle of the target user can be compensated, so that the target user can still accurately perceive the direction of the sound source.
[0114] Specifically, the target head-related transfer function HRTF corresponding to the virtual playback pose is obtained in the personalized acoustic transmission data, and the audio played by the audio equipment is rendered according to the target head-related transfer function HRTF, so as to simulate the playback of the audio at the virtual playback pose. Since the personalized acoustic transmission data includes at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user, each head-related transfer function HRTF in the personalized acoustic transmission data can reflect the personal auditory physiological characteristics of the target user. Therefore, by obtaining the target head-related transfer function HRTF through the virtual playback pose, and rendering the audio played by the audio equipment through the target head-related transfer function HRTF, the rendered audio is more consistent with the auditory physiological characteristics of the target user, and can better match the head rotation angle of the target user, so that the target user can accurately perceive the direction of the sound source even in the case of head rotation, thereby improving the sound quality performance of the audio equipment.
[0115] In an implementable embodiment, the head rotation angle comprises a horizontal rotation angle, the virtual speaker comprises a first speaker and a second speaker, and the virtual playing position pose comprises a first playing position pose and a second playing position pose; step S20 further comprises steps S21-S22:
[0116] Step S21: determining a target horizontal included angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal included angle;
[0117] Step S22: adjusting the angles of the first speaker and the second speaker in the horizontal direction according to the target horizontal included angle, to obtain the first playing position pose of the first speaker and the second playing position pose of the second speaker, so that the included angle between the first speaker and the second speaker in the horizontal direction is the target horizontal included angle.
[0118] It should be noted that the mapping relationship between the preset horizontal angle and the preset horizontal included angle can be determined in advance, and the present embodiment does not make specific limitation thereon, and the absolute value of the preset horizontal angle is negatively correlated with the preset horizontal included angle. For example, the mapping relationship between the preset horizontal transmission angle and the preset horizontal included angle can be referred to as shown in FIG. 1, Figure 4 , Figure 4 FIG. 1 shows a mapping relationship between a preset horizontal transmission angle and a preset horizontal included angle, Figure 4 which is one of the examples of the included angle change, and the present embodiment does not limit the specific value of the included angle. Figure 4 In FIG. 1, the horizontal coordinate is the head rotation angle, and the vertical coordinate is the included angle of the virtual speaker, Figure 4 the data points given in FIG. 1 represent the extreme values of the included angle change, and in Figure 4 the maximum value of the included angle is 120 degrees, and the minimum value is 60 degrees, when the head rotation angle is 0 degrees, the included angle is the maximum value of 120 degrees, when the head rotation angle is 90 degrees, the included angle is the minimum value of 60 degrees, when the head rotation angle rotates from 0 degrees to 90 degrees, the included angle changes from 120 degrees to 60 degrees, and when the head rotation angle rotates from 90 degrees to 180 degrees, the included angle changes from 60 degrees to 120 degrees, and the change is periodic. The included angle of the virtual speaker refers to the included angle between the first speaker and the second speaker.
[0119] For the dual-channel audio, when the head does not rotate (the head rotation angle is 0 degree), the included angle between the first loudspeaker and the second loudspeaker is 120 degrees, the angle of the first loudspeaker can be -60 degrees, and the angle of the second loudspeaker can be 60 degrees. When the head rotates, when rotating to the extreme left (negative 90 degrees) and the extreme right (positive 90 degrees), the included angle between the first loudspeaker and the second loudspeaker can be 60 degrees, the angle of the first loudspeaker can be -30 degrees, and the angle of the second loudspeaker can be 30 degrees. When the head rotation angle changes from 0 degree to 90 degrees, or from 0 degree to -90 degrees, the horizontal angle of the first loudspeaker can change from -60 degrees to -30 degrees, and the horizontal angle of the second loudspeaker can change from 60 degrees to 30 degrees.
[0120] In the embodiment, between -90 degrees and 90 degrees, -90 degrees can refer to the position of the head rotating to the extreme left, and 90 degrees can refer to the position of the head rotating to the extreme right. The higher the absolute value of the head rotation angle, the smaller the included angle between the first loudspeaker and the second loudspeaker, and the lower the absolute value of the head rotation angle, the larger the included angle between the first loudspeaker and the second loudspeaker. The angles of the first loudspeaker and the second loudspeaker are symmetrical. The higher the absolute value of the head rotation angle, the greater the amplitude of the head rotation. If the angles of the virtual loudspeakers are not adjusted, the sound field range constructed by the first loudspeaker and the second loudspeaker in the virtual loudspeaker is large, and there may be a situation that the sound image becomes virtual and the positioning feeling becomes poor. Therefore, the embodiment reduces the included angle between the first loudspeaker and the second loudspeaker, so as to reduce the sound field range constructed by the first loudspeaker and the second loudspeaker, so as to solve the problem of sound image becoming virtual and positioning feeling becoming poor. Therefore, in the embodiment, when the absolute value of the head rotation angle increases, the included angle between the first loudspeaker and the second loudspeaker is adjusted to be smaller, and when the absolute value of the head rotation angle decreases, the included angle between the first loudspeaker and the second loudspeaker is adjusted to be larger. The included angle here can be a horizontal included angle. For example, it can also be referred to as Figure 5 , Figure 5 In a, the included angle between the first loudspeaker and the second loudspeaker when the head rotation angle is 0 degree is indicated, the included angle of J1 is 120 degrees. In b, the included angle between the first loudspeaker and the second loudspeaker when the head rotation angle rotates to the extreme left is indicated, the included angle of J2 is 60 degrees. In c, the included angle between the first loudspeaker and the second loudspeaker when the head rotation angle rotates to the extreme right is indicated, the included angle of J2 is 60 degrees. Figure 5 In the arrows in a, b and c indicate the direction of the head rotation. For example, the arrow in a indicates the direction of the front, the head rotation angle is 0 degree, the arrow in b indicates the direction of the extreme left, the head rotation angle is 90 degrees, and the arrow in c indicates the direction of the extreme right, the head rotation angle is 90 degrees.
[0121] In an example, the target horizontal included angle corresponding to the horizontal rotation angle is determined in the mapping relationship between the preset horizontal angle and the preset horizontal included angle, the target horizontal included angle is equally divided to obtain a horizontal symmetry adjustment angle; the horizontal symmetry adjustment angle is used to respectively adjust the angles of the first loudspeaker and the second loudspeaker in the horizontal direction to obtain a first playing pose of the first loudspeaker and a second playing pose of the second loudspeaker, so that the included angle between the first loudspeaker and the second loudspeaker in the horizontal direction is the target horizontal included angle. The horizontal symmetry adjustment angle is half of the target horizontal included angle, and the horizontal symmetry adjustment angle is the absolute value of the included angle of the first loudspeaker relative to the preset central axis and the absolute value of the included angle of the second loudspeaker relative to the preset central axis. The preset central axis is a central axis in which the head is directed forward and horizontal to the ground. In this embodiment, the direction of the sound field range between the first loudspeaker and the second loudspeaker is unchanged after the angle adjustment of the first loudspeaker and the second loudspeaker, and only the size of the sound field range changes. The sound field range refers to the included angle region between the first loudspeaker and the second loudspeaker.
[0122] In other embodiments, after the target horizontal included angle is determined, a target adjustment ratio is determined according to the head rotation angle, the target horizontal included angle is divided according to the target adjustment ratio to obtain a first adjustment angle and a second adjustment angle, the orientation of the head is obtained, when the head is oriented to the right, the maximum of the first adjustment angle and the second adjustment angle is determined as a left adjustment angle and the minimum is determined as a right adjustment angle, when the head is oriented to the left, the maximum of the first adjustment angle and the second adjustment angle is determined as a right adjustment angle and the minimum is determined as a left adjustment angle, the angle of the first loudspeaker is adjusted according to the left adjustment angle, and the angle of the second loudspeaker is adjusted according to the right adjustment angle, so that the included angle between the first loudspeaker and the second loudspeaker in the horizontal direction is the target horizontal included angle. Further, when the head is turned to the extreme left, in order to improve the positioning feeling, the included angle between the second loudspeaker and the preset central axis is greater than the included angle between the first loudspeaker and the preset central axis, and further, the sound range between the first loudspeaker and the second loudspeaker is biased to the right, so that the user can accurately locate the sound source direction, and when the head is turned to the extreme right, in order to improve the positioning feeling, the included angle between the second loudspeaker and the preset central axis is less than the included angle between the first loudspeaker and the preset central axis, and further, the sound range between the first loudspeaker and the second loudspeaker is biased to the left, so that the user can accurately locate the sound source direction. The target adjustment ratio includes a target first adjustment ratio and a target second adjustment ratio, the sum of the target first adjustment ratio and the target second adjustment ratio is 1, the greater the head rotation angle, the greater the difference between the first adjustment ratio and the second adjustment ratio, and when the head rotation angle is 0, the first adjustment ratio and the second adjustment ratio can be the same. In this embodiment, the mapping relationship between the rotation angle interval and the preset adjustment ratio can be determined in advance, the target rotation angle interval in which the head rotation angle is located can be found in the mapping relationship, and the preset adjustment ratio corresponding to the target rotation angle interval can be taken as the target adjustment ratio. The sum of the first adjustment ratio and the second adjustment ratio in the preset adjustment ratio is 1.
[0123] In this embodiment, the target horizontal included angle is determined by the head rotation angle, and further, the angles of the first loudspeaker and the second loudspeaker are adjusted according to the target horizontal included angle, so as to compensate for the problems of virtual sound image and poor positioning feeling caused by head rotation by adjusting the angles of the virtual loudspeakers, and to improve the sound quality performance of the audio equipment.
[0124] In another possible embodiment, the head rotation angle includes a horizontal rotation angle and a vertical pitch angle, and step S20 further includes steps A10-A30:
[0125] Step A10, determining the target vertical included angle corresponding to the vertical pitch angle in the mapping relationship between the preset pitch angle and the preset vertical included angle;
[0126] Step A20, determining the target horizontal included angle corresponding to the horizontal rotation angle in the mapping relationship between the preset horizontal angle and the preset horizontal included angle;
[0127] Step A30, adjusting the angles of the first speaker in the virtual speaker and the second speaker in the virtual speaker in the vertical direction respectively according to the target vertical included angle, and adjusting the angles of the first speaker and the second speaker in the horizontal direction respectively according to the target horizontal included angle, to obtain a first playing pose of the first speaker and a second playing pose of the second speaker;
[0128] The included angle between the first playing pose and the second playing pose in the horizontal direction is the target horizontal included angle, and the included angle between the first playing pose and the second playing pose in the vertical direction is the target vertical included angle.
[0129] It should be noted that the mapping relationship between the preset pitch angle and the preset vertical included angle can be determined in advance, and the present embodiment does not make specific limitation thereon, and the absolute value of the preset pitch angle is negatively correlated with the preset vertical included angle. The target vertical included angle is the preset pitch angle corresponding to the vertical pitch angle.
[0130] For example, the target vertical included angle can be equally divided to obtain a vertical adjustment angle, and the angles of the first speaker in the virtual speaker and the second speaker in the virtual speaker in the vertical direction can be adjusted respectively according to the vertical adjustment angle. The target horizontal included angle can also be equally divided to obtain a horizontal adjustment angle, and the angles of the first speaker and the second speaker in the horizontal direction can be adjusted respectively according to the horizontal adjustment angle, to obtain a first playing pose of the first speaker and a second playing pose of the second speaker. Further, the included angle between the first playing pose and the second playing pose in the horizontal direction is the target horizontal included angle, and the included angle between the first playing pose and the second playing pose in the vertical direction is the target vertical included angle. In the present embodiment, the angle adjustment of the first speaker and the second speaker in the vertical direction can be symmetrical or asymmetrical, and the present embodiment does not make specific limitation thereon.
[0131] In the present embodiment, the target vertical included angle and the target horizontal included angle are determined by the head rotation angle, and then the angles of the first speaker and the second speaker in the vertical direction and the horizontal direction are adjusted, so that the angle of the virtual speaker can be adjusted more accurately, so as to improve the sound quality performance of the audio device. In the present embodiment, only the angles of the first speaker and the second speaker in the horizontal direction can be adjusted, or the angles of the first speaker and the second speaker in the horizontal direction and the vertical direction can be adjusted, thereby improving the flexibility of the angle adjustment of the virtual speaker.
[0132] In a feasible embodiment, the control method of the audio device further includes steps X10-X30:
[0133] Step X10, receiving a pose adjustment indication of the virtual speaker, determining a position adjustment parameter and an angle adjustment parameter of the virtual speaker from the pose adjustment indication;
[0134] Step X20, adjusting the pose of the virtual speaker according to the position adjustment parameter and the angle adjustment parameter, to obtain an adjusted pose of the virtual speaker in the audio device;
[0135] It should be noted that the pose adjustment indication is used to indicate the adjustment of the position and angle of the virtual speaker, and the pose adjustment indication can be initiated by the user. For example, when the audio device is an AR or VR device, the audio device can display the pose of the virtual speaker on the display screen, and the target user can adjust the position and / or angle of the virtual speaker on the display screen to achieve personalized playback experience. For example, the target user just wants to hear audio from the A direction, and the user can adjust the pose of the virtual speaker to the A direction on the display screen. The pose adjustment indication can indicate the pose adjustment of the first speaker and / or the second speaker.
[0136] The position adjustment parameter can include a first adjustment position of the first speaker and a second adjustment position of the second speaker, and the angle adjustment parameter can include a first adjustment angle of the first speaker and a second adjustment angle of the second speaker. The first adjustment position can be a target position of the first speaker, and the first adjustment angle can be a target angle of the first speaker. The second adjustment position can be a target position of the second speaker, and the second adjustment angle can be a target angle of the second speaker. The adjusted pose includes a first target pose of the first speaker and a second target pose of the second speaker.
[0137] For example, the current pose of the first speaker and the current pose of the second speaker are displayed on the display screen corresponding to the audio device, the pose adjustment indication of the first speaker and the second speaker is received from the user, the first adjustment position, the first adjustment angle, the second adjustment position and the second adjustment angle are determined from the pose adjustment indication, the pose of the first speaker is adjusted according to the first adjustment position and the first adjustment angle to determine the first target pose, and the pose of the second speaker is adjusted according to the second adjustment position and the second adjustment angle to determine the second target pose.
[0138] Step X30, rendering the audio played by the audio device based on the adjusted pose.
[0139] It should be noted that the audio played by the audio device is rendered according to the adjusted pose, which can provide the user with a personalized listening experience and increase the listening interest. In this embodiment, after the audio played by the audio device is rendered according to the adjusted pose, the virtual playback pose of the virtual speaker can still be determined according to the head rotation angle, and the target head-related transfer function HRTF corresponding to the virtual playback pose can be obtained in the personalized acoustic transfer data, and the audio played by the audio device is rendered according to the target head-related transfer function HRTF.
[0140] For example, the audio played by the audio device can be re-rendered according to the adjusted pose.
[0141] In a possible embodiment, step X30 further comprises step X31 or step X32:
[0142] Step X31, rendering the audio played by the audio device according to the adjusted pose of the virtual speaker and the personalized acoustic transfer data;
[0143] It should be noted that the adjusted pose includes a first target pose and a second target pose. For example, a first head-related transfer function HRTF corresponding to the first target pose can be obtained in the personalized acoustic transfer data, and a second head-related transfer function HRTF corresponding to the second target pose can be obtained in the personalized acoustic transfer data; the audio played by the audio device can be rendered according to the first head-related transfer function HRTF and the second head-related transfer function HRTF.
[0144] Step X32, obtaining an initial pose of the virtual speaker before receiving the pose adjustment instruction and an initial rendering parameter of the audio device for the audio; determining a pose deviation between the initial pose and the adjusted pose, adjusting the initial rendering parameter to obtain a target rendering parameter according to a sound source compensation coefficient corresponding to the pose deviation; and rendering the audio played by the audio device through the target rendering parameter.
[0145] It should be noted that the initial pose is the pose of the virtual speaker before receiving the pose adjustment instruction, and the initial pose can include a first initial pose of the first speaker and a second initial pose of the second speaker. In possible cases, the first initial pose can be the first playback pose in the virtual playback pose, and the second initial pose can be the second playback pose in the virtual playback pose, which is not limited in this embodiment.
[0146] The pose deviation is the deviation between the initial pose and the adjusted pose, and the pose deviation can include a first position deviation and a first angle deviation of the first speaker, and a second position deviation and a first angle deviation of the second speaker.
[0147] The sound source compensation coefficient can include a first compensation coefficient of the first loudspeaker and a compensation coefficient of the second loudspeaker, the first compensation coefficient can include a first position compensation coefficient and a first angle compensation coefficient, and the second compensation coefficient can include a second position compensation coefficient and a second angle compensation coefficient. The sound source compensation coefficient can be used to compensate for the parameter of the audio rendering error caused by the virtual loudspeaker due to the pose deviation. The sound source compensation coefficient can be determined based on a preset pose compensation mapping relationship, and the embodiment is not limited in this regard. The pose compensation mapping relationship includes the mapping relationship between a plurality of preset pose deviations and the respective corresponding preset compensation coefficients.
[0148] The initial rendering parameter is a parameter for the audio device to render the audio at the initial pose of the virtual loudspeaker, and the target rendering parameter is a target value for the audio device to render the audio after receiving the pose adjustment instruction. For example, the initial rendering parameter can include an initial binaural delay time, a filter coefficient, a cutoff frequency, and the like, and the embodiment is not limited in this regard.
[0149] For example, the initial pose of the virtual loudspeaker before receiving the pose adjustment instruction and the initial rendering parameter of the audio device for the audio are obtained; the pose deviation between the initial pose and the adjusted pose is calculated, the sound source compensation coefficient corresponding to the pose deviation is searched in the preset pose compensation mapping relationship, and the initial rendering parameter is adjusted to obtain the target rendering parameter; the audio played by the audio device is re-rendered through the target rendering parameter to simulate the audio played from the adjusted pose of the virtual loudspeaker, so that the target user can perceive that the audio is played from the adjusted pose, thereby facilitating the personalized listening experience.
[0150] The application also provides a control device of an audio device, please refer to Figure 6 The control device of the audio device comprises:
[0151] The monitoring module 10 is configured to monitor the head rotation angle of the target user and obtain the individual acoustic transmission data of the target user, wherein the individual acoustic transmission data comprises at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user.
[0152] The adjustment module 20 is configured to adjust the pose of the virtual loudspeaker in the audio device according to the head rotation angle to obtain a virtual playing pose of the virtual loudspeaker.
[0153] The rendering module 30 is configured to obtain a target head-related transfer function HRTF corresponding to the virtual playing pose from the individual acoustic transmission data, and render the audio played by the audio device according to the target head-related transfer function HRTF to simulate the audio played at the virtual playing pose.
[0154] The control device of the audio equipment provided in the embodiments of the present application adopts the control method of the audio equipment in the above embodiments, and aims to solve the problem of poor sound quality performance of the audio equipment caused by head rotation. Compared with the prior art, the control method of the audio equipment provided in the embodiments of the present application has the same beneficial effects as the control method of the audio equipment provided in the above embodiments, and other technical features in the control device of the audio equipment are the same as the features disclosed in the above method embodiments, which will not be repeated here.
[0155] The present application provides an audio equipment, which comprises at least one processor and a memory connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the control method of the audio equipment in the above embodiment one.
[0156] Reference will be made to the accompanying drawings Figure 7 which shows a structural schematic diagram of an audio equipment suitable for being used to implement the embodiments of the present application. The audio equipment in the embodiments of the present application can be a near-ear type audio equipment or a far-ear type audio equipment. The near-ear open type equipment includes but is not limited to AR equipment, VR equipment, smart audio glasses, neck-hung audio box, open earphone, mobile phone, tablet computer, etc. The far-ear open type equipment includes but is not limited to audio equipment of sound equipment, television, etc. Figure 7 The audio equipment shown is only an example and should not bring any limitation to the functions and use range of the embodiments of the present application.
[0157] As Figure 7As shown, the audio device can include a processing apparatus 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory 1002 or loaded from a storage apparatus 1003 into a random access memory 1004. Various programs and data required for the operation of the audio device are also stored in the random access memory 1004. The processing apparatus 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other by a bus 1005. An input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: input apparatuses 1007 including, for example, a touch panel, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output apparatuses 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage apparatus 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication apparatus 1009. The communication apparatus 1009 can allow the audio device to communicate wirelessly or by wire with other devices to exchange data. Although the audio device having various systems is shown in the figure, it should be understood that all the shown systems are not required to be implemented or possessed. More or less systems can be alternatively implemented or possessed.
[0158] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by a communication apparatus, or installed from the storage apparatus 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing apparatus 1001, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0159] The audio device provided by the present application adopts the control method of the audio device in the above-mentioned embodiments, and can solve the problem of poor sound quality performance of the audio device caused by head rotation. Compared with the prior art, the audio device provided by the present application has the same beneficial effects as the control method of the audio device provided by the above-mentioned embodiments, and other technical features in the audio device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.
[0160] It should be understood that parts of the present disclosure can be realized by hardware, software, firmware or a combination thereof. In the description of the above-mentioned embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0161] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0162] The embodiment provides a computer readable storage medium having computer readable program instructions stored thereon, and the computer readable program instructions are used for executing the control method of the audio device in the above embodiment one.
[0163] The computer readable storage medium provided by the embodiment of the present application may, for example, be a U disk, but is not limited to an electric, magnetic, optical, electromagnetic, infrared, or semiconductor device, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electric connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable EPROM (Electrical Programmable Read Only Memory, read-only memory) or a flash memory, an optical fiber, a portable compact disk CD-ROM (compact disc read-only memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiment, the computer readable storage medium can be any tangible medium containing or storing a program, which can be used by or in conjunction with an instruction execution device, device or apparatus. The program code contained on the computer readable storage medium can be transmitted by any suitable medium, including but not limited to: electric wire, optical cable, RF (Radio Frequency, radio frequency), etc., or any suitable combination of the above.
[0164] The above computer readable storage medium can be contained in the audio device; or can exist separately without being assembled into the audio device.
[0165] The computer readable storage medium described above carries one or more programs, when the one or more programs are executed by the audio device, cause the audio device to: monitor a head rotation angle of a target user, and acquire individual acoustic transmission data of the target user, wherein the individual acoustic transmission data comprises at least one head-related transfer function (HRTF) trained based on personal auditory physiological data of the target user; adjust a pose of a virtual loudspeaker in the audio device according to the head rotation angle, to obtain a virtual playing pose of the virtual loudspeaker; acquire a target head-related transfer function (HRTF) corresponding to the virtual playing pose from the individual acoustic transmission data, and render audio played by the audio device according to the target head-related transfer function (HRTF), to simulate playing of the audio at the virtual playing pose. The application solves the problem of poor sound quality performance of the audio device caused by head rotation.
[0166] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0167] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block in the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by special-purpose hardware-based devices, which perform the specified functions or operations, or combinations of hardware and computer instructions.
[0168] The modules described in the embodiments of the present disclosure can be implemented in the form of software or in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself.
[0169] The computer readable storage medium provided by the embodiments of the present application stores computer readable program instructions for executing the control method of the audio device, and is intended to solve the problem of poor sound quality performance of the audio device caused by head rotation. Compared with the prior art, the computer readable storage medium provided by the embodiments of the present application has the same beneficial effects as the control method of the audio device provided by the above embodiments, and will not be repeated here.
[0170] The embodiments of the present application also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the control method of the audio device are implemented.
[0171] The computer program product provided by the embodiments of the present application is intended to solve the problem of poor sound quality performance of the audio device caused by head rotation. Compared with the prior art, the computer program product provided by the embodiments of the present application has the same beneficial effects as the control method of the audio device provided by the above embodiments, and will not be repeated here.
[0172] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation based on the content of the present application specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent processing scope of the present application.
Claims
1. A control method of an audio device, characterized by, The method comprises: monitoring the head rotation angle of the target user and obtaining the individual acoustic transmission data of the target user, wherein the individual acoustic transmission data comprises at least one head-related transfer function HRTF trained based on the personal auditory physiological data of the target user, the head rotation angle is the rotation angle of the target user relative to a preset zero position direction, and the preset zero position direction is the direction of the target user's head facing forward; adjusting the pose of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playing pose of the virtual speaker; obtaining the target head-related transfer function HRTF corresponding to the virtual playing pose in the individual acoustic transmission data, and rendering the audio played by the audio device according to the target head-related transfer function HRTF to simulate playing the audio at the virtual playing pose; wherein the head rotation angle comprises a horizontal rotation angle, the virtual speaker comprises a first speaker and a second speaker, the virtual playing pose comprises a first playing pose of the first speaker and a second playing pose of the second speaker; and the step of adjusting the pose of the virtual speaker in the audio device according to the head rotation angle to obtain the virtual playing pose of the virtual speaker comprises: determining the target horizontal included angle corresponding to the horizontal rotation angle in the mapping relationship between the preset horizontal angle and the preset horizontal included angle, and equally dividing the target horizontal included angle to obtain a horizontal symmetric adjustment angle; adjusting the angles of the first speaker and the second speaker in the horizontal direction according to the horizontal symmetric adjustment angle to obtain the first playing pose of the first speaker and the second playing pose of the second speaker, so that the included angle of the first speaker and the second speaker in the horizontal direction is the target horizontal included angle; wherein the higher the absolute value of the head rotation angle, the smaller the target horizontal included angle, and the lower the absolute value of the head rotation angle, the larger the target horizontal included angle; wherein the horizontal symmetric adjustment angle is the absolute value of the included angle of the first speaker relative to a preset central axis, and is also the absolute value of the included angle of the second speaker relative to the preset central axis, and the preset central axis is a central axis of the head facing forward and horizontal to the ground. 2.The control method of an audio device of claim 1, wherein, The step of obtaining the individual acoustic transmission data of the target user comprises: training a preset model by using a obtained group auditory training data set to obtain a trained general model, wherein the group auditory training data set comprises multiple human auditory physiological data and respective corresponding HRTF labels; obtaining the personal auditory physiological data of the target user; obtaining a general decoder from the general model, and performing individual training on the general decoder by using the personal auditory physiological data to obtain an individual decoder, and generating the individual acoustic transmission data corresponding to the personal auditory physiological data based on the individual decoder. 3.The control method of an audio device of claim 2, wherein, The preset model comprises a preset encoder and a preset decoder; the step of training the preset model by using the obtained group auditory training data set to obtain a trained general model comprises: obtaining target human auditory physiological data from the group auditory training data set, inputting the target human auditory physiological data into the preset model, encoding the target human auditory physiological data by using the preset encoder in the preset model, inputting the encoded result into the preset decoder, and outputting an HRTF training result by using the preset decoder; determining a training loss between the HRTF training result and a target HRTF label of the target human auditory physiological data, and in a case where the training loss is greater than a preset loss threshold, re-obtaining new target human auditory physiological data from the group auditory training data set, and returning to the step of inputting the target human auditory physiological data into the preset model until the training loss is less than or equal to the preset loss threshold, thereby obtaining a trained general model. 4.The control method of an audio device of claim 2, wherein, The step of obtaining the personal auditory physiological data of the target user comprises: obtaining head structure parameters of the target user, the head structure parameters comprising auricle height, auricle width, ear canal inclination angle, head width, and shoulder-neck parameter; constructing a head geometric model of the target user by using the head structure parameters, and extracting personal auditory physiological data from the head geometric model. 5.The control method of an audio device of claim 2, wherein, The step of obtaining the personal auditory physiological data of the target user comprises: obtaining three-dimensional head scan data of the user, and determining personal auditory physiological data from the three-dimensional head scan data; and / or obtaining head measurement parameters of the user, obtaining target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters from a preset auditory physiological database, and obtaining personal auditory physiological data composed of a plurality of target sub-auditory physiological data, wherein the plurality of sub-measurement parameters included in the head measurement parameters are respectively external ear structure parameters, head-neck connection parameters, and head body parameters, and the preset auditory physiological database comprises preset sub-auditory physiological data corresponding to a plurality of preset sub-measurement parameters. 6.The control method of an audio device of claim 1, wherein, The head rotation angle further comprises a vertical pitch angle; The step of adjusting the pose of the virtual loudspeaker in the audio device according to the head rotation angle to obtain a virtual playing pose of the virtual loudspeaker comprises: determining a target vertical included angle corresponding to the vertical pitch angle in a mapping relationship between a preset pitch angle and a preset vertical included angle; determining a target horizontal included angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal included angle; adjusting the angles of a first loudspeaker in the virtual loudspeaker and a second loudspeaker in the virtual loudspeaker in the vertical direction according to the target vertical included angle, and adjusting the angles of the first loudspeaker and the second loudspeaker in the horizontal direction according to the target horizontal included angle, thereby obtaining a first playing pose of the first loudspeaker and a second playing pose of the second loudspeaker; and determining a target vertical included angle corresponding to the vertical pitch angle in a mapping relationship between a preset pitch angle and a preset vertical included angle; determining a target horizontal included angle corresponding to the horizontal rotation angle in a mapping relationship between a preset horizontal angle and a preset horizontal included angle; adjusting the angles of a first loudspeaker in the virtual loudspeaker and a second loudspeaker in the virtual loudspeaker in the vertical direction according to the target vertical included angle, and adjusting the angles of the first loudspeaker and the second loudspeaker in the horizontal direction according to the target horizontal included angle, thereby obtaining a first playing pose of the first loudspeaker and a second playing pose of the second loudspeaker. The first playing pose and the second playing pose have a target vertical included angle in a vertical direction. 7.The control method of an audio device of claim 1, wherein, The control method of the audio device further includes: receiving a pose adjustment instruction of the virtual speaker, determining a position adjustment parameter and an angle adjustment parameter of the virtual speaker from the pose adjustment instruction; adjusting the pose of the virtual speaker according to the position adjustment parameter and the angle adjustment parameter to obtain an adjusted pose of the virtual speaker in the audio device; rendering the audio played by the audio device based on the adjusted pose. 8.The control method of an audio device of claim 7, wherein, The step of rendering the audio played by the audio device based on the adjusted pose includes: rendering the audio played by the audio device according to the adjusted pose of the virtual speaker and individual acoustic transmission data; or obtaining an initial pose of the virtual speaker before the pose adjustment instruction is received and an initial rendering parameter of the audio device for the audio; determining a pose deviation between the initial pose and the adjusted pose, adjusting the initial rendering parameter to obtain a target rendering parameter according to a sound source compensation coefficient corresponding to the pose deviation; and rendering the audio played by the audio device through the target rendering parameter.
9. An audio device, comprising: The audio device includes at least one processor and a memory in communication with the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the control method of the audio device as claimed in any one of claims 1 to 8.
10. A storage medium, characterized by The storage medium is a computer readable storage medium, and the computer readable storage medium stores a program for implementing the control method of the audio device. The program for implementing the control method of the audio device is executed by the processor to implement the steps of the control method of the audio device as claimed in any one of claims 1 to 8.
Citation Information
Patent Citations
Audio processing method, audio playing device and computer readable storage medium
CN117956372A
Audio processing device and method, augmented reality device, equipment and storage medium
CN120188497A