Sound field expansion methods, audio devices, and computer-readable storage media
By acquiring the target transfer function and personalized acoustic transmission data to render audio and performing crosstalk cancellation, the problems of poor sound field effect and sound crosstalk in sound field extension are solved, achieving a high-quality sound field extension that is more in line with the physiological characteristics of users.
Patent Information
- Application Number
- CN202511264085.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-09-05
AI Technical Summary
In existing technologies, sound field extension functions are mainly implemented through the general HRTF algorithm, which results in poor sound field effects and sound crosstalk problems, especially in open audio devices.
The target transfer function between the test audio device at the extended target location and the user's two ears is obtained. Combined with the target user's personalized acoustic transmission data, including the head-related transfer function (HRTF) generated based on personal auditory physiological data, the original audio is rendered and crosstalk cancellation is performed to generate the target audio.
It enhances the naturalness and immersion of the sound field expansion effect, improves the positioning accuracy and spatial sense of the sound, and reduces sound crosstalk to ensure sound quality.
Smart Images

Figure CN120769218B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to a sound field expansion method, an audio device, and a computer-readable storage medium. Background Technology
[0002] Sound field extension refers to the acoustic phenomenon where the perceived sound field is wider than the actual location of the speaker. Sound field extension is similar to a virtual speaker that can extend the sound source to a wider location than the actual location of the speaker. In other words, the sound played by the sound source sounds to the human ear as if the sound is coming from a virtual speaker in a wider location.
[0003] In the field of audio processing technology, most actual audio signals are dual-channel stereo signals. Sound field expansion technology is based on dual-channel stereo, without adding channels or speakers. By processing the signal, it makes the listener feel that the sound comes from multiple directions, producing a simulated stereo sound field.
[0004] However, current sound field extension functions (i.e., virtual surround sound functions) are mainly implemented through the Head Related Transfer Function (HRTF) algorithm. When using a general HRTF for sound field extension, the sound field effect may not meet expectations because the general HRTF does not match the individual's head. Furthermore, crosstalk may also exist. Crosstalk refers to the phenomenon where a sound signal from one channel travels through the air and is received by another ear. In audio devices, where sound travels directly through the air, crosstalk is particularly prominent.
[0005] Therefore, how to effectively expand the sound field while reducing crosstalk and ensuring sound quality has become a pressing technical challenge. Summary of the Invention
[0006] The main objective of this application is to provide a sound field expansion method, audio device, and computer-readable storage medium, aiming to solve the technical problem of how to effectively expand the sound field while reducing crosstalk and ensuring sound quality.
[0007] To achieve the above objectives, this application provides a sound field expansion method, which includes the following steps:
[0008] Obtain the target transfer function between the test audio device at the extended target location and the user's two ears, and obtain the target user's personalized acoustic transmission data, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) generated based on the target user's personal auditory physiological data;
[0009] Personalized audio is obtained by rendering the original audio based on the personalized acoustic transmission data.
[0010] The personalized audio is subjected to crosstalk cancellation processing according to the target transfer function to obtain the target audio, and then the target audio is played.
[0011] In one embodiment, the step of obtaining the target transfer function between the test audio device at the extended target location and the user's two ears includes:
[0012] Obtain the artificial head transfer function and the free field transfer function;
[0013] Inverting the free field transfer function yields the inverse free field transfer function;
[0014] Multiplying the artificial head transfer function by the inverse free field transfer function yields the target transfer function between the test audio device at the extended target location and the user's ears.
[0015] In one embodiment, the step of obtaining the artificial head transfer function and the free field transfer function includes:
[0016] When the test audio device is placed at the extended target location and outputs a sound signal, the transfer function of the artificial head is measured through a preset microphone in the ear canal of the preset artificial head; and,
[0017] When the preset artificial head is removed and the test audio device outputs a sound signal, the free field transfer function is measured by a preset microphone placed at the left and right ear positions before the preset artificial head was removed.
[0018] In one embodiment, the step of performing crosstalk cancellation processing on the personalized audio according to the target transfer function to obtain the target audio includes:
[0019] Invert the target transfer function to obtain the target inverse transfer function;
[0020] The target audio is obtained by multiplying the personalized audio by the target inverse transfer function.
[0021] In one embodiment, the step of acquiring the target user's personalized acoustic transmission data includes:
[0022] By acquiring the group auditory training dataset, a pre-set model is trained to obtain a trained general model. The group auditory training dataset includes multiple human auditory physiological data and their respective corresponding HRTF labels. The general model includes a general encoder and a general decoder.
[0023] Acquire the target user's personal auditory physiological data;
[0024] A general decoder is obtained from the general model. The general decoder is then trained using the individual's auditory physiological data to obtain a personalized decoder. Personalized acoustic transmission data corresponding to the individual's auditory physiological data is then generated based on the personalized decoder.
[0025] In one embodiment, before the step of rendering the original audio based on the personalized acoustic transmission data to obtain personalized audio, the method further includes:
[0026] Obtain the amplitude spectrum and phase spectrum of the personalized acoustic transmission data, and obtain the preset covert sound threshold;
[0027] The amplitude spectrum is adjusted based on the preset covert sound threshold to obtain the adjusted amplitude spectrum, wherein the adjustment of the amplitude spectrum includes reducing the amplitude value of the target frequency point, and the target frequency point is the frequency point in the amplitude spectrum whose amplitude value is less than the preset covert sound threshold;
[0028] The adjusted amplitude spectrum and the adjusted phase spectrum are combined to obtain the adjusted personalized acoustic transmission data;
[0029] Based on the adjusted personalized acoustic transmission data, the step of rendering the original audio according to the personalized acoustic transmission data to obtain personalized audio is performed.
[0030] In one embodiment, before the step of obtaining the preset covert sound threshold, the method further includes:
[0031] Obtain the audio parameters of the original audio, wherein the audio parameters include the power density spectrum and the sound pressure level;
[0032] Based on the audio parameters, at least one masked audio frequency point in the original audio is identified, and the individual sound pressure level of each masked audio frequency point is obtained;
[0033] Determine the minimum value between the individual sound pressure level and the absolute hearing threshold, and set the minimum value as the preset covert sound threshold.
[0034] In one embodiment, after the step of obtaining the target transfer function between the test audio device at the extended target location and the user's two ears, the method further includes:
[0035] Obtain the ambient acoustic parameters of the target user's current environment, wherein the ambient acoustic parameters include reflected sound intensity, reverberation time, and dominant reflection direction;
[0036] Generate corresponding environmental compensation coefficients based on the aforementioned environmental acoustic parameters;
[0037] Multiply the environmental compensation coefficient by the target transfer function to obtain the compensated target transfer function;
[0038] Based on the compensated target transfer function, the step of performing crosstalk cancellation processing on the personalized audio according to the target transfer function to obtain the target audio is performed.
[0039] In addition, to achieve the above objectives, this application also provides an audio device, the audio device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the sound field expansion method as described above.
[0040] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of the sound field expansion method as described above.
[0041] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the sound field expansion method described above.
[0042] One or more technical solutions proposed in this application have at least the following technical effects:
[0043] This application addresses the technical challenges of poor sound field effects and crosstalk issues in existing technologies by acquiring the target transfer function between a test audio device at the target location and the user's two ears, and combining this with the target user's personalized acoustic transmission data. Specifically, the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) generated based on the target user's personal auditory physiological data. By using these personalized HRTFs to render the original audio, the generated personalized audio better matches the user's auditory perception, thereby improving the naturalness and immersion of the sound field expansion effect. Furthermore, crosstalk cancellation processing is applied to the personalized audio based on the target transfer function, reducing the degree to which the sound signal from one channel is received by the other ear, thereby enhancing the sound localization accuracy and spatial sense. Ultimately, the output target audio not only better matches the user's personalized physiological characteristics in terms of sound field expansion effect, effectively ensuring sound quality, but also undergoes crosstalk cancellation, achieving the goal of effectively expanding the sound field while reducing crosstalk and ensuring sound quality. Attached Figure Description
[0044] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating the first embodiment of the sound field expansion method of this application;
[0047] Figure 2 This is a schematic diagram illustrating an application scenario of an embodiment of the sound field expansion method of this application;
[0048] Figure 3 This is a schematic diagram of the crosstalk elimination process involved in the first embodiment of the sound field expansion method of this application;
[0049] Figure 4 This is a schematic diagram of a scenario involving the generation of personalized acoustic transmission data according to an embodiment of the sound field expansion method of this application;
[0050] Figure 5 This is a schematic diagram of the process for acquiring personal auditory physiological data according to an embodiment of the sound field expansion method of this application;
[0051] Figure 6 This is a schematic diagram of the hardware operating environment of the sound field expansion method device in the embodiments of this application.
[0052] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0053] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] In digital signal processing, most actual audio signals are two-channel stereo signals. Stereo extension technology, based on two-channel stereo, without adding channels or speakers, processes the signal to make the listener perceive sound coming from multiple directions, creating a simulated stereo sound field. Stereo extension technology has become an indispensable technology. It is mainly used for far-field sound sources, such as scenarios using speakers. In recent years, the use of near-ear open-back audio devices such as VR (Virtual Reality) and AR (Augmented Reality) has become increasingly widespread, and the demand for sound field extension functions in near-ear open-back audio devices has also gradually increased. Currently, stereo extension functions are mainly implemented through the Head Related Transfer Function (HRTF) algorithm. However, using a generic HRTF for sound field extension may result in a mismatch between the generic HRTF and the individual's head, leading to unsatisfactory sound field effects. Furthermore, for audio devices, crosstalk may occur, causing the sound image to narrow. Therefore, how to effectively extend the sound field while being suitable for open-back devices, reducing crosstalk, and ensuring sound quality has become particularly important.
[0055] Based on this, the main solution of this application is: to obtain the target transfer function between the test audio device at the extended target location and the user's two ears, and to obtain the target user's personalized acoustic transmission data, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) generated based on the target user's personal auditory physiological data; to render the original audio according to the personalized acoustic transmission data to obtain personalized audio; to perform crosstalk cancellation processing on the personalized audio according to the target transfer function to obtain target audio, and to play the target audio.
[0056] This application addresses the technical challenges of poor sound field effects and crosstalk issues in existing technologies by acquiring the target transfer function between a test audio device at the target location and the user's two ears, and combining this with the target user's personalized acoustic transmission data. Specifically, the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) generated based on the target user's personal auditory physiological data. By using these personalized HRTFs to render the original audio, the generated personalized audio better matches the user's auditory perception, thereby improving the naturalness and immersion of the sound field expansion effect. Furthermore, crosstalk cancellation processing is applied to the personalized audio based on the target transfer function, reducing the degree to which the sound signal from one channel is received by the other ear, thereby enhancing the sound localization accuracy and spatial sense. Ultimately, the output target audio not only better matches the user's personalized physiological characteristics in terms of sound field expansion effect, effectively ensuring sound quality, but also undergoes crosstalk cancellation, achieving the goal of effectively expanding the sound field while reducing crosstalk and ensuring sound quality.
[0057] It should be noted that the implementing entity of the sound field expansion method embodiments of this application can be an audio device capable of achieving the above-mentioned functions, such as AR helmets, VR helmets, smart audio glasses, neckband speakers, open-back headphones, and other near-ear open-back audio devices, as well as speakers, televisions, and other far-ear open-back audio devices. The embodiments of the sound field expansion method of this application do not impose specific limitations on this. Exemplarily, the embodiments of this application are described and illustrated using a near-ear open-back audio device as the implementing entity.
[0058] Based on this, this application proposes a sound field expansion method according to a first embodiment, referring to... Figure 1 As shown, the sound field expansion method includes the following steps S10~S30:
[0059] Step S10: Obtain the target transfer function between the test audio device at the extended target location and the user's two ears, and obtain the target user's personalized acoustic transmission data, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) generated based on the target user's personal auditory physiological data;
[0060] The test audio device can specifically be a device with a sound playback function; further, the test audio device can be a near-ear open-back audio device as the execution subject. For example, the various embodiments of this application are described and illustrated using the test audio device as a speaker. The extended target position is a virtual position where the user expects to perceive the sound source, used to simulate the sound source position in a wider sound effect scene.
[0061] Near-ear open-back audio devices, compared to speaker systems, have their speakers positioned much closer to the listener's ear. Furthermore, these devices are typically all-in-one units, meaning the distance between the playback device and the listener's ear is almost fixed. This means that, generally, the sound perceived by the human ear from near-ear open-back audio devices is near-field sound. (This can be combined with...) Figure 2 To understand, Figure 2 This is a schematic diagram of an application scenario provided in this embodiment. It is assumed that the relative position of the user's head to the device's speaker is as follows: Figure 2 As shown, when using near-ear open-back audio devices, even if the playback device position can be finely adjusted, it cannot provide the listening experience that the speaker position (shown as dotted in the diagram) can produce. Users can only experience the near-field listening sensation from the actual speaker. When users want to experience a wider sound effect, for example, when they want to hear from... Figure 2 The sound emitted from the speaker position indicated by the dashed line can be set as the target location for expansion. The target transfer function between the speaker at the target location and the user's ears represents the transfer function of sound from the speaker at that location to the human ear.
[0062] It should be noted that all transfer functions mentioned in this embodiment are acoustic transfer functions. An acoustic transfer function is the transfer function from a sound source to the reproduction region. The transfer function refers to the ratio of the Laplace transform (or z-transform) of the response (i.e., output) of a linear system under zero initial conditions to the Laplace transform of the excitation (i.e., input). In this embodiment, the target transfer function between the test audio device and the user's ears is the transfer function from the output sound source (i.e., speaker or loudspeaker) of the test audio device to the user's ears, used to reflect the changes in the output signal of the test audio device during its transmission to the user's ears.
[0063] Personalized acoustic transfer data can include at least one Head-Related Transfer Function (HRTF). An HRTF is a frequency-domain filtering function that describes the reflection, diffraction, and scattering effects of sound waves on physiological structures such as the head, auricle, and shoulders as sound propagates from a free field to both ears. HRTFs can be used to simulate the changes in sound waves propagating from the sound source location to the ear, giving the virtual sound source realistic spatial attributes. For the same target user, multiple HRTFs can exist; for example, personalized acoustic transfer data can include a first HRTF for the left ear and a second HRTF for the right ear.
[0064] Step S20: Render the original audio based on the personalized acoustic transmission data to obtain personalized audio;
[0065] The raw audio can be audio that has not undergone any acoustic processing or personalized acoustic optimization, such as audio input to an open-back audio device or local audio stored in an open-back audio device.
[0066] Since personalized acoustic transmission data includes at least one head-related transfer function (HRTF) generated based on the target user's personal auditory physiological data, each HRTF in the personalized acoustic transmission data can reflect the target user's personal auditory physiological characteristics. Therefore, this embodiment renders the original audio using personalized acoustic transmission data, making the rendered audio more consistent with the target user's auditory physiological characteristics, thereby improving the sound quality performance of the audio device.
[0067] Step S30: Perform crosstalk cancellation processing on the personalized audio according to the target transfer function to obtain the target audio, and play the target audio.
[0068] Understandably, since the speakers or horns of near-ear open-back audio devices are not ideal sound sources, and since these speakers or horns cannot be directly placed into the user's ear canals as playback devices, crosstalk often occurs during the transmission of the target audio. To avoid this problem, crosstalk cancellation processing is performed on the personal audio before the target audio is generated. This process cancels out the crosstalk problems that occur during the transmission of the target audio after playback. In other words, the target audio is obtained after crosstalk cancellation processing of the personal audio. It can cancel out the influence of the playback device itself and the influence of the user's head on the sound signal during sound transmission, while retaining the influence of the user's head on the sound transmission result. When the target audio enters the user's ears, the user can perceive that the source of the sound signal is an extended target location at a distance.
[0069] Understandably, since the target audio has undergone crosstalk cancellation processing based on the crosstalk cancellation algorithm provided in this embodiment before being output through the speaker, the influence of the near-ear open audio device itself and the user's head on the sound transmission result is eliminated. Therefore, when the target audio enters the user's ears, it can achieve the same effect as the personalized audio. If the personalized audio is the left channel, the target audio transmits the sound signal only to the left ear; if the personalized audio is the right channel, the target audio transmits the sound signal only to the right ear; if the personalized audio is stereo, the target audio outputs the left and right channel sound signals to the left and right ears respectively.
[0070] This embodiment solves the technical problems of poor sound field effects and sound crosstalk issues caused by the mismatch between the general HRTF and individual physiological characteristics in existing technologies by obtaining the target transfer function between the test audio device at the target location and the user's two ears, and combining it with the target user's personalized acoustic transmission data. Specifically, the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) generated based on the target user's personal auditory physiological data. By using these personalized HRTFs to render the original audio, the generated personalized audio can better match the user's auditory perception, thereby improving the naturalness and immersion of the sound field expansion effect. Furthermore, crosstalk cancellation processing is performed on the personalized audio according to the target transfer function to reduce the degree to which the sound signal from one channel is received by the other ear, thereby enhancing the sound positioning accuracy and spatial sense. Ultimately, the output target audio not only better matches the user's personalized physiological characteristics in terms of sound field expansion effect, effectively ensuring sound quality, but also undergoes crosstalk cancellation, achieving the goal of effectively expanding the sound field while reducing crosstalk and ensuring sound quality.
[0071] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, the step of obtaining the target transfer function between the test audio device at the extended target location and the user's two ears includes:
[0072] Step A10: Obtain the artificial head transfer function and the free field transfer function;
[0073] It should be noted that the target transfer function can be understood as the influence of the user's head contour on the sound signal transmission result. This embodiment derives two different acoustic transfer functions based on two different sound transmission scenarios, and then calculates the target transfer function. The artificial head transfer function is calculated when the test audio device is placed at the extended target position (i.e.,...). Figure 2 The acoustic transfer function (located at the position of the speaker in the middle dotted line) is measured by a preset microphone in the ear canal of the preset artificial head when the speaker outputs a sound signal. This function includes the influence of the test audio device and the preset artificial head on the sound transmission result. The free field transfer function is the acoustic transfer function measured by preset microphones placed at the left and right ear positions before the preset artificial head was removed, when the preset artificial head is removed and the test audio device outputs a sound signal. This function includes the influence of the test audio device on the sound transmission result.
[0074] In one possible implementation, the step of obtaining the artificial head transfer function and the free field transfer function includes:
[0075] Step A201: When the test audio device is placed at the extended target location and the test audio device outputs a sound signal, the transfer function of the artificial head is measured through a preset microphone in the ear canal of the preset artificial head; and,
[0076] Step A202: When the preset artificial head is removed and the test audio device outputs a sound signal, the free field transfer function is measured by the preset microphones placed at the left and right ear positions before the preset artificial head was removed.
[0077] It should be noted that in this embodiment, the preset artificial head is an auxiliary device constructed to simulate the user's head for assisting in the measurement of the acoustic transfer function. It can simulate the scenario of the user receiving sound signals emitted from the test audio. The preset artificial head is equipped with left and right ears and ear canals, and microphones for receiving sound signals can be pre-placed in the ear canals.
[0078] As an example, combined Figure 2 As shown in the application scenario, the test audio device is placed at the virtual speaker's designated location, i.e. Figure 2 At the location of the speaker in the middle dashed line, the acoustic transfer function from the sound source to the two ears of the preset artificial head is measured using two preset microphones in the ear canal of the preset artificial head, and this acoustic transfer function is denoted as H1.
[0079] As an example, combined Figure 2 As shown in the application scenario, first, two microphones identical to those in the preset artificial head ear canals in step A201 above are placed at the left and right ear positions of the preset artificial head. Then, the preset artificial head is removed, and the test audio device is still placed at the virtual speaker setting position. Figure 2 At the location of the speaker in the middle dashed line, the acoustic transfer function of the sound source when it is working in a free field is measured using two microphones that are not affected by the preset artificial head, and this acoustic transfer function is denoted as H2.
[0080] Step A20: Perform the inverse operation on the free field transfer function to obtain the inverse free field transfer function;
[0081] Step A30: Multiply the artificial head transfer function with the free field inverse transfer function to obtain the target transfer function between the test audio device at the extended target position and the user's ears.
[0082] In this embodiment, the free-field transfer function H2, which includes the influence of the test audio device on the sound transmission result, is first inverted to obtain the inverse free-field transfer function, denoted as H2'. Then, the artificial head transfer function H1, which includes the influence of the playback device and the preset artificial head on the sound transmission result, is multiplied by H2' to obtain the target transfer function H. It should be noted that H2' obtained after the inversion operation can eliminate the influence of the playback device on the sound transmission result. Multiplying it by H1 eliminates the part of H1 that includes the influence of the test audio device on the sound transmission result, retaining the influence of the preset artificial head on the sound transmission result as the target transfer function H.
[0083] In one possible implementation, the step of performing crosstalk cancellation processing on the personalized audio according to the target transfer function to obtain the target audio includes:
[0084] Step B10: Perform the inverse operation on the target transfer function to obtain the target inverse transfer function;
[0085] Step B20: Multiply the personalized audio with the target inverse transfer function to obtain the target audio.
[0086] As can be seen from the above steps, the target transfer function H represents the influence of the head contour on the sound transmission result. It should be understood that the target inverse transfer function obtained after inverting H is equivalent to an identity matrix, which represents the elimination of the influence of the head contour on the sound transmission result. The target audio obtained after applying it to the personalized audio can obviously cancel the influence of the head contour on the sound transmission result during the sound signal transmission, so that the sound signal received by the user's two ears can be consistent with the target audio, and it is a wider far-field sound effect from the virtual loudspeaker (i.e., the speaker at the extended target position), thus obtaining the extended sound field after crosstalk cancellation.
[0087] As an example, combined Figure 2 As shown in the application scenario, given the original audio X from a near-ear open-back audio device, it is processed by a personalized HRFT algorithm module to generate personalized acoustic transmission data M. This data is then processed by a binaural synthesis algorithm module, which renders X using the personalized acoustic transmission data M to obtain personalized audio, denoted as XM, achieving the effect of expanding the sound field. After further processing by a crosstalk cancellation algorithm module, the audio is output through SPK (speaker) to play the target audio into the listener's ear. (Refer to...) Figure 3 As shown, the basic implementation flow of the crosstalk cancellation algorithm module is as follows:
[0088] Step S11: Obtain the rendered personalized audio XM;
[0089] Step S12: Place the test audio device at the target position of the virtual speaker after the sound field is expanded (i.e., the expanded target position), and use the microphone in the preset artificial head ear canal to measure the transfer function from the sound source to the human ear, which is recorded as H1;
[0090] Step S13: Place two more microphones at the left and right ear positions of the preset artificial head, remove the preset artificial head, and measure the transfer function of the device when it is working in the free field, denoted as H2.
[0091] Step S14: Perform the inverse operation on H2, denoted as H2';
[0092] Step S15: Calculate the target transfer function H, specifically, H = H1H2', and inverse H, denoted as B;
[0093] Step S16: Generate the target audio Y. Specifically, Y = XMBH, and the target audio Y is the audio after crosstalk is eliminated.
[0094] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. Based on this, the step of obtaining the personalized acoustic transmission data of the target user includes:
[0095] Step C10: Using the obtained group auditory training dataset, train a preset model to obtain a trained general model. The group auditory training dataset includes multiple human auditory physiological data and their respective corresponding HRTF labels. The general model includes a general encoder and a general decoder.
[0096] It should be noted that the group auditory training dataset includes multiple human auditory physiological data sets. Specifically, a deep learning model is trained using a large amount of existing HRTF data (i.e., open-source HRTF data), resulting in the group auditory training dataset. This existing HRTF data can contain HRTF data from a large number of different individuals, covering a wide range of ages, ear shapes, and head shapes. In other words, multiple human auditory physiological data sets originate from different individuals. The preset model can include a preset encoder and a preset decoder, with each human auditory physiological data set corresponding to an HRTF label. The group auditory training dataset can be predetermined, and each human auditory physiological data set can reflect the auditory physiological characteristics of a human body. For example, human auditory physiological data can include data on ear canal resonant frequencies, the angle between the tragus and cymba conchae, and the complexity of auricular folds, etc. This embodiment does not specifically limit this. The method of acquiring human auditory physiological data can be the same as the method of acquiring the target user's individual auditory physiological data. Human auditory physiological data is data that the preset encoder can process. The general model includes a general decoder and a general encoder. After the preset model is trained, a general model consisting of a general decoder and a general encoder can be obtained.
[0097] For example, a group auditory training dataset is obtained, and a preset model is trained using this dataset to obtain a general model consisting of a general decoder and a general encoder. This general model can also be used to generate a general HRTF, which can simulate the changes in sound waves propagating from a sound source to the human ear. However, a general HRTF is difficult to adapt to individual differences in user hearing; therefore, rendering audio using a general HRTF still results in poor sound quality performance from audio devices. Therefore, in this embodiment, after determining the general model, the general decoder in the general model is fine-tuned to obtain a personalized decoder, so that personalized acoustic transmission data for the target user can be generated subsequently using the personalized decoder.
[0098] Step C20: Obtain the target user's personal auditory physiological data;
[0099] It should be noted that personal auditory physiological data can reflect the individual auditory physiological characteristics of the target user. Personal auditory physiological data can be obtained through 3D head scanning, from measurements of the target user's head, or through geometric modeling of the user's head. This embodiment does not impose specific limitations on these methods.
[0100] Step C30: Obtain a general decoder from the general model, perform personalized training on the general decoder using the personal auditory physiological data to obtain a personalized decoder, and generate personalized acoustic transmission data corresponding to the personal auditory physiological data based on the personalized decoder.
[0101] It should be noted that unsupervised training of the general decoder can be performed using individual auditory physiological data to personalize the general decoder. Since the general decoder has already learned the general mapping pattern between the physiological characteristics of the group and HRTF, unsupervised learning can be performed using individual auditory physiological data to make targeted adjustments to the general decoder. For example, the network parameters of the general decoder can be adjusted to obtain a personalized decoder. Here, the network parameters of the general decoder refer to the weights and biases of neurons in the general decoder.
[0102] For example, a general decoder is obtained from a general model. Using personal auditory physiological data, the general decoder is trained unsupervised to obtain a personalized decoder. Personal auditory physiological data can be input into a general encoder or a pre-set encoder to encode the data. The encoded result is then input into the personalized decoder, which outputs personalized acoustic transmission data. This personalized acoustic transmission data can include left-ear and right-ear acoustic transmission data. Both can include their corresponding head-related transfer functions (HRTFs). The left-ear and right-ear acoustic transmission data can be symmetrical or asymmetrical, adapted to the personalized characteristics of the target user's left and right ears.
[0103] Specifically, in other embodiments, the acoustic transmission data for the left and right ears can also be asymmetrical. For example, left and right ear personal physiological data can be obtained from individual auditory physiological data. The left ear personal physiological data is input into a universal encoder or a preset encoder to encode the data. The encoded result is then input into a personalized decoder, which outputs the left ear acoustic transmission data. Similarly, the right ear personal physiological data is input into a universal encoder or a preset encoder to encode the data. The encoded result is then input into a personalized decoder, which outputs the right ear acoustic transmission data. This improves the accuracy of personalized acoustic transmission data and can accommodate the differences between the left and right ears of the same user.
[0104] This embodiment first trains a general decoder, which facilitates the targeted adjustment of the network parameters of the general decoder based on individual auditory physiological data to obtain a personalized decoder. This personalized decoder is then combined to generate personalized acoustic transmission data, thereby improving the accuracy of audio rendering. The rendered audio conforms to the auditory physiological characteristics of the target user, providing the user with a good sound quality experience.
[0105] To better understand this embodiment, please refer to Figure 4 , Figure 4 The illustration shows the scenario of the personalization stage and the training stage of the general model in this embodiment. During the training stage of the general model, the group auditory training dataset 10 can be input into the preset encoder 20. The output of the preset encoder 20 is then input into the preset decoder 30. The preset decoder 30 can output the HRTF training result 40. The preset encoder 20 and preset decoder 30 can be trained until the trained general decoder is obtained. Personalized training of the general decoder can be performed using individual auditory physiological data 500 to obtain a personalized decoder 60.
[0106] In the personalization stage, acoustic measurements can be performed to obtain personal auditory physiological data 500. This personal auditory physiological data 500 can be used to train a general decoder, resulting in a personalized decoder 60. The personal auditory physiological data 500 is then input into the personalized decoder 60, which outputs personalized acoustic transmission data 70. Figure 4 The process of personalized training of a general decoder using 500 individual auditory physiological data is not shown in the figure.
[0107] In a feasible embodiment, step C10 further includes steps C101 to C102:
[0108] Step C101: Obtain target human auditory physiological data from the group auditory training dataset, input the target human auditory physiological data into the preset model, encode the target human auditory physiological data through the preset encoder in the preset model, input the encoding result into the preset decoder, and output the HRTF training result through the preset decoder.
[0109] Step C102: Determine the training loss between the HRTF training results and the target HRTF labels of the target human auditory physiological data. If the training loss is greater than the preset loss threshold, obtain new target human auditory physiological data from the group auditory training dataset and return to the step of inputting the target human auditory physiological data into the preset model until the training loss is less than or equal to the preset loss threshold, and obtain the trained general model.
[0110] It should be noted that the pre-defined encoder can encode target human auditory physiological data, converting high-dimensional target human auditory physiological data into low-dimensional auditory-related feature vectors. This allows for downsampling and encoding of the target human auditory physiological data. Specifically, the pre-defined encoder can encode features influencing the HRTF from the target human auditory physiological data, thereby reducing model complexity and improving training accuracy. The encoding result of the pre-defined encoder can be input into the pre-defined decoder, which decodes the result to generate the HRTF training result. The HRTF training result is the HRTF corresponding to the target human auditory physiological data output by the pre-defined decoder in the pre-defined model.
[0111] Training loss characterizes the difference between the target HRTF label and the HRTF training result. The larger the training loss, the greater the difference between the target HRTF label and the HRTF training result. The preset loss threshold can be determined based on the actual situation, and this embodiment does not impose specific limitations on it. If the training loss is greater than the preset loss threshold, it indicates that the difference between the target HRTF label and the HRTF training result is large. Therefore, it is necessary to acquire new target human auditory physiological data for training until the training loss is less than or equal to the preset loss threshold, thus obtaining a trained general model.
[0112] For example, target human auditory physiological data is input into a preset model, encoded by a preset encoder in the preset model, and the encoded result is input into a preset decoder. The preset decoder outputs the HRTF training result. The training loss between the HRTF training result and the target HRTF label of the target human auditory physiological data is calculated. If the training loss is greater than a preset loss threshold, new target human auditory physiological data is obtained from the group auditory training dataset, and the step of inputting the target human auditory physiological data into the preset model is returned until the training loss is less than or equal to the preset loss threshold, thus obtaining a general model that has been trained.
[0113] This embodiment trains a preset model using a group auditory training dataset to obtain a general model, which facilitates the subsequent use of a general decoder in the general model to determine individual acoustic transmission data, so as to provide a good sound quality experience for each user.
[0114] In a feasible embodiment, step C20 further includes steps C201 to C202:
[0115] Step C201: Obtain the head structure parameters of the target user, including auricle height, auricle width, ear canal tilt angle, head width, and shoulder and neck parameters.
[0116] Step C202: Construct a head geometric model of the target user using head structure parameters, and extract personal auditory physiological data from the head geometric model.
[0117] It should be noted that head width is the lateral distance between the left and right tragus, and head width determines the interaural time difference (ITD). Shoulder and neck parameters can include shoulder width, neck height, etc., and these parameters affect the reflection and scattering of low-frequency sound waves. Auricular height is the vertical distance from the tragus to the highest point of the helix, and auricular width can be the maximum horizontal distance from the anterior edge of the helix to the posterior edge of the antihelix. The ear canal tilt angle is the angle between the ear canal axis and the sagittal plane (plane of symmetry) of the head, which affects the direction of sound wave incidence.
[0118] A geometric model of the target user's head can be constructed using head structural parameters. 3D modeling software can be used to construct this geometric model based on these parameters. Features directly related to auditory function can be extracted directly from the head geometric model to obtain individual auditory physiological data. For example, this data may include ear canal resonant frequencies, the angle between the tragus and cymba conchae, and the ITD baseline value. This embodiment does not impose specific limitations on these parameters; they can be determined based on actual circumstances. In other embodiments, head structural parameters may also include jaw width, ear canal length, etc., which are not specifically limited in this embodiment.
[0119] For example, the head structure parameters of the target user can be obtained by taking a picture. Modeling software is then called to construct a geometric model of the target user's head based on the head structure parameters, and personal auditory physiological data can be extracted from the head geometric model.
[0120] This embodiment constructs a head geometry model of the target user using head structure parameters, which facilitates the extraction of personal auditory physiological data from the head geometry model, and further facilitates the determination of personalized acoustic transmission data, so as to improve the sound quality performance of audio devices.
[0121] In a possible embodiment, step C20 further includes step C203 and / or step C204:
[0122] Step C203: Obtain the user's three-dimensional head scan data and determine the individual's auditory physiological data from the three-dimensional head scan data;
[0123] Step C204: Obtain the user's head measurement parameters, and obtain the target sub-auditory physiological data corresponding to each sub-measurement parameter in the preset auditory physiological database to obtain personal auditory physiological data composed of multiple target sub-auditory physiological data. The multiple sub-measurement parameters included in the head measurement parameters are external ear structure parameters, head-neck connection parameters, and head body parameters. The preset auditory physiological database includes preset sub-auditory physiological data corresponding to multiple preset sub-measurement parameters.
[0124] It should be noted that 3D head scan data can be obtained by performing 3D structured light or laser scanning on the target user, and personal auditory physiological data related to hearing can be extracted from the 3D head scan data.
[0125] Head measurement parameters can be obtained by directly measuring the target user's head and the shoulders and neck that connect to it, for example, through multi-view photography. Head measurement parameters can include multiple sub-measurement parameters, which can be external ear structure parameters, head-neck connection parameters, and head body parameters. Each sub-measurement parameter can further include multiple sub-parameters. External ear structure parameters can include sub-parameters such as auricle contour, auricle folds, auricle size, ear canal length and diameter, and concha depth. Head-neck connection parameters can include sub-parameters such as shoulder width and neck length. Head body parameters can include sub-parameters such as head diameter and interauricular distance.
[0126] The preset auditory physiological database can be determined in advance based on actual conditions. It can include preset sub-auditory physiological data corresponding to multiple preset sub-measurement parameters. The target sub-auditory physiological data for each sub-measurement parameter can be found in the preset auditory physiological database.
[0127] For example, a 3D structured light and / or laser scan can be performed on the target user to obtain three-dimensional head scan data, and personal auditory physiological data related to hearing can be obtained from the three-dimensional head scan data.
[0128] It can also acquire the user's head measurement parameters. The target sub-auditory physiological data corresponding to each sub-measurement parameter in the head measurement parameters can be retrieved from a preset auditory physiological database, resulting in personal auditory physiological data composed of multiple target sub-auditory physiological data. Specifically, each sub-measurement parameter also includes multiple sub-parameters. The preset sub-auditory physiological data corresponding to each sub-parameter can be found in the preset auditory physiological database, resulting in personal auditory physiological data composed of multiple preset sub-auditory physiological data. The preset sub-measurement parameters also include multiple preset sub-parameters, and each preset sub-measurement parameter can correspond to multiple preset sub-auditory physiological data; this embodiment does not specifically limit this.
[0129] This embodiment can acquire personal auditory physiological data in multiple ways, thereby improving the flexibility of data acquisition. For example, it can also refer to... Figure 5 , Figure 5 Multiple methods for obtaining personal auditory physiological data are presented. Personal auditory physiological data 500 can be extracted from head geometric model 100, or from three-dimensional head scan data 200. Alternatively, target sub-auditory physiological data corresponding to each sub-measurement parameter in head measurement parameter 300 can be searched in preset auditory physiological database 400 to obtain personal auditory physiological data 500 composed of multiple target sub-auditory physiological data.
[0130] Based on the first, second, and / or third embodiments of this application, in the fourth embodiment of this application, the content that is the same as or similar to the above-described embodiments one, two, and three can be referred to the above description and will not be repeated hereafter. Furthermore, before the step of rendering the original audio based on the personalized acoustic transmission data to obtain personalized audio, the method further includes:
[0131] Step D10: Obtain the amplitude spectrum and phase spectrum of the personalized acoustic transmission data, and obtain the preset covert sound threshold.
[0132] Spectral analysis can be performed on individual acoustic transmission data to extract its amplitude and phase spectra. The amplitude spectrum reflects the magnitude of different frequency components, while the phase spectrum reflects the phase information of different frequency components.
[0133] The preset masking sound threshold can be pre-set or determined based on the masking effect characteristics of the human ear. For example, it can be set as the minimum intensity level at which a specific sound (the masked sound) can be perceived in the presence of masking sound. Masking sound refers to a sound that can affect the perception of other sounds. It can be in the form of pure tone, polyphony, or noise. Masking sound increases the hearing threshold of the masked sound, making sounds that can be heard in a quiet environment difficult to perceive when masking sound is present.
[0134] Step D20: Adjust the amplitude spectrum based on the preset covert sound threshold to obtain the adjusted amplitude spectrum, wherein the adjustment of the amplitude spectrum includes reducing the amplitude value of the target frequency point, and the target frequency point is the frequency point in the amplitude spectrum whose amplitude value is less than the preset covert sound threshold;
[0135] The algorithm iterates through all frequency points in the amplitude spectrum, reducing the amplitude values of target frequencies whose amplitude values are below a preset covert sound threshold. This adjustment aims to reduce the audio energy at these frequencies, minimizing their impact on auditory perception during subsequent audio processing. For example, the amplitude values of these target frequencies can be reduced to a preset low level, or adjusted according to a specific attenuation function, to ensure that the energy of the audio signal at these frequencies does not interfere with the user's auditory experience. In this way, the frequency response characteristics of the HRTF can be optimized to better align with the auditory perception patterns of the human ear.
[0136] It should be noted that for frequency points that are greater than or equal to the covert sound threshold, the amplitude value can remain unchanged, or a corresponding optimization adjustment method can be set based on actual needs. This embodiment does not impose specific restrictions on this.
[0137] Step D30: Combine the adjusted amplitude spectrum and the phase spectrum to obtain the adjusted personalized acoustic transmission data;
[0138] Step D40: Based on the adjusted personalized acoustic transmission data, perform the step of rendering the original audio according to the personalized acoustic transmission data to obtain personalized audio.
[0139] The adjusted amplitude spectrum is recombined with the acquired phase spectrum to obtain the adjusted personalized acoustic transfer data. This combination can be achieved by synthesizing the adjusted amplitude spectrum and the original phase spectrum in a complex form, thereby generating a new personalized acoustic transfer function. The adjusted personalized acoustic transfer data can better adapt to the auditory needs of the target user, while reducing the influence of covert sounds and improving the accuracy and effect of audio processing. In this way, the generated sound field is more in line with the auditory perception of the human ear.
[0140] In one possible implementation, prior to the step of obtaining the preset covert sound threshold, the method further includes:
[0141] Step E10: Obtain the audio parameters of the original audio, wherein the audio parameters include the power density spectrum and the sound pressure level;
[0142] The raw audio is analyzed to extract its audio parameters. These parameters can include the power density spectrum and sound pressure level. The power density spectrum reflects the energy distribution of the audio signal at different frequencies, while the sound pressure level represents the intensity of the audio signal. By analyzing the power density spectrum of the raw audio, we can understand the energy distribution of the audio signal at various frequencies; by measuring the sound pressure level, we can determine the overall intensity of the audio signal.
[0143] Step E20: Identify at least one masking frequency point in the original audio based on the audio parameters, and obtain the individual sound pressure level of each masking frequency point;
[0144] The masking sound frequency can specifically be a frequency that meets the preset masking sound characteristics, such as a local peak point in the power density spectrum that maintains the maximum amplitude within a preset bandwidth (e.g., 0.5 barcks), or a frequency point whose sound pressure level is greater than the absolute hearing threshold of the human ear.
[0145] For each masked audio frequency, its individual sound pressure level (SPL) is further obtained, i.e., the SPL at that frequency. The individual SPL reflects the intensity of the audio signal at that frequency.
[0146] Step E30: Determine the minimum value between the individual sound pressure level and the absolute hearing threshold, and set the minimum value as the preset covert sound threshold.
[0147] For each identified masking frequency, its individual sound pressure level is compared to its absolute hearing threshold. The absolute hearing threshold is the lowest sound intensity that the human ear can perceive; it is a fixed value, usually expressed in decibels (dB). By comparing the individual sound pressure level with the absolute hearing threshold, the minimum value between the two is determined. This minimum value will be set as the preset masking sound threshold.
[0148] This embodiment determines the minimum value between the individual sound pressure level of each masking frequency point and the absolute hearing threshold as the preset masking sound threshold. This threshold determination comprehensively considers the energy of the masking frequency point and the hearing sensitivity of the human ear, ensuring that the masking sound can effectively mask other sounds while avoiding over-processing and distortion caused by excessively strong masking sound. Furthermore, the masking sound threshold determined based on the original audio more accurately reflects the actual masking capability of the audio, ensuring that the set threshold matches the actual masking characteristics of the current audio to be output, further improving the quality and effect of subsequent audio processing.
[0149] In one possible implementation, after the step of obtaining the target transfer function between the test audio device at the extended target location and the user's two ears, the method further includes:
[0150] Step F10: Obtain the ambient acoustic parameters of the target user's current environment, wherein the ambient acoustic parameters include reflected sound intensity, reverberation time, and dominant reflection direction;
[0151] After obtaining the target transfer function between the test audio device at the target location and the user's ears, the ambient acoustic parameters of the target user's current environment are further obtained. These parameters include reflected sound intensity, reverberation time, and dominant reflection direction. Specifically, a microphone array can be placed in the target user's current environment, and a set of test signals can be played through the microphone array to acquire the room impulse response (RIR) in real time. The ambient acoustic parameters are then calculated based on the acquired room impulse response.
[0152] Reflected sound intensity refers to the energy of reflected sound in the environment within a specific time window, such as the reflected sound energy within 5-50ms after the test signal is emitted; reverberation time refers to the time required for sound to decay to a certain level in the environment, such as the time required for the test signal to decay by 60dB; dominant reflection direction refers to the main propagation direction of reflected sound, such as the direction in which the reflected sound is strongest.
[0153] Step F20: Generate the corresponding environmental compensation coefficient based on the environmental acoustic parameters;
[0154] After obtaining the environmental acoustic parameters, corresponding environmental compensation coefficients are generated based on these parameters. The generation of environmental compensation coefficients aims to adjust the target transfer function according to the environmental acoustic characteristics in order to compensate for the impact of the environment on audio propagation and perception.
[0155] Specifically, different correspondences between environmental acoustic parameters and environmental compensation coefficients can be pre-defined through experimental measurement or experience, and the environmental compensation coefficients corresponding to the environmental acoustic parameters can be obtained based on these correspondences. Alternatively, a prediction model can be pre-trained with environmental acoustic parameters as input and environmental compensation coefficients as output. After obtaining the environmental acoustic parameters, the environmental acoustic parameters can be input into the prediction model, and the corresponding environmental compensation coefficients can be output.
[0156] It should be noted that the transfer function can be represented as a complex matrix, and therefore, the environmental compensation coefficient can be a coefficient matrix.
[0157] Step F30: Multiply the environmental compensation coefficient by the target transfer function to obtain the compensated target transfer function;
[0158] Step F40: Based on the compensated target transfer function, perform the crosstalk cancellation process on the personalized audio according to the target transfer function to obtain the target audio.
[0159] Based on the compensated target transfer function, a crosstalk cancellation process is performed on the personalized audio according to the target transfer function to obtain the target audio. The purpose of crosstalk cancellation is to eliminate the crosstalk effect caused by mutual interference between speakers during audio signal propagation, thereby improving the positioning accuracy and spatial sense of the audio. Performing crosstalk cancellation based on the compensated target transfer function can more accurately consider the influence of environmental acoustic characteristics on audio propagation, thus ensuring that the target audio played subsequently can meet the acoustic transmission characteristics of the current spatial environment, enabling higher quality audio effects in different environments. For example, in an environment with strong reflected sound, crosstalk cancellation through the compensated target transfer function can reduce the interference of reflected sound on audio positioning, allowing users to more accurately perceive the direction of the audio source and enhance the immersion and realism of the audio.
[0160] Furthermore, embodiments of this application also propose an audio device, the audio device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the sound field expansion method as described above.
[0161] refer to Figure 6 The diagram illustrates a structural schematic of an audio device suitable for implementing the embodiments of this application. The audio device in the embodiments of this application may also include, but is not limited to, mobile terminals such as AR headsets, VR headsets, headphones, mobile phones, servers, laptops, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), etc., as well as fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The audio device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0162] like Figure 6 As shown, the audio device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the audio device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), test audio equipment, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows audio devices to communicate wirelessly or wiredly with other devices to exchange data. While audio devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0163] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0164] The audio device provided in this application, employing the sound field expansion method described in the above embodiments, can solve the technical problem of how to effectively expand the sound field while reducing crosstalk and ensuring sound quality. Compared with the prior art, the beneficial effects of the audio device provided in this application are the same as those of the sound field expansion method provided in the above embodiments, and other technical features of this audio device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0165] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0166] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0167] In addition, to achieve the above objectives, embodiments of this application also provide a readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the sound field expansion method in the above embodiments.
[0168] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0169] The aforementioned computer-readable storage medium may be included in an audio device or may exist independently without being assembled into an audio device.
[0170] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an audio device, cause the audio device to: acquire a target transfer function between a test audio device at an extended target location and the user's two ears, and acquire personalized acoustic transmission data of the target user, wherein the personalized acoustic transmission data includes at least one head-related transfer function (HRTF) generated based on the target user's personal auditory physiological data; render original audio based on the personalized acoustic transmission data to obtain personalized audio; perform crosstalk cancellation processing on the personalized audio according to the target transfer function to obtain target audio, and play the target audio.
[0171] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0173] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the modules themselves.
[0174] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described sound field expansion method. This solves the technical problem of how to effectively expand the sound field while reducing crosstalk and ensuring sound quality. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the sound field expansion method provided in the above embodiments, and will not be repeated here.
[0175] Furthermore, embodiments of this application also propose a computer program product, including a computer program that, when executed by a processor, implements the steps of the sound field expansion method as described above.
[0176] The specific implementation of the computer program product in this application is basically the same as the embodiments of the sound field expansion method described above, and will not be repeated here.
[0177] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0178] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0179] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software sensor. This computer software sensor is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause an audio device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0180] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A sound field extension method, characterized by, The sound field expansion method comprises the following steps: obtaining a target transfer function between a test audio device at an expansion target position and user binaural ears, and obtaining individual acoustic transmission data of a target user, wherein the individual acoustic transmission data comprises at least one head-related transfer function (HRTF) generated based on personal auditory physiological data of the target user; obtaining an amplitude spectrum and a phase spectrum of the individual acoustic transmission data, and obtaining audio parameters of original audio, wherein the audio parameters comprise a power density spectrum and a sound pressure level; identifying at least one masking audio point in the original audio based on the audio parameters, and obtaining individual sound pressure levels of the masking audio points; determining a minimum value of the individual sound pressure levels and the absolute auditory threshold, and setting the minimum value as a preset masking sound threshold; adjusting the amplitude spectrum based on the preset masking sound threshold to obtain an adjusted amplitude spectrum, wherein the adjustment of the amplitude spectrum comprises reducing the amplitude value of a target frequency point, the target frequency point being a frequency point in the amplitude spectrum with an amplitude value less than the preset masking sound threshold; combining the adjusted amplitude spectrum and the phase spectrum to obtain the adjusted individual acoustic transmission data; rendering the original audio according to the adjusted individual acoustic transmission data to obtain individual audio; performing crosstalk cancellation processing on the individual audio according to the target transfer function to obtain target audio, and playing the target audio.
2. The sound field extension method of claim 1, wherein, The step of obtaining the target transfer function between the test audio device at the expansion target position and the user binaural ears comprises: obtaining an artificial head transfer function and a free field transfer function; performing inverse operation on the free field transfer function to obtain a free field inverse transfer function; multiplying the artificial head transfer function and the free field inverse transfer function to obtain the target transfer function between the test audio device at the expansion target position and the user binaural ears.
3. The sound field extension method of claim 2, wherein, The step of obtaining the artificial head transfer function and the free field transfer function comprises: when the test audio device is placed at the expansion target position and the test audio device outputs a sound signal, measuring the artificial head transfer function through a preset microphone in a preset artificial head ear canal; and when the preset artificial head is removed and the test audio device outputs a sound signal, measuring the free field transfer function through preset microphones placed at left and right ear positions before the artificial head is removed.
4. The sound field extension method of claim 2, wherein, The step of performing crosstalk cancellation processing on the individual audio according to the target transfer function to obtain target audio comprises: performing inverse operation on the target transfer function to obtain a target inverse transfer function; multiplying the individual audio and the target inverse transfer function to obtain the target audio.
5. The sound field expansion method of claim 1, wherein, The step of obtaining the individual acoustic transmission data of the target user comprises: training a preset model to obtain a trained general model through a group auditory training data set, wherein the group auditory training data set comprises multiple human auditory physiological data and respective corresponding HRTF labels, and the general model comprises a general encoder and a general decoder; obtaining personal auditory physiological data of the target user; Obtaining a general decoder from the general model, training the general decoder by the personal auditory physiological data to obtain a personalized decoder, and generating individualized acoustic transmission data corresponding to the personal auditory physiological data based on the personalized decoder.
6. The sound field extension method according to any one of claims 1 to 5, wherein, After the step of obtaining the target transfer function between the test audio device and the user's binaural ears at the target position, the method further comprises: Obtaining an environmental acoustic parameter of an environment in which the target user is currently located, wherein the environmental acoustic parameter comprises a reflection intensity, a reverberation time, and a dominant reflection direction; Generating a corresponding environmental compensation coefficient based on the environmental acoustic parameter; Multiplying the environmental compensation coefficient and the target transfer function to obtain a compensated target transfer function; Based on the compensated target transfer function, the step of performing crosstalk cancellation processing on the individualized audio according to the target transfer function to obtain target audio.
7. An audio device, comprising: Comprise: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is executed by the processor to implement the sound field expansion method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a sound field expansion program, and the sound field expansion program is executed by the processor to implement the steps of the sound field expansion method according to any one of claims 1 to 6.
Citation Information
Patent Citations
A method and device for vectorizing translation personality characteristics of a translator
CN109670180A
Cross sound elimination method and device, audio equipment and computer readable storage medium
CN115278474A
Audio optimization system and method for in-vehicle screen projection
CN119012095A
Audio processing method and electronic equipment
CN120416754A