Audio processing method and electronic equipment

By obtaining user physiological parameters to generate personalized head-related transfer functions, the problem that audio processing in the prior art cannot provide each user with accurate immersive sound effects, and improves the personalization and accuracy of the audio experience.

CN120416754APending Publication Date: 2025-08-01HONOR DEVICE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410098738.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing audio processing technologies are difficult to provide each user with a personalized and accurate immersive sound experience because there are differences in physiological characteristics of each user, which makes it impossible for the general head-dependent transfer function to accurately simulate the position of the sound source in three-dimensional space.

Method used

By obtaining the physiological parameters of the head and ear areas of the target user, combining the mean and position of the pre-stored sample user head-related transfer function, a target head-related transfer function adapted to the target user is generated, and the audio is rendered to generate audio adapted to the target user.

Benefits of technology

It realizes the generation of personalized audio based on user physiological characteristics, improves the user's spatial sound experience, and ensures the accuracy of the position perception of the sound source in three-dimensional space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416754A_ABST
    Figure CN120416754A_ABST
Patent Text Reader

Abstract

The invention provides an audio processing method and electronic equipment, and the method comprises the steps: obtaining a first user image which comprises a target region of a target user; determining a first size of the target area; according to the first size, the first head-related transfer function and the first position, a target head-related transfer function is obtained, and the first head-related transfer function is a mean value of head-related transfer functions of N sample users pre-stored in the electronic equipment; in response to the playing operation of the first audio, according to the target head-related transfer function, the first audio in the electronic equipment is subjected to audio rendering, a second audio is generated, the second audio is the audio with the first sound effect, and the hearing position of a sound source of the second audio is the first position of the simulated three-dimensional space. Therefore, the audio data adaptive to the physiological features of the user can be generated for the user, so that the user can obtain better spatial sound experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing, and particularly to an audio processing method and an electronic device. Background Art

[0002] With the rapid development of audio processing technology, users have put forward higher requirements for the audio sound effect experience. For example, when an electronic device establishes a communication connection with a headset and the electronic device plays audio through the headset, users want an immersive sound effect experience with a sense of space.

[0003] Currently, an electronic device can render audio based on a general binaural playback spatial sound technology to obtain audio with a sense of space sound effect, thereby meeting the user's immersive sound effect experience.

[0004] However, since there are differences in the physiological characteristics of each user, it is difficult to provide a completely accurate and personalized sound experience for each user using a general technology. Summary of the Invention

[0005] This application provides an audio processing method and an electronic device, which can generate audio data adapted to the user's physiological characteristics as much as possible, so that the user can obtain a better spatial sound experience.

[0006] In a first aspect, this application provides an audio processing method, which is applied to an electronic device. The method includes:

[0007] Obtain a first user image, where the first user image includes a target area, and the target area refers to an image area corresponding to one or more parts of the target user's body; determine a first size of the target area; according to the first size, a first head-related transfer function, and a first position, obtain a target head-related transfer function, where the first head-related transfer function is the average of the head-related transfer functions of N sample users pre-stored in the electronic device, the first position is the position of a sound source in a pre-stored simulated three-dimensional space in the electronic device, and the target head-related transfer function is used to indicate the head-related transfer function of the target user; in response to a playback operation on a first audio, perform audio rendering on the first audio in the electronic device according to the target head-related transfer function to generate a second audio, where the second audio is an audio with a first sound effect, and the perceived position of the sound source of the second audio in the first sound effect is the first position in the simulated three-dimensional space.

[0008] In the above method, by obtaining the first user image, the first size corresponding to the target area can be determined. Since the first user image includes the target area of the target user, and the target area includes the head area and / or the ear area, the electronic device determining the first size of the head area and / or the ear area can prepare data for generating a target head-related transfer function that matches the target user.

[0009] Based on this, the target head-related transfer function that matches the target user can be obtained according to the first dimension, the first head-related transfer function, and the first position, and the first audio in the electronic device can be audio-rendered according to the target head-related transfer function to generate the second audio that is adapted to the target user and whose sound source is at the first position in the simulated three-dimensional space. Thus, the electronic device can overcome the physiological difference problem, generate audio adapted to the target user, and facilitate the target user to obtain a better spatial sound experience.

[0010] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes:

[0011] Calibrating the target head-related transfer function to obtain a calibrated target head-related transfer function. The target head-related transfer function is associated with the first position, and the calibrated target head-related transfer function is associated with the second position. The second position is a calibration position for making the perceived position of the sound source of the audio be at the first position; audio-rendering the first audio in the electronic device according to the target head-related transfer function to generate the second audio includes: audio-rendering the first audio in the electronic device according to the calibrated target head-related transfer function to generate the third audio. The third audio is an audio with a second sound effect. The similarity between the perceived position of the sound source in the second sound effect and the first position is higher than the similarity between the perceived position of the sound source in the first sound effect and the first position.

[0012] In the above method, after obtaining the target head-related transfer function, the target head-related transfer function can also be calibrated to obtain a calibrated target head-related transfer function. In this way, the electronic device can audio-render the first audio in the electronic device according to the calibrated target head-related transfer function to generate the third audio that matches the target user and has a second sound effect. Since the calibrated target head-related transfer function is associated with the second position, and the second position is a calibration position for making the perceived position of the sound source of the audio be at the first position, then, audio-rendering the first audio in the electronic device according to the calibrated target head-related transfer function to generate the third audio can make the position accuracy of the sound source of the third audio perceived by the target user high, enable the perceived position of the sound source of the second audio to be at the first position, avoid the situation where the perceived position of the sound source of the second audio by the user is far from the first position, and ensure the playback quality of the third audio.

[0013] Combined with the first aspect, in some implementation manners of the first aspect, calibrating the target head-related transfer function to obtain a calibrated target head-related transfer function includes:

[0014] Input the target head-related transfer function into a neural network model to obtain a corrected target head-related transfer function. The neural network model is used to predict the relationship between the position associated with the head-related transfer function and the perceived position of the sound source in the sound effect of the audio rendered by the head-related transfer function.

[0015] In the above method, the corrected target head-related transfer function can be obtained more quickly through the neural network model.

[0016] Combined with the first aspect, in some implementation manners of the first aspect, the generation process of the neural network model includes:

[0017] Obtain P sample head-related transfer functions, P sample positions, and P sample perceived positions. The P sample head-related transfer functions are respectively in one-to-one correspondence with the P sample positions and the P sample perceived positions. The P sample perceived positions are respectively used to indicate the true perceived positions of the sound sources in the sound effects of the audio rendered by the P sample head-related transfer functions. P is a positive integer greater than or equal to 2; input the P sample head-related transfer functions and the P sample positions into the original neural network model to obtain P predicted positions, and the P predicted positions are respectively used to indicate the predicted perceived positions of the sound sources in the sound effects of the audio rendered by the P sample head-related transfer functions; train the original neural network model according to the differences between the P sample perceived positions and the corresponding P predicted positions to obtain the neural network model.

[0018] In the above method, the original neural network model can be trained according to the differences between the P sample perceived positions and the corresponding P predicted positions, so as to obtain the neural network model.

[0019] Combined with the first aspect, in some implementation manners of the first aspect, the first position is determined according to the first angle value. According to the first size, the first head-related transfer function, and the first position, obtaining the target head-related transfer function includes:

[0020] Obtain the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value. The first angle value includes the angle value of the azimuth angle and / or the angle value of the elevation angle. The angle value of the azimuth angle is used to represent the included angle between the line connecting the sound source of the audio in the simulated three-dimensional space and the target user, and the extension line of the front of the target user on the horizontal plane. The angle value of the elevation angle is used to represent the included angle between the line connecting the sound source of the audio in the simulated three-dimensional space and the target user, and the extension line of the front of the target user on the vertical plane.

[0021] In the above method, the accurate target head-related transfer function can be obtained by combining the first angle value, so as to facilitate generating the second audio corresponding to the specific position according to the target head-related transfer function.

[0022] In combination with the first aspect, in some implementations of the first aspect, the number of the first angle values is M groups, and the M groups of first angle values are used to indicate M positions in the simulated three-dimensional space. M is a positive integer greater than or equal to 2. Obtaining the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value includes:

[0023] Obtaining M head-related transfer functions according to the first size, the first head-related transfer function, and the M groups of first angle values, where the M head-related transfer functions are respectively associated with the M groups of first angle values; determining the second head-related transfer function among the M head-related transfer functions as the target head-related transfer function, and the first head-related transfer function is the head-related transfer function corresponding to a group of first angle values selected based on the operation of the target user among the M head-related transfer functions.

[0024] In the above method, multiple groups of target head-related transfer functions can be obtained by combining multiple groups of first angle values, and the user can select from the multiple groups of target head-related transfer functions according to the angle values, so as to facilitate the generation of the second audio required by the user and ensure the user experience.

[0025] In combination with the first aspect, in some implementations of the first aspect, the number of the first angle values is M groups, and the M groups of first angle values are used to indicate M positions in the simulated three-dimensional space. M is a positive integer greater than or equal to 2. Obtaining the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value includes:

[0026] Obtaining the third head-related transfer function according to the first size, the first head-related transfer function, and the second angle value, where the second angle value includes a group of first angle values selected based on the operation of the target user among the M groups of first angle values; determining the third head-related transfer function as the target head-related transfer function.

[0027] In the above method, the user can select from multiple groups of target head-related transfer functions according to the angle values, so as to facilitate the generation of the target head-related transfer function required by the user. Based on this, the second audio required by the user can be generated, and the user experience can be ensured.

[0028] In combination with the first aspect, in some implementations of the first aspect, obtaining the target head-related transfer function according to the first size, the first head-related transfer function, and the first position includes:

[0029] Inputting the first size, the first head-related transfer function, and the first position into a generative diffusion network to obtain the target head-related transfer function. The generative diffusion network is used to generate a head-related transfer function adapted to the target user by combining the first size and the first position of the target user on the basis of the average data corresponding to the head-related transfer functions of multiple users.

[0030] In the above method, a target head-related transfer function can be quickly generated through a generative diffusion network.

[0031] In combination with the first aspect, in some implementation manners of the first aspect, the generation process of the generative diffusion network includes:

[0032] Obtain N sample user images, where each sample user image in the N sample user images includes a target region, and N is a positive integer greater than or equal to 2; determine N groups of sample sizes corresponding to the N target regions; input the N groups of sample sizes into the original generative diffusion network to obtain N sample head-related transfer functions; determine the average data corresponding to the N head-related transfer functions as the first head-related transfer function; and train the original generative diffusion network according to the relationships between the N sample head-related transfer functions and the first head-related transfer function respectively to obtain the generative diffusion network.

[0033] In the above method, the original generative diffusion network can be trained according to the relationships between the N sample head-related transfer functions and the first head-related transfer function respectively, so as to obtain an accurate generative diffusion network.

[0034] In combination with the first aspect, in some implementation manners of the first aspect, the target region includes the head region, the ear region, and / or the image region corresponding to the torso region, and the head region includes other head regions except the ear region; the first size corresponding to the head region includes at least one of head circumference, head width, head depth, head height, neck width, neck height, neck depth, anterior cranial offset, shoulder width, and shoulder circumference; the first size corresponding to the ear region includes at least one of auricle height, auricle width, ear cavity height, ear cavity width, ear cavity depth, cymba conchae height, cochlea height, intertragal width, auricle rotation angle, auricle flare angle, auricle inferior offset, and auricle right offset; the first size corresponding to the torso region includes at least one of torso width, torso height, torso depth, height, and sitting height.

[0035] In combination with the first aspect, in some implementation manners of the first aspect, the electronic device is applied to an audio system, and the audio system includes the electronic device and a wearable device. After the first audio in the electronic device is audio-rendered according to the target head-related transfer function to generate the second audio, the method further includes:

[0036] Send the second audio to the wearable device so that the wearable device plays the second audio for the target user.

[0037] In the above method, after the electronic device obtains the second audio, it can send the second audio to the wearable device, so that when the wearable device receives the second audio, it can play the second audio. Based on this, the target user can hear the second audio.

[0038] In a second aspect, the present application provides an audio processing device, which is configured to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0039] In a third aspect, the present application provides an electronic device, which includes: one or more processors, and a memory; the memory is coupled to the one or more processors, and the memory is configured to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0040] In a fourth aspect, the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are configured to call computer instructions to cause the electronic device to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0041] In a fifth aspect, the present application provides a computer-readable storage medium, the computer-readable storage medium includes instructions, and when the instructions run on an electronic device, the electronic device is caused to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0042] In a sixth aspect, the present application provides a computer program product, and when the computer program product runs on a computer, the computer is caused to execute the method in the first aspect and any possible implementation manner of the first aspect.

[0043] It can be understood that the beneficial effects of the above second aspect to the sixth aspect can refer to the relevant descriptions in the first aspect above, and will not be elaborated here. Description of the Drawings

[0044] Figure 1 A schematic diagram of a generative model provided by the prior art;

[0045] Figure 2 A schematic diagram of a scenario of an audio processing method provided by an embodiment of the present application;

[0046] Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0047] Figure 4 A schematic diagram of a human-computer interaction interface provided by an embodiment of the present application;

[0048] Figure 5 A schematic diagram of a human-computer interaction interface provided by an embodiment of the present application;

[0049] Figure 6Schematic diagram of a human-computer interaction interface provided by an embodiment of the present application;

[0050] Figure 7 Schematic flow diagram of an audio processing method provided by an embodiment of the present application;

[0051] Figure 8 Schematic diagram of the training process of a generative diffusion network provided by an embodiment of the present application;

[0052] Figure 9 Schematic diagram of the generation process of a head-related transfer function provided by an embodiment of the present application;

[0053] Figure 10 Schematic flow diagram of an audio processing method provided by an embodiment of the present application;

[0054] Figure 11 Schematic diagram of the training process of a neural network model provided by an embodiment of the present application;

[0055] Figure 12 Schematic diagram of the structure of an audio processing device provided by an embodiment of the present application. Detailed implementation manners

[0056] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or similar expressions thereof refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a alone, b alone, or c alone may represent: a alone, b alone, c alone, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b, and c, where a, b, and c may be single or multiple. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0057] For ease of understanding, some examples are given to illustrate the related concepts of the embodiments of the present application for reference.

[0058] 1. Spatial audio technology based on binaural playback.

[0059] Spatial audio technology based on binaural playback is a technology that utilizes the propagation characteristics of sound to simulate the sound source position and sound field effects in a three-dimensional space. This technology plays processed sound signals specifically in each ear of the user, enabling the user to perceive that the sound comes from different positions (direction and distance).

[0060] 2. Head-Related Transfer Function (HRTF).

[0061] In the field of audio processing, HRTF is used to indicate the influence of the head and ears on sound propagation and localization. HRTF is a function that describes the transformation that occurs to a sound signal during the process of passing through the head and ears. Using HRTF to perform convolution processing on audio can convert an ordinary sound source into a sound that seems to come from a specific position.

[0062] 3. Head-Related Impulse Response Function (HRIR)

[0063] HRTF is obtained by performing a time-domain transformation on the Head-Related Impulse Response Function (HRIR).

[0064] 4. Mean HRIR

[0065] Mean HRIR is the average of the HRIRs of multiple users. It represents an average response pattern corresponding to the head and ears. Performing a time-domain transformation on this average value can obtain the mean HRTF.

[0066] 5. Personalized HRIR (personal HRTF).

[0067] Personalized HRIR is used to indicate the personalized head-related impulse response obtained according to the parameters of an individual's head and ears.

[0068] Personalized HRIR can more accurately simulate the characteristics of an individual in terms of sound propagation and localization. Performing a time-domain transformation on this personalized HRIR can obtain the personalized HRTF.

[0069] 6. Generative Diffusion Network

[0070] Generative models are a type of machine learning model used to generate new data samples by learning the distribution probability from training data.

[0071] The diffusion network is a generative model that can be used to generate data of types such as images, speech, and text.

[0072] The diffusion network can be a network obtained by combining a convolutional neural network and a residual network (convolutional neural network - residual network, CNN + ResNet) or a U-Net.

[0073] As Figure 1 shown is the inference process when the diffusion network is a U-Net network. During the inference process, the diffusion network can predict the noise at each step (t-step noise) and add random white noise to generate diverse results. This method can make the generated results more diverse and realistic.

[0074] Among them, during the training stage of the diffusion network, it is trained by continuously adding noise until white noise is finally obtained:

[0075] In the diffusion network, noise can be added to the training dataset, and this process is carried out iteratively. More noise will be gradually added in each iterative step.

[0076] That is to say, at the beginning of training, the difference between the input data and the target data is small. As training continues, the added noise will gradually increase, making the input data closer to white noise (a random signal without obvious patterns or structures). Such training can better enable the network to learn the characteristics of the target distribution.

[0077] Among them, in each training step, the noise signal can be introduced into the diffusion network by randomly adding t-step noise.

[0078] Then, the diffusion network uses the current noise as the input to generate an output.

[0079] Based on this, the loss function can also be defined by Bayesian to evaluate the difference between the generated results and the real samples. The goal of the loss function is to make the results generated by the network as close to the real data as possible while maintaining the diversity of the generated results between the generated results and the real samples.

[0080] In summary, the diffusion network in the generative network trains the model to learn to generate data close to white noise by iteratively adding noise. During the training and inference stages, by predicting the noise and combining random noise, the generated results are made diverse and as close as possible to the expected generated data.

[0081] 7. Neural network model.

[0082] A neural network model is a computational model composed of multiple neurons, inspired by the human brain's nervous system. The neural network model simulates and solves various tasks by learning a large amount of data. The neural network model is also known as a deep learning model, and usually consists of multiple layers of neuron structures.

[0083] Neural network models include various types, commonly including feedforward neural networks (FNN), recurrent neural networks (RNN), convolutional neural networks (CNN), generative adversarial networks (GAN), etc.

[0084] 8. Auditory perception position.

[0085] The auditory perception position of an audio is used to indicate the user's perception of the sound position in space when listening to the audio.

[0086] With the rapid development of audio processing technology, users have higher requirements for the audio sound effect experience. For example, when an electronic device establishes a communication connection with a headset and plays audio through the headset, users want an immersive sound effect experience with a sense of space.

[0087] Currently, an electronic device can render audio based on the general binaural reproduction spatial sound technology to obtain audio with a sense of space sound effect, so as to meet the user's immersive sound effect experience.

[0088] Specifically, in the spatial sound technology based on binaural reproduction, an electronic device usually uses the average HRTF to simulate the position and effect of a sound source in space, so as to transmit the sound signal processed by HRTF to the user.

[0089] However, there are differences in each user's physiological characteristics (such as the size of the ears, head, and / or torso), so each user's HRTF is also different, resulting in it being difficult for the general HRTF to provide a completely accurate and personalized sound experience for each user.

[0090] That is to say, it is difficult to achieve an extreme spatial sound experience by using general technology to render audio.

[0091] Currently, there are also some methods to approximately approximate the user's HRTF, but the effect of this method is difficult to be stable compared with the effect of the general HRTF.

[0092] For example, an electronic device plays a spatial audio effect based on the average HRTF and an angular value of 30 degrees in azimuth in a virtual three-dimensional space (30 degrees to the right of the user). The sound source of the audio heard by one user may come from the direction 30 degrees to the right of the user, while the sound source of the audio heard by another user may come from the direction 90 degrees to the right of the user, which brings great differences in the user experience for activities such as watching movies and listening to music.

[0093] In the face of the above problems, the present application can provide an audio processing method, an electronic device, a chip system, a computer-readable storage medium, and a computer program product. The electronic device can obtain a user image including the head region and / or ear region of a target user, and obtain a first size of the head region, ear region, and / or torso region of the target user. Thus, the electronic device can combine the first size of the target user to obtain target head-related impulse data adapted to the target user. Therefore, when receiving a playback operation for the audio to be played, the electronic device can perform audio rendering on the audio to be played according to the target head-related impulse data adapted to the target user to generate an audio with a sound source of the audio in a first position in a simulated three-dimensional space and adapted to the target user; thus, the electronic device can overcome the problem of physiological differences and generate an audio adapted to the target user, facilitating the target user to obtain a better spatial sound experience.

[0094] Among them, the above-mentioned target user can be the user using the electronic device, and the target user can be the owner of the electronic device.

[0095] The above-mentioned electronic device can be an electronic device with the function of connecting a wearable device. The above-mentioned audio processing method is applied to the electronic device, and the electronic device is applied to an audio system. The audio system includes the electronic device and the wearable device, and the audio generated by the above-mentioned audio processing method can be played to the target user by the wearable device.

[0096] For example, the electronic device can be a mobile phone, a tablet computer, a vehicle-mounted device, a laptop computer, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart car, a smart TV, a robot, and other devices.

[0097] For example, the wearable device can be a wireless earphone or a wired earphone.

[0098] When the wearable device is a wireless earphone, the electronic device is a device with Bluetooth function, and the electronic device can be connected to the wireless earphone through the Bluetooth function.

[0099] It should be noted that, in some possible implementation manners, the electronic device may also be referred to as a terminal device, a user equipment (UE), etc., and the embodiments of the present application do not limit this.

[0100] Please refer to Figure 2 , Figure 2 which shows a schematic diagram of a scenario of an audio processing method provided by an embodiment of the present application.

[0101] As Figure 2 shown, an audio processing method provided by an embodiment of the present application is applicable to an audio system including an electronic device and a wearable device.

[0102] Among them, the electronic device may be a device for generating audio.

[0103] Specifically, as Figure 2 shown, the electronic device may first obtain a user image of the target user, where the user image includes a front image and a side image of the user, and the user image may include the head region, ear region, and / or torso region of the target user.

[0104] Based on this, the electronic device may also obtain the physiological parameters of the user from the user image, and obtain a target head-related transfer function according to the physiological parameters of the user, the mean value of the head-related transfer functions of N sample users pre-stored in the electronic device, and the first position. In this way, the electronic device may store the target head-related transfer function.

[0105] When the electronic device receives an operation instruction from the target user to play audio, the electronic device may perform audio rendering on the corresponding audio in the electronic device according to the target head-related transfer function to generate audio adapted to the target user.

[0106] After the electronic device obtains the audio adapted to the target user, it may send the audio to the wearable device.

[0107] Among them, the wearable device may be a device for receiving the audio generated by the electronic device and playing the audio.

[0108] After the wearable device receives the audio sent by the electronic device, it may play the audio.

[0109] The embodiments of the present application do not limit the types of the electronic device and the wearable device.

[0110] For example, the electronic device is a mobile phone and the wearable device is a wired earphone.

[0111] Another example is that the electronic device is a tablet computer and the wearable device is a wireless earphone.

[0112] It should be understood that the above are examples of scenarios and do not limit the scenarios of this application in any way.

[0113] For ease of explanation, Figure 3 in [the description], the electronic device 100 is taken as an example of a mobile phone for illustration.

[0114] As Figure 3 shown, in some embodiments, the electronic device 100 may include a processor 101, a communication module 102, a display screen 103, a camera 104, sensors 105, an internal memory 106, a USB interface 107, an external memory interface 108, a charging management module 109, a power management module 110, and a battery 111, etc.

[0115] Among them, the processor 101 may include one or more processing units. For example, the processor 101 may include an application processor (AP), a modem processor, a graphics processor, an image signal processor (ISP), a controller, a memory, a video stream codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors 101.

[0116] The controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0117] A memory may also be provided in the processor 101 for storing instructions and data.

[0118] The communication module 102 may include antenna 1 and antenna 2, a mobile communication module, and / or a wireless communication module.

[0119] Among them, the wireless communication module can provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module can be one or more devices integrating at least one communication processing module. The wireless communication module receives electromagnetic waves via the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 101. The wireless communication module can also receive the signals to be sent from the processor 101, frequency-modulate them, amplify them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0120] Among them, the display screen 103 is used to display images, videos, etc. in the human-computer interaction interface. The display screen 103 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 103, where N is a positive integer greater than 1.

[0121] Optionally, the electronic device 100 may further include peripheral devices, such as a mouse, a button, an indicator light, a keyboard, a speaker, a microphone, etc.

[0122] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100.

[0123] In other embodiments, the electronic device 100 may include more or fewer components than shown, or combine some components, separate some components, or arrange the components differently. The components shown may be implemented in hardware, software, or a combination of software and hardware.

[0124] Based on the above description, combined with Figures 4 - 6 , details the audio processing method of the electronic device to implement this application. For the convenience of explanation, Figures 4 - 6 In the figure, the electronic device is a mobile phone, and the shooting direction of the mobile phone is a vertical screen direction.

[0125] See also Figures 4 - 6 , Figures 4 - 6 A schematic diagram of a human-computer interaction interface provided by an embodiment of the present application is shown.

[0126] The phone can display Figure 4 The interface 11 shown in (a) is used to obtain a front image of the first user image.

[0127] The interface 11 may include: a first prompt message and a control 1001 . The first prompt message may be a text message “Please take a frontal image”. The control 1001 is used to trigger the mobile phone to take a frontal image of the target user.

[0128] Upon receiving the user's Figure 4 After the control 1001 shown in (a) is triggered (such as click, double-click or long press operation), the mobile phone can collect the user's front image and Figure 4 The interface 11 shown in (a) in FIG. 1 changes to display as shown in FIG. Figure 4 The interface 12 shown in (b) is used to obtain a side image of the first user image.

[0129] The interface 12 may include: a second prompt message and a control 1001 , the second prompt message may be a text message “Please take a side image”, and the control 1001 is used to trigger the mobile phone to take a front image of the target user.

[0130] Upon receiving the user's Figure 4 After the control 1001 shown in (b) is triggered, the mobile phone can collect the side image of the user.

[0131] Based on this, the mobile phone can obtain the front image and the side image included in the first user image, thereby starting to determine the first size of the head area, ear area, and / or torso area of the target user.

[0132] Among them, combined Figure 4 (a) and Figure 4In (b) thereof, the first dimension corresponding to the head region includes at least one of head circumference, head width (L1), head depth (L3), head height (L2), neck width, neck height (L7), neck depth (L8), anterior cranial offset (L13), shoulder width (L12), and shoulder circumference.

[0133] The first dimension corresponding to the ear region includes at least one of auricle height, auricle width, ear cavity height, ear cavity width, ear cavity depth, cymba conchae height, fossa of helix height, intertragic width, auricle rotation angle, auricle flare angle, auricle inferior offset, and auricle right offset.

[0134] The first dimension corresponding to the torso region includes at least one of torso width (L9), torso height (L10), torso depth (L11), height, and sitting height.

[0135] The above-mentioned other Figure 4 in (a) and Figure 4 The first dimensions corresponding to the head region and the torso region that are not marked in (b) thereof can be obtained by calculating from the marked first dimensions above, which will not be elaborated here.

[0136] In addition, the above-mentioned other Figure 4 in (a) and Figure 4 the first dimension corresponding to the ear region that is not marked in (b) can be measured and calculated according to the Figure 4 ear region in (b) thereof, which will not be elaborated here.

[0137] In some embodiments, the acquisition of the first user image and the method for determining the first dimension can adopt the acquisition method of data in a public database (such as the CIPIC database) for researching audio processing and human auditory models, which will not be specifically elaborated here.

[0138] Thus, the mobile phone can obtain the target head-related transfer function according to the first dimension, the first head-related transfer function, and the first position.

[0139] In a specific implementation manner, the first position includes a set of first angle values, and this set of first angle values includes the angle value of an azimuth angle (azim) and / or the angle value of an elevation angle (elve).

[0140] Among them, the angle value of the azimuth angle is used to represent the included angle on the horizontal plane between the line connecting the position of the audio sound source in the simulated three-dimensional space and the target user, and the extension line of the front of the target user. The angle value of the elevation angle is used to represent the included angle on the vertical plane between the line connecting the position of the audio sound source in the simulated three-dimensional space and the target user, and the extension line of the front of the target user.

[0141] For example, the angular value of the azimuth angle is 0 degrees, and the angular value of the elevation angle is 20 degrees.

[0142] After receiving the operation of the user triggering the control 1001 shown in (b) of Figure 4 , the side image of the user can be collected.

[0143] Based on this, the mobile phone can obtain the front image and the side image included in the first user image. Thus, the mobile phone can start to determine the first size of the head region and / or the ear region of the target user, and according to the first size, the first head-related transfer function, and a set of first angular values, obtain a set of head-related transfer functions, and this set of head-related transfer functions is the target head-related transfer function.

[0144] After obtaining the target head-related transfer function, the mobile phone can change the interface 12 shown in (b) of Figure 4 to display the interface 13 shown in (a) of Figure 5 , and the interface 13 is used to display the generation status of the head-related transfer function.

[0145] In another specific implementation manner, the first position includes M sets of first angular values, and the M sets of first angular values include the angular values of M azimuth angles and / or the angular values of M elevation angles.

[0146] After receiving the operation of the user triggering the control 1001 shown in (b) of Figure 4 , the mobile phone can start to determine the first size of the head region and / or the ear region of the target user, and according to the first size, the first head-related transfer function, and M sets of first angular values, obtain M sets of head-related transfer functions.

[0147] Among them, the M sets of head-related transfer functions correspond to the M sets of first angular values one by one.

[0148] Based on this, the mobile phone can change the interface 12 shown in (b) of Figure 4 to display the interface 14 shown in (b) of Figure 5 , and the interface 14 is used to display the generation status of the head-related transfer function, and the control 1002, and the control 1002 is used to trigger the selection of the M sets of first angular values.

[0149] After receiving the operation of the user triggering the control 1002 shown in (b) of Figure 5 , the mobile phone can change the interface 14 shown in (b) of Figure 5 to display the interface 15 shown in Figure 6 .

[0150] Among them, the interface 15 may include: an azimuth angle value adjustment box 1003 and an elevation angle value adjustment box 1004. The azimuth angle value adjustment box 1003 is used to adjust the angle value of the azimuth angle, and the elevation angle value adjustment box 1004 is used to adjust the angle value of the elevation angle.

[0151] After receiving the operation in which the user triggers the azimuth angle value adjustment box 1003 shown in Figure 6 , the mobile phone can adjust the angle value of the azimuth angle; after receiving the operation in which the user triggers the elevation angle value adjustment box 1004 shown in Figure 6 , the mobile phone can adjust the angle value of the elevation angle.

[0152] For example, when receiving the operation in which the user adjusts the angle value of the azimuth angle to 0 degrees through the azimuth angle value adjustment box 1003 and adjusts the angle value of the elevation angle to 20 degrees through the elevation angle value adjustment box 1004, the mobile phone can determine the head-related transfer function corresponding to the azimuth angle value of 0 degrees and the elevation angle value of 20 degrees as the target head-related transfer function.

[0153] In another specific implementation manner, the first position includes M groups of first angle values, and the M groups of first angle values include M azimuth angle values and / or M elevation angle values.

[0154] After receiving the operation in which the user triggers the control 1001 shown in Figure 4 (b) herein, the mobile phone can collect the user's side image.

[0155] Based on this, the mobile phone can obtain the front image and the side image included in the first user image. Thus, the mobile phone can start to determine the first size of the head area and / or the ear area of the target user. After obtaining the first size, the mobile phone can change from the interface 12 shown in Figure 4 (b) herein to display the interface 15 shown in Figure 6 .

[0156] After receiving the operation in which the user triggers the azimuth angle value adjustment box 1003 shown in Figure 6 , the mobile phone can adjust the angle value of the azimuth angle; after receiving the operation in which the user triggers the elevation angle value adjustment box 1004 shown in Figure 6 , the mobile phone can adjust the angle value of the elevation angle.

[0157] Based on this, the mobile phone can determine the azimuth angle value and / or the elevation angle value according to the operations of the target user on the azimuth angle value adjustment frame 1003 and the elevation angle value adjustment frame 1004, so as to determine the second angle value. Thus, the mobile phone can obtain a set of head-related transfer functions according to the first size, the first head-related transfer function, and the second angle value, and this set of head-related transfer functions is the target head-related transfer function.

[0158] For example, when receiving the operation that the user adjusts the azimuth angle value to 0 degrees through the azimuth angle value adjustment frame 1003 and adjusts the elevation angle value to 20 degrees through the elevation angle value adjustment frame 1004, the mobile phone can determine that the first angle value includes the azimuth angle value of 0 degrees and the elevation angle value of 20 degrees. Thus, the mobile phone can obtain a set of head-related transfer functions according to the first size, the first head-related transfer function, and the second angle value, and this set of head-related transfer functions is the target head-related transfer function.

[0159] Based on the above description, the mobile phone can obtain the target head-related transfer function. Thus, the mobile phone can perform audio rendering on the first audio in the electronic device according to the target head-related transfer function to generate the second audio, and the second audio is adapted to the target user and has the sound effect that the sound source of the audio is at the first position in the simulated three-dimensional space.

[0160] In summary, the mobile phone combines the first size of the target user to obtain the target head-related impulse data adapted to the target user. Thus, when receiving the playback operation of the audio to be played, the mobile phone can render the audio to be played according to the target head-related impulse data adapted to the target user to generate the audio adapted to the target user and having the sound effect that the sound source of the audio is at the first position in the simulated three-dimensional space.

[0161] Based on the above-described scenario, the audio processing method provided in the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings and application scenarios.

[0162] Among them, the audio processing method provided in the embodiments of the present application can be applied to an audio system including an electronic device and a wearable device, and is executed by the electronic device.

[0163] Please refer to Figure 7 , Figure 7 which shows a schematic flow chart of the audio processing method provided in an embodiment of the present application.

[0164] As Figure 7 shown, the audio processing method provided in the present application may include:

[0165] S201. Obtain the first user image.

[0166] Among them, the first user image includes a target area, which refers to the image area corresponding to one or more parts of the target user's body.

[0167] In some embodiments, the target area includes the image areas corresponding to the head area, the ear area, and / or the torso area. The head area includes other head areas except the ear area.

[0168] It should be understood that the head area, ear area, and / or torso area of each user are different. When the electronic device generates audio based on the average HRTF and a first position, the position of the sound source in the generated audio may be different from the first position.

[0169] Based on this, the electronic device can collect the first user image of the target user including the head area, the ear area, and / or the torso area, which is convenient for the electronic device to obtain accurate audio in combination with the sizes corresponding to the head area, the ear area, and / or the torso area.

[0170] Among them, there are various situations for the target area, for example:

[0171] Situation 1: The target area includes the image area corresponding to the head area.

[0172] For Situation 1, the target area includes the head area, which is convenient for the electronic device to obtain an accurate head-related transfer function according to the head physiological parameters corresponding to the head area.

[0173] Situation 2: The target area includes the image area corresponding to the ear area.

[0174] For Situation 2, the target area includes the ear area, which is convenient for the electronic device to obtain the image area corresponding to the accurate head-related transfer function according to the ear physiological parameters corresponding to the ear area.

[0175] Situation 3: The target area includes the image areas corresponding to the head area and the ear area respectively.

[0176] For Situation 3, the target area includes the image areas corresponding to the head area and the ear area respectively, which is convenient for the electronic device to obtain a more accurate head-related transfer function by combining the head physiological parameters corresponding to the head area and the ear physiological parameters corresponding to the ear area.

[0177] It can be understood that, compared with Situation 1 and Situation 2, the electronic device can obtain a head-related transfer function with higher accuracy based on Situation 3.

[0178] Situation 4: The target area includes the image areas corresponding to the head area, the ear area, and the torso area respectively.

[0179] For Case 4, the target area includes the head area, the ear area, and the torso area, so that the electronic device can combine the head physiological parameters corresponding to the head area, the ear physiological parameters corresponding to the ear area, and the physiological parameters corresponding to the torso area to obtain a high-accuracy head-related transfer function.

[0180] It can be understood that, compared with Case 1, Case 2, and Case 3, the electronic device can obtain a head-related transfer function with a higher accuracy rate based on Case 4.

[0181] Among them, the first user image may include a front image and a side image of the target user, and the side image includes an image on the right side and / or the left side of the target user.

[0182] In some embodiments, the electronic device may obtain the first user image including the head area and / or the ear area in the following manner:

[0183] The electronic device may display an image acquisition interface, and the image acquisition interface may display prompt information to guide the target user to capture an image including the head area and / or the ear area, so that the electronic device can obtain the first user image.

[0184] Among them, the first user image may include a front image and / or a side image of the target user, and the side image may include a side image and / or a left side image.

[0185] Exemplarily, the image acquisition interface may refer to Figure 4 the interface shown; the prompt information for reminding the user to capture a front image may refer to Figure 4 the prompt information in (a); the prompt information for reminding the user to capture a side image may refer to Figure 4 the prompt information in (b).

[0186] This application does not limit the specific implementation manner of the image acquisition interface; this application does not limit the specific implementation manner of the prompt information.

[0187] In some embodiments, when the electronic device plays audio for the first time, it may display an image acquisition interface and prompt the target user to acquire the first user image through the image acquisition interface.

[0188] In some other embodiments, when the electronic device is powered on for the first time, it may display an image acquisition interface and prompt the target user to acquire the first user image through the image acquisition interface.

[0189] In some other embodiments, at any time after the electronic device is powered on for the first time, it may display an image acquisition interface based on the operation of the user indicating to open the image acquisition interface, and prompt the target user to acquire the first user image through the image acquisition interface.

[0190] For the case of displaying an image acquisition interface for an operation of opening an image acquisition interface based on a user instruction, the electronic device may include a setting interface, and the setting interface may include an audio setting control. When receiving an operation of the user on the audio setting control, the electronic device may display the image acquisition interface.

[0191] S202. Determine the first size of the target area.

[0192] Combined with Case 1, when the target area includes the image area corresponding to the head area, the first size includes the first size corresponding to the head area.

[0193] Among them, the first size corresponding to the head area may include at least one of: head circumference, head width, head depth, head height, neck width, neck height, neck depth, anterior cranial offset, shoulder width, and shoulder circumference.

[0194] This application does not limit the specific parameters of the first size corresponding to the head area.

[0195] Combined with Case 2, when the target area includes the image area corresponding to the ear area, the first size includes the first size corresponding to the ear area.

[0196] Among them, the first size corresponding to the ear area may include at least one of: auricle height, auricle width, ear cavity height, ear cavity width, ear cavity depth, cymba conchae height, fossa of helix height, intertragic width, auricle rotation angle, auricle flare angle, auricle inferior offset, and auricle right offset.

[0197] This application does not limit the specific parameters of the first size corresponding to the ear area.

[0198] Combined with Case 3, when the target area includes the image areas corresponding to the head area and the ear area respectively, the first size includes the first size corresponding to the head area and the first size corresponding to the ear area.

[0199] Combined with Case 4, when the target area includes the image areas corresponding to the head area, the ear area, and the torso area respectively, the first size includes the first size corresponding to the head area, the first size corresponding to the ear area, and the first size corresponding to the torso area.

[0200] Among them, the first size corresponding to the torso area may include at least one of: torso width, torso height, torso depth, height, and sitting height.

[0201] In a specific embodiment, the target area includes the head area and the ear area, and the first size includes the head circumference and auricle height of the target user.

[0202] When the target area includes the head area and the ear area and the first dimension includes the head circumference and auricle height of the target user, it is convenient for the electronic device to obtain a more accurate head-related transfer function by combining the head circumference and auricle height of the target user.

[0203] S203. Obtain a target head-related transfer function according to the first dimension, the first head-related transfer function, and the first position.

[0204] Wherein, the first head-related transfer function is the mean value of the head-related transfer functions of N sample users pre-stored in the electronic device, and the first position is the position of the sound source in the simulated three-dimensional space pre-stored in the electronic device.

[0205] The simulated three-dimensional space can be understood as a space for simulating audio playback, and this space can be an unobstructed space.

[0206] It can be understood that the mean value of the head-related transfer functions of N sample users is obtained in advance. After obtaining the mean value of the head-related transfer functions of N sample users (i.e., mean HRTF), this average data can be stored in the electronic device for the electronic device to use when obtaining the target head-related transfer function (i.e., personalized HRTF).

[0207] Wherein, obtaining the mean value of the head-related transfer functions of N sample users can be performed by the electronic device executing the audio processing method of this application, or can be performed by another electronic device.

[0208] In some embodiments, the mean value of the head-related transfer functions of N sample users can be obtained through the following steps:

[0209] Step 2031. Obtain N sample user images. Each sample user image in the N sample user images includes the target area of the sample user, and N is a positive integer greater than or equal to 2.

[0210] Step 2032. Determine N groups of sample sizes corresponding to the N target areas.

[0211] Step 2033. Input the N sample parameters into the original generative diffusion network to obtain the mean value of the head-related transfer functions of N sample users.

[0212] Wherein, the N sample user images can include various types of user images such as male user images, female user images, elderly user images, and child user images.

[0213] Wherein, the specific implementation manner of the sample size can refer to the description of the above first dimension, and will not be specifically elaborated here.

[0214] It should be noted that for N sample user images, one sample user image can be input into the original generative diffusion network each time, and a sample head-related transfer function is output. Then, the average value is calculated with the previously obtained sample head-related transfer function, so as to obtain the mean value of the N sample head-related transfer functions.

[0215] For example, for the first time, sample image 1 is input into the original generative diffusion network to obtain sample head-related transfer function 1; for the second time, sample image 2 is input into the original generative diffusion network to obtain sample head-related transfer function 2, and the average data of sample head-related transfer function 1 and sample head-related transfer function 2 is obtained to get average data 1; for the third time, sample image 3 is input into the original generative diffusion network to obtain sample head-related transfer function 3, and the average data of sample head-related transfer function 3 and average data 1 is obtained to get average data 2, and so on, the mean value of the N sample head-related transfer functions can be obtained.

[0216] In some embodiments, the first position is determined according to the first angle value. According to the first size, the first head-related transfer function, and the first position, obtaining the target head-related transfer function includes:

[0217] Obtaining the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value.

[0218] Wherein, the first angle value includes the angle value of the azimuth angle and / or the angle value of the elevation angle. The angle value of the azimuth angle is used to indicate the direction of the sound source on the horizontal plane in the simulated three-dimensional space, and the angle value of the elevation angle is used to represent the direction of the sound source on the vertical plane in the simulated three-dimensional space.

[0219] It should be noted that the electronic device can obtain the target head-related impulse response function according to the first size, the first head-related transfer function, and the first position. By performing time-domain conversion on the target head-related impulse response function, the target head-related transfer function can be obtained.

[0220] Specifically, the angle value of the azimuth angle is used to represent the included angle between the line connecting the position of the sound source of the audio in the simulated three-dimensional space and the target user and the extension line of the front of the target user on the horizontal plane, and the angle value of the elevation angle is used to represent the included angle between the line connecting the position of the sound source of the audio in the simulated three-dimensional space and the target user and the extension line of the front of the target user on the vertical plane.

[0221] Specifically, in combination with the first angle value, the electronic device can obtain the target head-related transfer function corresponding to a more specific angle, thereby preparing data for generating a more accurate second audio for the electronic device.

[0222] For example, when the angular value of the azimuth is 0 degrees and the angular value of the elevation is 20 degrees, the angular value of the azimuth associated with the target head-related transfer function generated by the electronic device is 0 degrees, and the angular value of the elevation is 20 degrees.

[0223] Among them, the electronic device obtains the target head-related transfer function according to the first size, the first head-related transfer function, and the first angular value, which may include:

[0224] Input the first size, the first head-related transfer function, and the first position into the generative diffusion network to obtain the target head-related transfer function.

[0225] Among them, the generative diffusion network is used to generate a head-related transfer function adapted to the target user based on the average data corresponding to the head-related transfer functions of multiple users, combined with the first size and the first position of the target user.

[0226] Specifically, the electronic device can, based on the first head-related transfer function, construct a head-related transfer function of the target user that matches the target user in combination with the first head-related transfer function. The position associated with the head-related transfer function of the target user is the first position.

[0227] For example, when the first size includes the head circumference of the head region and the auricle height of the ear region, the angular value of the azimuth is 0 degrees, and the angular value of the elevation is 20 degrees, the electronic device inputs the user's head circumference, the user's auricle height, the angular value of the azimuth 0 degrees, the angular value of the elevation 20 degrees, and the first head-related transfer function into the generative diffusion network to obtain the target head-related transfer function.

[0228] Among them, the electronic device can obtain the generative diffusion network in advance and store the generative diffusion network in the electronic device and / or a storage device communicatively connected to the electronic device.

[0229] In some embodiments, the generative diffusion network can be pre-generated by the server and stored in the electronic device. Combined with Figure 8 , the generation process of the server generative diffusion network includes:

[0230] Step 2034: Obtain N sample user images. Each sample user image among the N sample user images includes the target region of the sample user. N is a positive integer greater than or equal to 2.

[0231] Step 2035: Determine N groups of sample sizes corresponding to the N target regions.

[0232] Step 2036: Input the N groups of sample sizes into the original generative diffusion network to obtain N sample head-related transfer functions.

[0233] Step 2037: Determine the average data corresponding to the N head-related transfer functions as the first head-related transfer function.

[0234] Step 2038: Train the original generative diffusion network according to the relationships between the N sample head-related transfer functions and the first head-related transfer function respectively to obtain a generative diffusion network.

[0235] Among them, the acquisition of the N sample user images and the method for determining the N groups of sample sizes corresponding to the N target regions can adopt the data acquisition method in a public database (such as the CIPIC database) for researching audio processing and human auditory models.

[0236] For any one of the N groups of sample sizes, the server can splice (concat) the multiple sample sizes included in a group of sample sizes and perform parameter analysis on the multiple sample sizes through the self-attention mechanism. Through analysis, it is found that some parameters are highly important for generating the head-related transfer function, and some parameters are less important for generating the head-related transfer function. Therefore, the multiple sample sizes can be sorted according to the importance degree, and the vector corresponding to the importance degree is input into the original generative diffusion network to obtain a sample head-related transfer function.

[0237] Based on the above description, it can be seen that the process of training the original generative diffusion network is to average the head-related transfer functions of multiple sample users (eliminating directionality, that is, personal characteristics) and learn the differences between the mean of the head-related transfer functions of multiple sample users and the head-related transfer function of each sample user. By calculating the loss between the mean of the head-related transfer functions of multiple sample users and the head-related transfer function of each sample user (the specific loss calculation can adopt L1norm), the original generative diffusion network can learn the transformation situation between the head-related transfer function of a sample user and the mean of the head-related transfer functions of multiple sample users.

[0238] The above principle is: For the predicted white noise corresponding to different diffusion steps, the original generative diffusion network can calculate the loss for the differences between the predicted white noise and the noise at the real t step and the head-related transfer function of each sample user and the head-related transfer functions of multiple sample users.

[0239] Since the process of training the original generative diffusion network can learn the transformation situation between the mean of the head-related transfer functions of multiple users and the head-related transfer function of each user, therefore, combined with Figure 9, in the electronic device, the generative diffusion network can restore the head-related transfer function of the target user based on the mean of the head-related transfer functions of multiple users, combined with the first size of the target user.

[0240] Specifically, combined with Figure 9 , the electronic device can splice multiple first sizes included in the first size of a group of target users, perform parameter analysis on the multiple first sizes through the self-attention mechanism, sort the multiple first sizes according to the importance level, and finally input the vector corresponding to the importance level into the generative diffusion network. At the same time, the electronic device can also encode the mean of the head-related transfer functions of multiple sample users through an encoder (concat), and input the first position into the generative diffusion network, so that the electronic device can obtain the target head-related transfer function.

[0241] Among them, when the number of the first angle values is M groups, and the M groups of first angle values are used to indicate M positions in the simulated three-dimensional space, and M is a positive integer greater than or equal to 2, the process of obtaining the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value can be implemented in two ways.

[0242] Implementation method 1: Obtaining the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value includes:

[0243] Obtaining M head-related transfer functions according to the first size, the first head-related transfer function, and M groups of first angle values, and the M head-related transfer functions are respectively associated with the M groups of first angle values; determining the second head-related transfer function among the M head-related transfer functions as the target head-related transfer function, and the first head-related transfer function is the head-related transfer function corresponding to a group of first angle values selected based on the operation of the target user among the M head-related transfer functions.

[0244] Specifically, after the electronic device obtains M head-related transfer functions according to the first size, the first head-related transfer function, and M groups of first angle values, it can display an angle selection interface. The angle selection interface can display an azimuth angle value adjustment selection box and an elevation angle value adjustment box. Through the two angle value adjustment boxes, the target user can select the angle values of the azimuth angle and the elevation angle. Thus, the electronic device can select the head-related transfer function corresponding to a group of first angle values selected by the target user's operation.

[0245] Exemplarily, the angle selection interface can refer to Figure 6 the interface shown; the angle selection box for the azimuth angle and the angle selection box for the elevation angle can refer to Figure 6 the description of the azimuth angle value adjustment box 1003 and the elevation angle value adjustment box 1004 in, and will not be specifically described here.

[0246] Implementation method 2: Obtaining a target head-related transfer function according to a first dimension, a first head-related transfer function, and a first angle value, including:

[0247] According to the first dimension, the first head-related transfer function, and a second angle value, obtain a third head-related transfer function, where the second angle value includes a group of first angle values selected based on the operation of the target user from M groups of first angle values; determine the third head-related transfer function as the target head-related transfer function.

[0248] Specifically, after the electronic device obtains the first dimension, it can display an angle selection interface. The angle selection interface can display an azimuth angle value adjustment box and an elevation angle value adjustment box. Through the two angle value adjustment boxes, the target user can select the angle values of the azimuth angle and the elevation angle. Thus, the electronic device can determine a group of angle values selected by the target user as the second angle value. Based on this, the electronic device can obtain a third head-related transfer function according to the first dimension, the first head-related transfer function, and the second angle value.

[0249] Exemplarily, for the angle selection interface, reference can be made to Figure 6 the interface shown; for the angle selection box of the azimuth angle and the angle selection box of the elevation angle, reference can be made to Figure 6 the description of the azimuth angle value adjustment box 1003 and the elevation angle value adjustment box 1004 in

[0250] In some other embodiments, the first position is determined according to the first angle value and a preset distance. Obtaining a target head-related transfer function according to the first dimension, the first head-related transfer function, and the first position, including:

[0251] Obtain a target head-related transfer function according to the first dimension, the first head-related transfer function, the first angle value, and the preset distance.

[0252] Wherein, the preset distance is preset and stored in the electronic device, and the preset distance is the distance between the sound source and the target user.

[0253] Specifically, by combining the first angle value and the preset distance, the electronic device can obtain a more specific angle and the target head-related transfer function corresponding to the specific distance, thereby preparing data for generating a more accurate second audio for the electronic device.

[0254] For example, when the angle value of the azimuth angle is 0 degree, the angle value of the elevation angle is 20 degrees, and the preset distance is 10 meters, the sound source position of the second audio generated by the electronic device is at the position with an angle value of 0 degree, an elevation angle value of 20 degrees, and a distance of 10 meters.

[0255] Among them, for the process in which the electronic device inputs the first size, the first head-related transfer function, the first angle value, and the preset distance into the generative diffusion network to obtain the target head-related transfer function, reference may be made to the process in which the electronic device inputs the first size, the first head-related transfer function, and the first angle value into the generative diffusion network to obtain the target head-related transfer function, which will not be elaborated here.

[0256] S204. In response to the play operation on the first audio, according to the target head-related transfer function, perform audio rendering on the first audio in the electronic device to generate a second audio.

[0257] Among them, the second audio is an audio with the first sound effect, and the perceived position of the sound source of the second audio in the first sound effect is the first position in the simulated three-dimensional space.

[0258] Among them, the electronic device can receive the play operation on the first audio in various ways.

[0259] In some embodiments, the first electronic device includes an icon of a music player application. After receiving an operation (such as a click, double-click, or long-press operation, etc.) triggered by the target user on the icon of the music player application, the first electronic device can display the interface of the music player application, and the music player application includes the first audio. The target user can perform the play operation on the first audio on the interface of the music player application, so that the electronic device can receive the play operation on the first audio.

[0260] In other embodiments, a voice assistant is added to the first electronic device. After receiving a specific wake-up word input by the user's voice, the voice assistant can be woken up. And after receiving an instruction input by the user's voice to open the music player application and play the first audio, the first electronic device can start the music player application and receive the play operation on the first audio.

[0261] For example, the voice assistant can be YOYO. After receiving the user's voice input of "Hello, YOYO", YOYO can be woken up, and after receiving the voice instruction input by the user "Open the music player application and play the first audio", the first electronic device can open the music player application and receive the play operation on the first audio.

[0262] This application does not limit the specific implementation manner of receiving the play operation on the first audio.

[0263] Among them, audio rendering is a process for processing the original audio signal into a signal suitable for playing.

[0264] When receiving a user's playback operation on a first audio, the electronic device can perform audio rendering on the first audio in the electronic device according to the target head-related transfer function to generate a second audio.

[0265] Based on the above description, the target head-related transfer function is obtained by combining the physiological parameters of the head and / or ears of the target user. Therefore, by performing audio rendering on the first audio in the electronic device through the target head-related transfer function, the electronic device can obtain a second audio that is adapted to the target user and whose sound source is at a first position in a simulated three-dimensional space.

[0266] Based on the above description, in some embodiments, the first position is determined according to a first angle value, the first angle value includes the angle value of the azimuth angle and the angle value of the elevation angle, and the target head-related transfer function associates the angle value of the azimuth angle and the angle value of the elevation angle.

[0267] For example, when the angle value of the azimuth angle is 0 degrees and the angle value of the elevation angle is 20 degrees, the second audio generated by the electronic device is an audio with an azimuth angle of 0 degrees and an elevation angle of 20 degrees.

[0268] In other embodiments, the first position is determined according to a first angle value, the first angle value includes the angle value of the azimuth angle and the angle value of the elevation angle, and the target head-related transfer function associates the angle value of the azimuth angle, the angle value of the elevation angle, and the distance.

[0269] For example, the first position is determined according to a first angle value, the first angle value includes the angle value of the azimuth angle and the angle value of the elevation angle. When the angle value of the azimuth angle is 0 degrees, the angle value of the elevation angle is 20 degrees, and the preset distance is 10 meters, the second audio generated by the electronic device is an audio with an azimuth angle of 0 degrees, an elevation angle of 20 degrees, and a distance of 10 meters from the target user.

[0270] In addition, since the electronic device is applied to an audio system, the audio system includes the electronic device and a wearable device. After the electronic device obtains the second audio, it can send the second audio to the wearable device, so that when the wearable device receives the second audio, it can play the second audio. Based on this, the target user can hear the second audio.

[0271] In some embodiments, the wearable device is a headset, the target head-related transfer function includes the head-related transfer function corresponding to the left ear and the head-related transfer function corresponding to the right ear, the second audio generated by the electronic device includes the audio corresponding to the left ear and the audio corresponding to the right ear, and when the electronic device sends the second audio to the headset, it can send the audio corresponding to the left ear to the headset corresponding to the left ear and send the audio corresponding to the right ear to the headset corresponding to the right ear.

[0272] When it is determined at the first position according to the first angle value, the audio corresponding to the left ear and the audio corresponding to the right ear can be processed through the following formula:

[0273] outL = s * hrir_l(angle); outR = s * hrir_r(angle).

[0274] Among them, s is used to represent the audio signal corresponding to the first audio, hrir_l is used to represent the head-related transfer function corresponding to the left ear, hrir_r represents the head-related transfer function corresponding to the right ear, and angle is used to represent the first angle value.

[0275] For example, when the first angle value is determined according to the angle value of the azimuth angle, and the angle value of the azimuth angle is 0 degrees, that is, the sound source is directly in front of the target user, then the audio corresponding to the left ear and the audio corresponding to the right ear can be processed through the following formula:

[0276] outL = s * hrir_l(azim, 0); outR = s * hrir_r(azim, 0).

[0277] It should be understood that when the electronic device performs audio rendering processing on the first audio based on the target head-related transfer function, the gain can also be adjusted, that is, gain attenuation processing is performed, so that the amplitude of the audio signal is appropriately reduced to avoid sound distortion or damage to the wearable device caused by too large an amplitude of the audio signal.

[0278] The gain attenuation processing formula is as follows: factor = 1 / (1 + x 2 ), where factor is used to represent the multiple of the audio signal gain attenuation, and x is used to represent the preset distance. The larger the value of x, the smaller the factor. Therefore, the volume of the audio signal corresponding to the first audio can be controlled by adjusting the value of x.

[0279] When adding gain attenuation processing, the audio corresponding to the left ear and the audio corresponding to the right ear can be processed through the following formula:

[0280] outL = factor * s * hrir_l(angle); outR = factor * s * hrir_r(angle).

[0281] The audio processing method of the present application enables the electronic device to determine the first size corresponding to the target area by obtaining the first user image. Since the first user image includes the target area of the target user, and the target area includes the head area and / or the ear area, the electronic device's determination of the first size of the head area and / or the ear area can prepare data for generating a target head-related transfer function that matches the target user. Based on this, the electronic device can obtain a target head-related transfer function that matches the target user according to the first size, the first head-related transfer function, and the first position, and perform audio rendering on the first audio in the electronic device according to the target head-related transfer function to generate a second audio that is adapted to the target user and has the sound source at the first position in the simulated three-dimensional space. Thus, the electronic device can overcome the problem of physiological differences and generate audio adapted to the target user, facilitating the target user to obtain a better spatial sound experience.

[0282] Based on the above Figure 7 description of S203, considering that the second audio obtained by performing audio rendering on the first audio according to the target head-related transfer function generated in S203 may not necessarily enable the target user to obtain an extreme sound experience, therefore, after obtaining the target head-related transfer function, the electronic device can also correct the target head-related transfer function to obtain a corrected target head-related transfer function.

[0283] Next, in combination with Figure 10 , the specific implementation process of the audio processing method of the present application will be introduced in detail.

[0284] Among them, the electronic device is applied to an audio system, and the audio system includes the electronic device and a wearable device. After the electronic device obtains the second audio, it can send the second audio to the wearable device so that the wearable device plays the second audio for the target user.

[0285] Please refer to Figure 10 , Figure 10 which shows a schematic flowchart of the audio processing method provided by an embodiment of the present application.

[0286] As Figure 10 shown, the audio processing method provided by the present application may include:

[0287] S301. Obtain a first user image, where the first user image includes a target area of a target user, and the target area includes a head area and / or an ear area.

[0288] S302. Determine the first size of the target area.

[0289] S303. Obtain a target head-related transfer function based on the first dimension, the first head-related transfer function, and the first position. The first head-related transfer function is the mean of the head-related transfer functions of N sample users pre-stored in the electronic device, and the first position is a position in the pre-stored simulated three-dimensional space in the electronic device.

[0290] Among them, S301, S302, and S303 are respectively similar to the implementation manners of S201, S202, and S203 in the Figure 7 illustrated embodiment, and will not be elaborated here.

[0291] S304. Correct the target head-related transfer function to obtain a corrected target head-related transfer function.

[0292] Among them, the target head-related transfer function is associated with the first position, and the corrected target head-related transfer function is associated with the second position. The second position is a correction position for making the perceived position of the sound source of the audio be at the first position.

[0293] It should be understood that the target head-related transfer function is generated by the generative diffusion network according to the first dimension of the target region in the first user image and is physical data. The sound source position of the rendered audio obtained by performing audio rendering according to the target head-related transfer function is also physical data. There are usually differences between physical data and the data perceived by the user. Then, when the first audio is rendered according to the target head-related transfer function to obtain the second audio, the position of the sound source of the second audio perceived by the target user may not be the first position, resulting in the target user not being able to experience the best sound.

[0294] For example, the first position is determined according to the first angle value. When the angle value of the azimuth included in the first angle value is 10 degrees, the first angle value associated with the target head-related transfer function includes the angle value of the azimuth of 10 degrees. After performing audio rendering on the audio according to the corrected target head-related transfer function, the angle value of the azimuth of the sound source heard by the target user may be 20 degrees.

[0295] Based on this, the electronic device corrects the target head-related transfer function to obtain a second angle value associated with the corrected target head-related transfer function, where the angle value of the azimuth included is 0 degrees. Then, after performing audio rendering on the audio according to the target head-related transfer function, the angle value of the azimuth of the sound source heard by the target user may be 10 degrees, making the angle value of the azimuth of the sound source heard by the target user the same as the first angle value.

[0296] For example, the first position is determined according to the first angle value. When the azimuth angle value included in the first angle value is 20 degrees, the first angle value associated with the target head-related transfer function includes the azimuth angle value of 20 degrees. After the audio is rendered according to the target head-related transfer function, the azimuth angle value of the sound source heard by the target user may be 30 degrees.

[0297] Based on this, the electronic device corrects the target head-related transfer function, and the second angle value associated with the corrected target head-related transfer function includes the azimuth angle value of 10 degrees. Then, after the audio is rendered according to the corrected target head-related transfer function, the azimuth angle value of the sound source heard by the target user may be 20 degrees, so that the azimuth angle value of the sound source heard by the target user is the same as the first angle value.

[0298] Based on this, after obtaining the target head-related transfer function, the electronic device can correct the target head-related transfer function, so that after the audio is rendered according to the target head-related transfer function subsequently, the position of the sound source perceived by the target user is the first position.

[0299] In some embodiments, correcting the target head-related transfer function to obtain the corrected target head-related transfer function includes:

[0300] Inputting the target head-related transfer function into a neural network model to obtain the corrected target head-related transfer function.

[0301] Wherein, the neural network model is used to predict the relationship between the position associated with the head-related transfer function and the perceived position of the sound source in the sound effect of the audio rendered by the head-related transfer function.

[0302] Specifically, after obtaining the target head-related transfer function, the electronic device can input the target head-related transfer function into the neural network model, and the neural network model can predict the true position of the sound source in the sound effect of the audio rendered by the head-related transfer function according to the first position associated with the target head-related transfer function.

[0303] Wherein, the electronic device can obtain the neural network model in advance and store the neural network model in the electronic device and / or a storage device communicatively connected to the electronic device.

[0304] In some embodiments, the neural network model can be pre-generated by the server and stored in the electronic device. Combined with Figure 11 , the generation process of the neural network model includes:

[0305] Step 3031: Obtain P sample head-related transfer functions, P sample positions, and P sample auditory perception positions. The P sample head-related transfer functions correspond one-to-one with the P sample positions and the P sample auditory perception positions respectively. The P sample auditory perception positions are respectively used to indicate the true auditory perception positions of sound sources in the sound effects of the audio rendered by the P sample head-related transfer functions, where P is a positive integer greater than or equal to 2.

[0306] Step 3032: Input the P sample head-related transfer functions and the P sample positions into the original neural network model to obtain P predicted positions, which are respectively used to indicate the predicted auditory perception positions of sound sources in the sound effects of the audio rendered by the P sample head-related transfer functions.

[0307] Step 3033: Train the original neural network model according to the differences between the P sample auditory perception positions and the corresponding P predicted positions to obtain the neural network model.

[0308] In some embodiments, combined with Figure 11 , the sample position is determined according to the sample angle value, and the sample angle value includes the angle value of the azimuth angle and / or the angle value of the elevation angle. The server can input the P sample head-related transfer functions and the P sample angle values into the original neural network model, and through classification by a classifier (softmax), obtain P maximum angle values (maxangle), that is, P predicted angle values.

[0309] For example, the value of P can be 12, and the server can train the original neural network model through 12 sample head-related transfer functions, 12 sample positions, 12 sample auditory perception positions, and 12 predicted positions.

[0310] Based on the above description, it can be seen that the process of training the original neural network model is the process of learning the differences between the predicted positions and the sample auditory perception positions. Through this process, the original neural network model can learn the differences between the given sample positions and the predicted positions.

[0311] Since the training process of the original neural network model can learn the differences between the predicted positions and the sample auditory perception positions, therefore, the neural network model in the electronic device can obtain the corrected target head-related transfer function based on the target head-related transfer function.

[0312] In some embodiments, the sample position is determined by a sample angle value, and the sample auditory perception position is determined according to the sample auditory perception angle value feedback by the sample user. By playing the audio rendered through the sample head-related transfer function, the sample auditory perception angle value feedback by the sample user is determined. If the sample auditory perception angle value is within the corresponding range of the currently given sample angle value, it is considered that the audio generation quality meets the quality condition; otherwise, the current difference is simulated through a neural network model.

[0313] For example, 15 sample angle values can be obtained to get the sample auditory perception angle value feedback by the sample user (the feedback angle value of the user can be obtained while recording the HRTF).

[0314] In addition, when simulating the current difference through a neural network model, the neural network model can be used for prediction to obtain a predicted angle value. Subsequently, the mapping relationship between the predicted angle value of the corresponding point and the sample auditory perception angle value is learned, and the weighting coefficient is determined, which is convenient for calibration through the weighting coefficient when calibrating the target head-related transfer function, so that the generated HRTF is more consistent with the subjective feeling of the user by customizing the corresponding HRTF.

[0315] Illustratively, when the angle value of the azimuth angle included in the first angle value associated with the target head-related transfer function is 10 degrees, before calibration, after audio rendering of the first audio according to the target head-related transfer function, the angle value of the azimuth angle of the sound source of the second audio perceived by the target user may be 20 degrees.

[0316] After the electronic device obtains the target head-related transfer function, the target head-related transfer function is input into the neural network model. When the weighting coefficient determined by the neural network model is 10 degrees, the angle value of the azimuth angle included in the second angle value associated with the calibrated target head-related transfer function obtained by the electronic device may be 0 degrees.

[0317] Based on this, after audio rendering of the first audio according to the calibrated target head-related transfer function, the angle value of the azimuth angle of the sound source of the second audio heard by the target user may be 10 degrees, and the angle value of the azimuth angle of the sound source heard by the target user is the same as the first angle value.

[0318] S305. According to the calibrated target head-related transfer function, perform audio rendering on the first audio in the electronic device to generate a third audio.

[0319] Wherein, the third audio is an audio adapted to the target user and having a second sound effect. The similarity between the auditory perception position of the sound source in the second sound effect and the first position is higher than the similarity between the auditory perception position of the sound source in the first sound effect and the first position.

[0320] Wherein, S305 is respectively associated withFigure 7 The implementation of S204 in the illustrated embodiment is similar and will not be elaborated here.

[0321] S306. Send the third audio to the wearable device so that the wearable device plays the third audio for the target user.

[0322] Since the electronic device is applied to the audio system, and the audio system includes the electronic device and the wearable device, after the electronic device obtains the third audio, it can send the third audio to the wearable device, so that when the wearable device receives the third audio, it can play the third audio. Based on this, the target user can hear the third audio.

[0323] In this application, after the electronic device obtains the target head-related transfer function, it can also correct the target head-related transfer function to obtain the corrected target head-related transfer function. In this way, the electronic device can perform audio rendering on the first audio in the electronic device according to the corrected target head-related transfer function to generate the third audio that matches the target user and has the second sound effect. Since the corrected target head-related transfer function is associated with the second position, and the second position is the correction position for making the perceived position of the sound source of the audio be at the first position, then, performing audio rendering on the first audio in the electronic device according to the corrected target head-related transfer function to generate the third audio can make the position accuracy of the sound source of the third audio perceived by the target user high, and can make the perceived position of the sound source of the second audio be at the first position, avoiding the situation that the position of the sound source of the second audio perceived by the user is far from the first position, and can ensure the playback quality of the third audio.

[0324] Moreover, after the electronic device obtains the third audio, it can send the third audio to the wearable device, so that when the wearable device receives the third audio, it can play the third audio. Based on this, the target user can hear the third audio.

[0325] Exemplarily, this application provides an audio processing device.

[0326] Please refer to Figure 12 , Figure 12 which shows a schematic block diagram of the audio processing device provided by an embodiment of this application.

[0327] As Figure 12 shown, the audio processing device 400 can exist independently or be integrated in other devices, and can communicate with the above-mentioned first electronic device to implement the operations corresponding to the first electronic device in any of the above method embodiments. The audio processing device 400 provided by the embodiments of this application includes: an acquisition module 401, a determination module 402, a obtaining module 403, and a generation module 404.

[0328] An acquisition module 401, configured to acquire a first user image, where the first user image includes a target area, and the target area refers to an image area corresponding to one or more parts of the target user's body.

[0329] A determination module 402, configured to determine a first size of the target area.

[0330] An obtaining module 403, configured to obtain a target head-related transfer function according to the first size, a first head-related transfer function, and a first position. The first head-related transfer function is the mean of the head-related transfer functions of N sample users pre-stored in the electronic device, and the first position is the position of a sound source in a pre-stored simulated three-dimensional space in the electronic device. The target head-related transfer function is used to indicate the head-related transfer function of the target user.

[0331] A generation module 404, configured to, in response to a play operation on a first audio, perform audio rendering on the first audio in the electronic device according to the target head-related transfer function to generate a second audio. The second audio is an audio with a first sound effect, and the perceived position of the sound source in the first sound effect of the second audio is the first position in the simulated three-dimensional space.

[0332] In some embodiments, the obtaining module 403 is specifically configured to:

[0333] Correct the target head-related transfer function to obtain a corrected target head-related transfer function. The target head-related transfer function is associated with the first position, and the corrected target head-related transfer function is associated with a second position. The second position is a correction position for making the perceived position of the sound source of the audio be at the first position.

[0334] In some embodiments, the generation module 403 is specifically configured to:

[0335] Perform audio rendering on the first audio in the electronic device according to the corrected target head-related transfer function to generate a third audio. The third audio is an audio with a second sound effect, and the similarity between the perceived position of the sound source in the second sound effect and the first position is higher than the similarity between the perceived position of the sound source in the first sound effect and the first position.

[0336] In some embodiments, the obtaining module 403 is specifically configured to:

[0337] Input the target head-related transfer function into a neural network model to obtain a corrected target head-related transfer function. The neural network model is used to predict the relationship between the position associated with the head-related transfer function and the perceived position of the sound source in the sound effect of the audio rendered by the head-related transfer function.

[0338] In some embodiments, the audio processing apparatus further includes a network processing module. The network processing module is specifically configured to:

[0339] Obtain P head-related transfer functions of samples, P sample positions, and P perceived listening positions of samples. The P head-related transfer functions of samples correspond one-to-one with the P sample positions and the P perceived listening positions of samples respectively. The P perceived listening positions of samples are respectively used to indicate the true perceived listening positions of sound sources in the sound effects of the audio rendered by the P head-related transfer functions of samples, where P is a positive integer greater than or equal to 2; input the P head-related transfer functions of samples and the P sample positions into the original neural network model to obtain P predicted positions, and the P predicted positions are respectively used to indicate the predicted perceived listening positions of sound sources in the sound effects of the audio rendered by the P head-related transfer functions of samples; train the original neural network model according to the differences between the P sample positions and the corresponding P predicted positions to obtain the neural network model.

[0340] In some embodiments, the first position is determined according to the first angle value. The obtaining module 403 is specifically configured to:

[0341] According to the first size, the first head-related transfer function, and the first angle value, obtain the target head-related transfer function. The first angle value includes the angle value of the azimuth angle and / or the angle value of the elevation angle. The angle value of the azimuth angle is used to represent the included angle between the connection line of the sound source of the audio in the simulated three-dimensional space and the target user, and the extension line of the front of the target user on the horizontal plane. The angle value of the elevation angle is used to represent the included angle between the connection line of the sound source of the audio in the simulated three-dimensional space and the target user, and the extension line of the front of the target user on the vertical plane.

[0342] In some embodiments, the number of the first angle values is M groups, and the M groups of first angle values are used to indicate M positions in the simulated three-dimensional space. M is a positive integer greater than or equal to 2. The obtaining module 403 is specifically configured to:

[0343] According to the first size, the first head-related transfer function, and the M groups of first angle values, obtain M head-related transfer functions. The M head-related transfer functions are respectively associated with the M groups of first angle values; determine the second head-related transfer function among the M head-related transfer functions as the target head-related transfer function. The first head-related transfer function is the head-related transfer function corresponding to a group of first angle values selected based on the operation of the target user among the M head-related transfer functions.

[0344] In some embodiments, the number of the first angle values is M groups, and the M groups of first angle values are used to indicate M positions in the simulated three-dimensional space. M is a positive integer greater than or equal to 2. The obtaining module 403 is specifically configured to:

[0345] According to the first size, the first head-related transfer function, and the second angle value, obtain the third head-related transfer function. The second angle value includes a group of first angle values selected based on the operation of the target user among the M groups of first angle values; determine the third head-related transfer function as the target head-related transfer function.

[0346] In some embodiments, the obtaining module 403 is specifically configured to:

[0347] Input the first size, the first head-related transfer function, and the first position into the generative diffusion network to obtain the target head-related transfer function. The generative diffusion network is used to generate a head-related transfer function adapted to the target user based on the average data corresponding to the head-related transfer functions of multiple users, in combination with the first size and the first position of the target user.

[0348] In some embodiments, the audio processing device further includes a network processing module. The network processing module is specifically configured to:

[0349] Obtain N sample user images, where each sample user image in the N sample user images includes a target area, and N is a positive integer greater than or equal to 2; determine N groups of sample sizes corresponding to the N target areas; input the N groups of sample sizes into the original generative diffusion network to obtain N sample head-related transfer functions; determine the average data corresponding to the N head-related transfer functions as the first head-related transfer function; and train the original generative diffusion network according to the relationships between the N sample head-related transfer functions and the first head-related transfer function to obtain the generative diffusion network.

[0350] In some embodiments, the first size corresponding to the head region includes at least one of head circumference, head width, head depth, head height, neck width, neck height, neck depth, anterior cranial offset, shoulder width, and shoulder circumference; the first size corresponding to the ear region includes at least one of auricle height, auricle width, ear cavity height, ear cavity width, ear cavity depth, cymba conchae height, cochlear height, intertragal width, auricle rotation angle, auricle flare angle, auricle inferior offset, and auricle right offset; the first size corresponding to the torso region includes at least one of torso width, torso height, torso depth, height, and sitting height.

[0351] In some embodiments, the obtaining module 403 is specifically configured to: The electronic device is applied to an audio system, and the audio system includes the electronic device and a wearable device. After audio rendering the first audio in the electronic device according to the target head-related transfer function to generate the second audio, the method further includes:

[0352] Send the second audio to the wearable device so that the wearable device plays the second audio for the target user.

[0353] Exemplarily, the present application provides an electronic device, including a processor; when the processor executes the computer code or instructions in the memory, the electronic device is caused to execute the audio processing method in the foregoing embodiments.

[0354] Exemplarily, the present application provides an electronic device, including one or more processors; a memory; and one or more computer programs, where the one or more computer programs are stored on the memory, and when the computer programs are executed by the one or more processors, the electronic device is enabled to execute the audio processing method in the foregoing embodiments.

[0355] It can be understood that, in order for the electronic device to implement the above functions, it includes corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the manner of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.

[0356] In this embodiment, the electronic device can be divided into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0357] In the case of dividing each functional module corresponding to each function, the electronic device involved in the above embodiment may further include: an acquisition module, a determination module, a obtaining module, and a generation module. Among them, the acquisition module, the determination module, the obtaining module, and the generation module cooperate with each other, and can be used to support the electronic device to execute the above steps, and / or for other processes of the technology described herein.

[0358] It should be noted that all relevant contents of the steps involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be repeated here.

[0359] The electronic device provided in this embodiment is used to execute the above audio processing method, and thus can achieve the same effect as the above implementation method.

[0360] Exemplarily, the present application provides a chip system, the chip system includes a processor, which is used to call and run a computer program from a memory, so that an electronic device installed with the chip system executes the audio processing method in the foregoing embodiments.

[0361] Exemplarily, the present application provides a computer-readable storage medium storing code or instructions, which, when run on an electronic device, cause the electronic device to implement the audio processing method in the foregoing embodiments when executed.

[0362] Exemplarily, the present application provides a computer program product, which, when run on a computer, causes the electronic device to implement the audio processing method in the foregoing embodiments.

[0363] Among them, the electronic device, computer-readable storage medium, computer program product, or chip system provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0364] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0365] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0366] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0367] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0368] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs. The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An audio processing method, characterized in that, The method is applied to an electronic device, and the method includes: Obtain a first user image, where the first user image includes a target area, and the target area refers to an image area corresponding to one or more parts of the target user's body; Determine a first size of the target area; According to the first size, a first head-related transfer function, and a first position, obtain a target head-related transfer function, where the first head-related transfer function is the mean of the head-related transfer functions of N sample users pre-stored in the electronic device, the first position is the position of a sound source in a pre-stored simulated three-dimensional space in the electronic device, and the target head-related transfer function is used to indicate the head-related transfer function of the target user; In response to a play operation on a first audio, according to the target head-related transfer function, perform audio rendering on the first audio in the electronic device to generate a second audio, where the second audio is an audio with a first sound effect, and the perceived position of the sound source of the second audio is the first position in the simulated three-dimensional space.

2. The method according to claim 1, characterized in that, The method further includes: Correct the target head-related transfer function to obtain a corrected target head-related transfer function. The target head-related transfer function is associated with the first position, and the corrected target head-related transfer function is associated with a second position, where the second position is a correction position for making the perceived position of the sound source of the audio be at the first position; Performing audio rendering on the first audio in the electronic device according to the target head-related transfer function to generate a second audio includes: Performing audio rendering on the first audio in the electronic device according to the corrected target head-related transfer function to generate a third audio, where the third audio is an audio with a second sound effect, and the similarity between the perceived position of the sound source in the second sound effect and the first position is higher than the similarity between the perceived position of the sound source in the first sound effect and the first position.

3. The method according to claim 2, wherein The correcting the target head-related transfer function to obtain a corrected target head-related transfer function includes: Input the target head-related transfer function into a neural network model to obtain the corrected target head-related transfer function, where the neural network model is used to predict the relationship between the position associated with the head-related transfer function and the perceived position of the sound source in the sound effect of the audio rendered by the head-related transfer function.

4. The method according to claim 3, characterized in that, The generation process of the neural network model includes: Obtain P sample head-related transfer functions, P sample positions, and P sample perceived positions. The P sample head-related transfer functions are respectively in one-to-one correspondence with the P sample positions and the P sample perceived positions. The P sample perceived positions are respectively used to indicate the true perceived positions of the sound sources in the sound effects of the audio rendered by the P sample head-related transfer functions, and P is a positive integer greater than or equal to 2; Input the P sample head-related transfer functions and the P sample positions into an original neural network model to obtain P predicted positions, where the P predicted positions are respectively used to indicate the predicted perceived positions of the sound sources in the sound effects of the audio rendered by the P sample head-related transfer functions; Training the original neural network model according to the differences between the P perceptual positions of the P samples and the corresponding P predicted positions to obtain the neural network model.

5. The method according to any one of claims 1 to 4, characterized in that, The first position is determined according to a first angle value. The obtaining of the target head-related transfer function according to the first size, the first head-related transfer function, and the first position includes: Obtaining the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value. The first angle value includes the angle value of the azimuth angle and / or the angle value of the elevation angle. The angle value of the azimuth angle is used to represent the included angle on the horizontal plane between the line connecting the sound source of the audio in the simulated three-dimensional space and the target user and the extension line of the front of the target user. The angle value of the elevation angle is used to represent the included angle on the vertical plane between the line connecting the sound source of the audio in the simulated three-dimensional space and the target user and the extension line of the front of the target user.

6. The method according to claim 5, wherein The number of the first angle values is M groups. The M groups of first angle values are used to indicate M positions in the simulated three-dimensional space. M is a positive integer greater than or equal to 2. The obtaining of the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value includes: Obtaining M head-related transfer functions according to the first size, the first head-related transfer function, and the M groups of first angle values. The M head-related transfer functions are respectively associated with the M groups of first angle values; Determining the second head-related transfer function among the M head-related transfer functions as the target head-related transfer function. The first head-related transfer function is the head-related transfer function corresponding to a group of the first angle values selected based on the operation of the target user among the M head-related transfer functions.

7. The method according to claim 6, wherein The number of the first angle values is M groups. The M groups of first angle values are used to indicate M positions in the simulated three-dimensional space. M is a positive integer greater than or equal to 2. The obtaining of the target head-related transfer function according to the first size, the first head-related transfer function, and the first angle value includes: Obtaining a third head-related transfer function according to the first size, the first head-related transfer function, and a second angle value. The second angle value includes a group of the first angle values selected based on the operation of the target user among the M groups of first angle values; Determining the third head-related transfer function as the target head-related transfer function.

8. The method according to any one of claims 1 to 7, characterized in that, The obtaining of the target head-related transfer function according to the first size, the first head-related transfer function, and the first position includes: Inputting the first size, the first head-related transfer function, and the first position into a generative diffusion network to obtain the target head-related transfer function. The generative diffusion network is used to generate a head-related transfer function adapted to the target user based on the average data corresponding to the head-related transfer functions of multiple users, in combination with the first size and the first position of the target user.

9. The method according to claim 8, characterized in that The generation process of the generative diffusion network includes: Obtain N sample user images, where each sample user image among the N sample user images includes the target region, and N is a positive integer greater than or equal to 2; Determine N sets of sample sizes corresponding to the N target regions; Input the N sets of sample sizes into the original generative diffusion network to obtain N sample head-related transfer functions; Determine the average data corresponding to the N head-related transfer functions as the first head-related transfer function; Train the original generative diffusion network according to the relationships between the N sample head-related transfer functions and the first head-related transfer function respectively to obtain the generative diffusion network.

10. The method according to any one of claims 1 to 9, characterized in that The target region includes the head region, the ear region, and / or the image region corresponding to the torso region, and the head region includes other head regions except the ear region; The first size corresponding to the head region includes at least one of: head circumference, head width, head depth, head height, neck width, neck height, neck depth, anterior cranial offset, shoulder width, and shoulder circumference; The first size corresponding to the ear region includes at least one of: auricle height, auricle width, ear cavity height, ear cavity width, ear cavity depth, cymba conchae height, fossa of helix height, intertragal width, auricle rotation angle, auricle flare angle, auricle inferior offset, and auricle right offset; The first size corresponding to the torso region includes at least one of: torso width, torso height, torso depth, height, and sitting height; 11. The method according to any one of claims 1 to 10, characterized in that, The electronic device is applied to an audio system, the audio system includes the electronic device and a wearable device, after generating a second audio by performing audio rendering on a first audio in the electronic device according to the target head-related transfer function, the method further includes: Send the second audio to the wearable device so that the wearable device plays the second audio for the target user.

12. An electronic device, characterized in that, The electronic device includes: one or more processors, and a memory; The memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 11.

13. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions, when the instructions run on an electronic device, causing the electronic device to execute the method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Sound field expansion method, audio equipment and computer readable storage medium

    CN120769218A

  • Sound field expansion methods, audio devices, and computer-readable storage media

    CN120769218B