Audio adjustment method and device, equipment, storage medium and program product

By generating a user's hearing feature model and head-related transfer function, and generating an audio adjustment strategy based on user feedback, the problem that existing audio processing technologies cannot meet personalized needs is solved, and personalized adjustment of audio quality is achieved.

CN121940686APending Publication Date: 2026-04-28BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2025-12-15
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing audio processing technologies struggle to meet the personalized audio quality needs of different users, especially due to differences in users' hearing abilities leading to inconsistent audio experiences.

Method used

By generating a user's hearing feature model and head-related transfer function, an audio adjustment strategy is generated based on user feedback information and the head-related transfer function to adjust the audio to meet user needs, including adjusting volume, frequency, and sound source location.

Benefits of technology

It enables personalized audio adjustments based on users' hearing characteristics, improving audio quality satisfaction and meeting the individual needs of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940686A_ABST
    Figure CN121940686A_ABST
Patent Text Reader

Abstract

The invention relates to an audio adjustment method and device, equipment, a storage medium and a program product, and relates to the technical field of audio adjustment, and the method comprises the steps: receiving feedback information of a user for a test audio, generating a user hearing feature model corresponding to the user based on the feedback information, and responding to a received head-related transfer function corresponding to the user, the audio adjustment strategy is generated based on the user hearing feature model and the head-related transfer function, the audio is adjusted based on the audio adjustment strategy to obtain the target audio, and the audio is adjusted based on the related information corresponding to the user, so that the obtained audio can meet the requirement of the user for audio quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of audio adjustment technology, and in particular to an audio adjustment method, apparatus, device, storage medium, and program product. Background Technology

[0002] As technology advances, users have increasingly higher demands for audio quality. To improve audio quality, audio processing techniques are used. However, due to differences in hearing abilities among users, current audio processing technologies struggle to meet the diverse audio quality requirements of different users.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this disclosure provides an audio adjustment method, apparatus, device, storage medium, and program product.

[0005] According to a first aspect of the present disclosure, an audio adjustment method is provided, comprising: In response to receiving feedback from the user on the test audio, a user hearing feature model corresponding to the user is generated based on the feedback information; In response to receiving the head-related transfer function corresponding to the user, an audio adjustment strategy is generated based on the user's hearing feature model and the head-related transfer function; The audio is adjusted based on the audio adjustment strategy to obtain the target audio.

[0006] In one embodiment of this disclosure, the feedback information includes feedback information for the test audio volume and feedback information for the test audio frequency; Based on feedback information, a user-specific hearing feature model is generated, including: Based on feedback information regarding the volume and frequency of the test audio, a user's binaural hearing threshold curve is generated. The binaural hearing threshold curve is used to represent the hearing threshold of each ear at different frequencies.

[0007] In one embodiment of this disclosure, the method further includes: Retrieve the header-related transfer function for the user.

[0008] In one embodiment of this disclosure, obtaining the header-related transfer function corresponding to the user includes: The head-related transfer function is determined based on the user's human body geometry features.

[0009] In one embodiment of this disclosure, determining a head-related transfer function based on the user's human body geometric features includes: The user's human geometric features are input into a pre-trained intelligent model to obtain the head-related transfer function.

[0010] In one embodiment of this disclosure, determining a head-related transfer function based on the user's human body geometric features includes: The user's human geometric features are acquired using depth cameras and / or 3D scanning; The head-related transfer function is determined based on the user's human body geometry features.

[0011] In one embodiment of this disclosure, in response to receiving a head-related transfer function corresponding to a user, an audio adjustment strategy is generated based on the user's hearing feature model and the head-related transfer function, including: Audio adjustment strategies are generated based on signal processing algorithms, user hearing feature models, and head-related transfer functions.

[0012] In one embodiment of this disclosure, adjusting audio based on an audio adjustment strategy to obtain target audio includes: Based on the audio adjustment strategy, at least one of the following is adjusted: the position of the audio in the virtual space, the time of transmission to different ears of the user, the frequency of the audio transmitted to different ears of the user, and the level difference of the audio transmitted to different ears of the user, to obtain the target audio.

[0013] According to a second aspect of the present disclosure, an audio adjustment device is provided, comprising: The first generation module is used to respond to the user's feedback information on the test audio and generate a user hearing feature model corresponding to the user based on the feedback information. The second generation module is used to generate an audio adjustment strategy based on the user's hearing feature model and the head-related transfer function in response to receiving the head-related transfer function corresponding to the user. The adjustment module is used to adjust the audio based on the audio adjustment strategy to obtain the target audio.

[0014] In one embodiment of this disclosure, the first generation module includes: The first generation unit is used to generate a user's binaural hearing threshold curve based on feedback information regarding the volume of the test audio and feedback information regarding the frequency of the test audio. The binaural hearing threshold curve is used to represent the hearing threshold of each ear of the user at different frequencies.

[0015] In one embodiment of this disclosure, the apparatus further includes: The acquisition module is used to obtain the header-related transmission functions corresponding to the user.

[0016] In one embodiment of this disclosure, the acquisition module includes: The determination unit is used to determine the head-related transfer function based on the user's human body geometric features.

[0017] In one embodiment of this disclosure, the determining unit includes: The first determining subunit is used to input the user's human body geometric features into a pre-trained intelligent model to obtain the head-related transfer function.

[0018] In one embodiment of this disclosure, the determining unit includes: Acquisition subunit, used to acquire the user's human geometric features based on depth camera and / or 3D scan; The second determining subunit is used to determine the head-related transfer function based on the user's human body geometric features.

[0019] In one embodiment of this disclosure, the second generation module includes: The second generation unit is used to generate audio adjustment strategies based on signal processing algorithms, user hearing feature models, and head-related transfer functions.

[0020] In one embodiment of this disclosure, the adjustment module includes: The adjustment unit is used to adjust at least one of the following based on the audio adjustment strategy: the position of the audio in the virtual space, the time of transmission to different ears of the user, the frequency of the audio transmitted to different ears of the user, and the level difference of the audio transmitted to different ears of the user, so as to obtain the target audio.

[0021] According to a third aspect of the present disclosure, an electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to implement any of the audio adjustment methods described in the first aspect above.

[0022] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, which, when the instructions in the storage medium are executed by a processor of a terminal, enables the terminal to perform any of the audio adjustment methods of the first aspect described above.

[0023] According to a fifth aspect of the present disclosure, a computer program product is provided, wherein a computer program or computer instructions are loaded and executed by a processor to enable a computer to perform any of the audio adjustment methods of the first aspect described above.

[0024] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: This disclosure responds to receiving user feedback on test audio, generates a user hearing feature model corresponding to the user based on the feedback, responds to receiving a head-related transfer function corresponding to the user, generates an audio adjustment strategy based on the user hearing feature model and the head-related transfer function, adjusts the audio based on the audio adjustment strategy, and obtains the target audio. By adjusting the audio based on the user's relevant information, the obtained audio can meet the user's audio quality requirements.

[0025] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0027] Figure 1 This is a flowchart of an audio adjustment method according to an exemplary embodiment of the present disclosure. Figure 1 .

[0028] Figure 2 This is a flowchart of an audio adjustment method according to an exemplary embodiment of the present disclosure. Figure 2 .

[0029] Figure 3 This is a flowchart of an audio adjustment method according to an exemplary embodiment of the present disclosure. Figure 3 .

[0030] Figure 4 This is a flowchart of an audio adjustment method according to an exemplary embodiment of the present disclosure. Figure 4 .

[0031] Figure 5 This is a schematic diagram of the structure of an audio adjustment model according to an exemplary embodiment of the present disclosure.

[0032] Figure 6 This is a flowchart of an audio adjustment method according to an exemplary embodiment of the present disclosure. Figure 5 .

[0033] Figure 7 This is a flowchart of an audio adjustment method according to an exemplary embodiment of the present disclosure. Figure 6 .

[0034] Figure 8 This is a block diagram illustrating an audio adjustment device according to an exemplary embodiment of the present disclosure.

[0035] Figure 9This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0036] Exemplary embodiments of this disclosure will be described in detail herein, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.

[0037] The embodiments described below, which are examples of some of the embodiments of this disclosure, do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0038] The specific implementation methods of the embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0039] Figure 1 This is a flowchart of an audio adjustment method according to an exemplary embodiment of the present disclosure. Figure 1 The audio adjustment method can be applied to terminals, including but not limited to mobile phones, computers, smart wearable devices, smart home devices, smart cars, VR devices, and AR devices. It can also be applied to server-side applications such as local servers and cloud servers, which can be deployed on a single computer or a cluster of multiple computers.

[0040] like Figure 1 As shown, it includes the following steps.

[0041] S110, in response to receiving feedback information from the user on the test audio, generates a user hearing feature model corresponding to the user based on the feedback information.

[0042] In some embodiments, the test audio may include audio used to test a user's hearing. It should be noted that the test audio may include a combination of audio frequencies and volumes. For example, the volume difference between multiple audio frequencies in the test audio is 10 dB, and the frequency difference between multiple audio frequencies in the test audio is 100 Hz.

[0043] In some embodiments, the feedback information may include information indicating a user's response to the test audio. It should be noted that the feedback information may include feedback sent by the user through language and gestures. For example, when a user hears feedback information at a specific volume and frequency, they can send feedback information about the test audio by tapping the terminal or sending voice messages to the terminal.

[0044] In some embodiments, the method of this disclosure may further include: sending adjusted test audio to the user in response to not receiving feedback from the user on the test audio. It should be noted that the adjusted test audio may include test audio with adjusted volume and adjusted frequency. For example, In some embodiments, the adjusted test audio may include test audio with volume adjusted based on a preset step size and test audio with frequency adjusted based on a preset step size.

[0045] In some embodiments, a user hearing feature model may include a dataset reflecting the user's hearing. It should be noted that the hearing feature model may include models identified by mathematical functions, graphs, etc. For example, the hearing feature model in this disclosure embodiment may be a hearing curve.

[0046] In some embodiments, hearing curves can be used to represent the hearing threshold of each frequency in each ear of a user, the type and degree of hearing loss, and the difference or asymmetry between the left and right ears.

[0047] In some embodiments, this disclosure tests a user's hearing based on test audio and generates a user hearing feature model based on the user's feedback information, which can more accurately reflect the user's true hearing level and facilitate personalized adjustments to the audio received by the user.

[0048] S120, in response to receiving the head-related transfer function corresponding to the user, generates an audio adjustment strategy based on the user's hearing feature model and the head-related transfer function.

[0049] In some embodiments, the head-related transfer function may include the frequency domain transmission characteristics of the user's corresponding sound waves from the midpoint of a free field to the user's periosteum. Because different users have different physiological characteristics, the head-related transfer functions for different users will also be different.

[0050] In some embodiments, generating an audio adjustment strategy based on a user's hearing feature model and a head-related transfer function may include adjusting the user's hearing feature model using the head-related transfer function.

[0051] In some embodiments, the user's spatial audio parameters may first be adjusted based on a head-related transfer function. For example, adjusting the user's spatial audio parameters may include generating adjusted spatial audio parameters based on signal processing algorithms and equalizer, gain control, and spatial audio rendering techniques.

[0052] In some embodiments, generating an audio adjustment strategy based on a user's hearing feature model and a head-related transfer function may include adjusting the hearing curve based on the head-related transfer function.

[0053] For example, the hearing curve may include the curve corresponding to the transfer function, and the transfer function may include:

[0054] Where B is the bandwidth. , Q is the center frequency of the hearing test, and Q is the gain value. For digital center frequency, , This is the gain value.

[0055] S130, adjust the audio based on the audio adjustment strategy to obtain the target audio.

[0056] In some embodiments, adjusting audio based on an audio adjustment strategy may include adjusting the audio volume, frequency, and sound source location. The sound source location may include the location of the sound source in real space and virtual three-dimensional space. It should be noted that this disclosure may pre-acquire the audio and then adjust it based on the audio adjustment strategy to obtain the target audio.

[0057] This disclosure responds to receiving user feedback on test audio, generates a user hearing feature model corresponding to the user based on the feedback, responds to receiving a head-related transfer function corresponding to the user, generates an audio adjustment strategy based on the user hearing feature model and the head-related transfer function, adjusts the audio based on the audio adjustment strategy, and obtains the target audio. By adjusting the audio based on the user's relevant information, the obtained audio can meet the user's audio quality requirements.

[0058] Figure 2 This is a flowchart of a data processing method according to an exemplary embodiment of the present disclosure. Figure 2 . Figure 2 In steps S220 and S230 and Figure 1 Steps S120 and S130 correspond to each other and will not be repeated here, as follows: Figure 2 As shown, in Figure 1 In addition to the implementation process shown, the following steps are also included: S210 generates a user's binaural hearing threshold curve based on feedback information regarding the volume of the test audio and feedback information regarding the frequency of the test audio. The binaural hearing threshold curve is used to represent the hearing threshold of each ear of the user at different frequencies.

[0059] In some embodiments, the hearing threshold may include the lowest audio volume that a user can access.

[0060] In some embodiments, as described above, testing a user's hearing based on test audio can include testing with test audio at different volumes and test audio at different frequencies. Therefore, this disclosure can obtain feedback information regarding the volume of the test audio and feedback information regarding the frequency of the test audio.

[0061] In some embodiments, the binaural hearing threshold curve has been described in detail in the above embodiments and will not be repeated here.

[0062] In this embodiment of the disclosure, a user's binaural hearing threshold curve is generated based on the feedback information of the test audio volume and the feedback information of the test audio frequency. This fully obtains the binaural hearing threshold curves corresponding to different ears of the user, enabling separate consideration of the hearing of different ears of the user and improving the flexibility of audio adjustment.

[0063] Figure 3 This is a flowchart of a data processing method according to an exemplary embodiment of the present disclosure. Figure 3 . Figure 3 In steps S310, S330 and S340 and Figure 1 Steps S110 to S130 correspond to each other and will not be repeated here, as follows: Figure 3 As shown, in Figure 1 In addition to the implementation process shown, the following steps are also included: S320, obtain the header-related transmission function corresponding to the user.

[0064] In some embodiments, obtaining the head-related transfer function corresponding to a user may include determining the head-related transfer function based on the user's human geometric features.

[0065] In some embodiments, the terminal can directly obtain the header-related transfer function corresponding to the user. Alternatively, the terminal can obtain the header-related transfer function sent by other devices.

[0066] In some embodiments, the user's corresponding head-related transfer function can be obtained at a preset period. Since the user's physical characteristics may change, the user's corresponding head-related transfer function can be updated based on the preset period.

[0067] In this embodiment, the head-related transfer function (HRF) corresponding to the user is obtained, and then an audio adjustment strategy is generated based on the HRF and the user's hearing feature model. The audio is then adjusted based on the audio adjustment strategy to obtain the target audio. Because the obtained HRF is user-specific, the audio can be adjusted based on user characteristics, making the adjusted audio more in line with the user's needs.

[0068] Figure 4 This is a flowchart of a data processing method according to an exemplary embodiment of the present disclosure. Figure 4 . Figure 4 In the middle steps S410, S430 and S440 and Figure 1 Steps S310, S330, and S340 correspond to each other and will not be repeated here, as follows: Figure 4 As shown, in Figure 3 In addition to the implementation process shown, the following steps are also included: S420 inputs the user's human geometric features into a pre-trained intelligent model to obtain the head-related transfer function.

[0069] In some embodiments, the user's human body geometry may include shape data of the user's head, ears, and / or torso. It should be noted that the shape data of the user's head, ears, and / or torso may include size data and relative position data. The size data may include three-dimensional size data, and the relative position data may include the relative position data between the head and ears, and the relative position data between the head and torso.

[0070] In some embodiments, inputting the user's human geometric features into a pre-trained intelligent model to obtain a head-related transfer function may include inputting human geometric feature data into a pre-trained intelligent model and having the intelligent model output a head-related transfer function.

[0071] In some embodiments, the pre-trained intelligent model in this disclosure may include a neural network intelligent model, an adversarial network, and a variational autoencoder.

[0072] In some embodiments, the method of this disclosure can further match the user's human geometric features with human geometric features in a preset human geometric feature library, and use the head-related transfer function associated with the matched human geometric features as the user's corresponding human geometric features. It should be noted that the method of this disclosure can also include establishing an association relationship based on historical human geometric features and historical head-related transfer functions. A human geometric feature library is then established based on the established association relationship.

[0073] In some embodiments, the intelligent model can also be trained based on a human geometric feature library to obtain a pre-trained intelligent model.

[0074] In this embodiment of the disclosure, by inputting the user's human geometric features into a pre-trained intelligent model to obtain a head-related transfer function, the obtained head-related transfer function can be more closely matched with the user, thereby making the obtained target audio more able to meet the user's needs.

[0075] Figure 5 This is a flowchart of a data processing method according to an exemplary embodiment of the present disclosure. Figure 5 . Figure 5 In the middle steps S510, S540 and S550 and Figure 1 Steps S310, S330, and S340 correspond to each other and will not be repeated here, as follows: Figure 5 As shown, in Figure 3 In addition to the implementation process shown, the following steps are also included: The S520 acquires the user's human geometric features based on a depth camera and / or 3D scanning.

[0076] In some embodiments, a depth camera may include a 3D camera and an RGB-D camera. The depth camera can output the distance to each pixel in the scene, i.e., the pixel depth. Compared with a regular 2D camera, it can simultaneously obtain X / Y plane coordinates + depth Z, and then calculate three-dimensional coordinates for tasks such as 3D measurement, environmental perception, and 3D reconstruction.

[0077] It should be noted that a depth camera can include one that projects a structured pattern using an infrared projector, then captures the distorted pattern using an infrared camera, and calculates depth using triangulation. A depth camera can also include a binocular vision camera, which images the same scene from both sides, calculating depth based on parallax. Alternatively, a depth camera can emit modulated infrared light, directly measuring distance based on the time-of-flight of light.

[0078] In some embodiments, by acquiring the user's human geometric features through a depth camera, it is possible to obtain the depth and relative positional relationship of multiple organs of the user, thereby making the determined human geometric features more accurate.

[0079] In some embodiments, the user's human geometric features can also be acquired through 3D scanning. It should be noted that 3D scanning uses sensors such as optical, laser, or acoustic sensors to acquire the three-dimensional coordinates (X / Y / Z) of an object's surface, and can simultaneously acquire color / texture, outputting point clouds or triangular meshes for reconstruction, which can then be used for measurement, reverse engineering, inspection, and visualization. Its operation is analogous to a camera, but it records "distance" rather than color.

[0080] In some embodiments, 3D scanning methods may include: laser triangulation: projecting laser lines / points, with a camera observing deformation; high accuracy over short distances. Structured light: projecting coded gratings to solve fringe deformation; fast speed, suitable for static or semi-static scenes. Time-of-flight (ToF): measuring the round-trip time of light pulses to obtain distance; suitable for medium to large scenes and mobile platforms. Passive stereo vision: binocular / multi-viewing using parallax to recover depth; highly dependent on lighting and texture. Depth cameras: such as RGB-D cameras, conveniently acquiring depth maps; widely used for short-range applications. CT scans: achieving 3D reconstruction of internal structures through X-ray tomography; used for non-destructive testing / medical applications.

[0081] S530 determines the head-related transfer function based on the user's human body geometric features.

[0082] In some embodiments, the above-described embodiments have already determined the header-related transfer function based on user geometric features, which will not be repeated here.

[0083] In some embodiments, the user's human geometric features can be acquired based on a depth camera and a 3D scan, respectively, and then the final head-related transfer function can be determined based on the scanning results of the depth camera and the scanning results of the 3D scan.

[0084] In this embodiment of the disclosure, the user's human body geometric features are acquired through a depth camera and / or 3D scanning, and a head-related transfer function is obtained through the human body geometric features. This makes the obtained head-related transfer function more compatible with the user, thereby making the obtained target audio more able to meet the user's needs.

[0085] Figure 6 This is a flowchart of a data processing method according to an exemplary embodiment of the present disclosure. Figure 6 . Figure 6 In the middle steps S610 and S630 and Figure 1 Steps S110 and S130 correspond to each other and will not be repeated here, as follows: Figure 6 As shown, in Figure 1 In addition to the implementation process shown, the following steps are also included: S620, in response to receiving the head-related transfer function corresponding to the user, generates an audio adjustment strategy based on the signal processing algorithm, the user's hearing feature model, and the head-related transfer function.

[0086] In some embodiments, signal processing methods may include equalizers, gain control methods, and spatial audio rendering techniques. It should be noted that an equalizer is an algorithm used to adjust the relative intensity of different frequency components in an audio signal. Exemplary equalizers may include parametric equalizers, shelf equalizers, notch filters, and dynamic equalizers. Parametric equalizers allow users to adjust the center frequency, bandwidth (Q value), and gain, offering high flexibility and suitability for precise tuning. Shelf equalizers affect all frequencies above or below a specific frequency band, such as low-frequency shelves and high-frequency shelves. Notch filters are used to eliminate noise or interference at specific frequencies. Dynamic equalizers combine dynamic processing and equalization functions, automatically adjusting the frequency response based on the signal level. Gain control methods may include the process of adjusting the amplitude of the audio signal. This ensures the audio signal is within a suitable level range, avoiding distortion, clipping, or an excessively low signal-to-noise ratio. Common gain control methods include static gain adjustment, automatic gain control, compressors, amplitude transformers, and expanders. Spatial audio rendering technology is a technique that simulates the propagation and localization of sound in three-dimensional space, capable of perceiving the specific method, distance, and environment from which sound originates. Spatial audio rendering technology can include ambient acoustics, binaural rendering, object-based audio, channel-based audio, and wavefield synthesis. Ambient acoustics is an omnidirectional audio coding format that can represent sound fields in any direction in three-dimensional space, suitable for VR / AR and immersive audio. Binaural rendering can generate audio signals with directional information for each ear, typically combined with HRTF, to achieve 3D hearing through headphones. Object-based audio treats each sound as an independent object, containing its audio signal and spatial location information, dynamically adjusted by the renderer according to the playback environment.

[0087] In some embodiments, an audio adjustment strategy is generated based on a signal processing algorithm, a user hearing feature model, and a head-related transfer function. This may include adjusting the user hearing feature model based on the head-related transfer function using a signal processing algorithm to obtain an adjusted user hearing feature model. The adjusted user hearing feature model is then used as the audio adjustment strategy. For example, the binaural hearing threshold curve may be adjusted based on the head-related transfer function.

[0088] In this embodiment of the disclosure, after obtaining the user's hearing feature model and head-related transfer function, an audio adjustment strategy can be generated based on the user's hearing feature model and head-related transfer function. Adjusting the audio based on the audio feature strategy can make the adjusted audio better meet the user's needs.

[0089] Figure 7 This is a flowchart of a data processing method according to an exemplary embodiment of the present disclosure. Figure 7 . Figure 7 In the middle steps S710 and S720 and Figure 1 Steps S110 and S120 correspond to each other and will not be repeated here, as follows: Figure 7 As shown, in Figure 1 In addition to the implementation process shown, the following steps are also included: S730 adjusts at least one of the following based on the audio adjustment strategy: the position of the audio in the virtual space, the time of transmission to different ears of the user, the frequency of the audio transmitted to different ears of the user, and the level difference of the audio transmitted to different ears of the user, to obtain the target audio.

[0090] In some embodiments, the location of audio in virtual space can include its position within the virtual space of spatial audio, games, and videos. It should be noted that spatial audio refers to audio content or technology that accurately presents the direction, distance, motion trajectory, and environmental interaction effects of sound sources in three-dimensional space. This type of audio is not merely left-right channel stereo, but extends to richer dimensions such as horizontal plane, vertical direction, depth, and even environmental reflections and reverberation.

[0091] In some embodiments, the audio can be real-world audio, which is sent to the user by the terminal device after being received. During the transmission to the user, the location of the audio in the real world can be simulated in virtual space.

[0092] In some embodiments, since users have different hearing abilities, the timing, frequency, and audio level of the audio sent to different users' ears can be adjusted when adjusting the audio.

[0093] In some embodiments, to enable users to accurately determine the location of audio in virtual space, the location of audio in virtual space can be adjusted based on an audio adjustment strategy. For example, if a user's hearing in their left ear is weaker than in their right ear, the audio in the virtual space can be shifted towards the left ear.

[0094] In some embodiments, the level difference of audio transmitted to different ears of a user can include inter-ear sound level difference and inter-ear intensity difference. Inter-ear sound level difference (ILD) or inter-ear intensity difference (IID) refers to the difference in sound pressure level (i.e., volume intensity) when the same sound signal reaches a person's left and right ears. Inter-ear sound level difference is an indispensable element in constructing a realistic binaural spatial audio experience, especially in applications such as stereo, virtual reality (VR), 3D audio, and binaural audio.

[0095] In this embodiment of the disclosure, an audio adjustment strategy is used to adjust at least one of the following: the location of the audio in the virtual space, the time it takes to be transmitted to different ears of the user, the frequency of the audio transmitted to different ears of the user, and the level difference of the audio transmitted to different ears of the user, so that the user can receive audio that is more in line with the user's hearing characteristics.

[0096] Based on the same inventive concept, this disclosure also provides an audio adjustment device, as shown in the following embodiment. Since the principle by which this device embodiment solves the problem is similar to that of the above-described method embodiment, the implementation of this device embodiment can refer to the implementation of the above-described method embodiment, and repeated details will not be described again.

[0097] Figure 8 This is a block diagram illustrating an audio adjustment device according to an exemplary embodiment of the present disclosure. (Refer to...) Figure 8 The device 800 includes: The first generation module 810 is used to respond to the user's feedback information on the test audio and generate a user hearing feature model corresponding to the user based on the feedback information. The second generation module 820 is used to generate an audio adjustment strategy based on the user's hearing feature model and the head-related transfer function in response to receiving the head-related transfer function corresponding to the user. The adjustment module 830 is used to adjust the audio based on the audio adjustment strategy to obtain the target audio.

[0098] This disclosure responds to receiving user feedback on test audio, generates a user hearing feature model corresponding to the user based on the feedback, responds to receiving a head-related transfer function corresponding to the user, generates an audio adjustment strategy based on the user hearing feature model and the head-related transfer function, adjusts the audio based on the audio adjustment strategy, and obtains the target audio. By adjusting the audio based on the user's relevant information, the obtained audio can meet the user's audio quality requirements.

[0099] In one embodiment of this disclosure, the first generation module 810 includes: The first generation unit is used to generate a user's binaural hearing threshold curve based on feedback information regarding the volume of the test audio and feedback information regarding the frequency of the test audio. The binaural hearing threshold curve is used to represent the hearing threshold of each ear of the user at different frequencies.

[0100] In one embodiment of this disclosure, the apparatus further includes: The acquisition module is used to obtain the header-related transmission functions corresponding to the user.

[0101] In one embodiment of this disclosure, the acquisition module includes: The determination unit is used to determine the head-related transfer function based on the user's human body geometric features.

[0102] In one embodiment of this disclosure, the determining unit includes: The first determining subunit is used to input the user's human body geometric features into a pre-trained intelligent model to obtain the head-related transfer function.

[0103] In one embodiment of this disclosure, the determining unit includes: Acquisition subunit, used to acquire the user's human geometric features based on depth camera and / or 3D scan; The second determining subunit is used to determine the head-related transfer function based on the user's human body geometric features.

[0104] In one embodiment of this disclosure, the second generation module 820 includes: The second generation unit is used to generate audio adjustment strategies based on signal processing algorithms, user hearing feature models, and head-related transfer functions.

[0105] In one embodiment of this disclosure, the adjustment module 830 includes: The adjustment unit is used to adjust at least one of the following based on the audio adjustment strategy: the position of the audio in the virtual space, the time of transmission to different ears of the user, the frequency of the audio transmitted to different ears of the user, and the level difference of the audio transmitted to different ears of the user, so as to obtain the target audio.

[0106] This disclosure responds to receiving user feedback on test audio, generates a user hearing feature model corresponding to the user based on the feedback, responds to receiving a head-related transfer function corresponding to the user, generates an audio adjustment strategy based on the user hearing feature model and the head-related transfer function, adjusts the audio based on the audio adjustment strategy, and obtains the target audio. By adjusting the audio based on the user's relevant information, the obtained audio can meet the user's audio quality requirements.

[0107] Figure 9 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. For example, device 900 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.

[0108] Reference Figure 9 The device 900 may include one or more of the following components: a processing component 902, a memory 904, a power supply component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, and a communication component 916.

[0109] Processing component 902 typically controls the overall operation of device 900, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 902 may include one or more processors 920 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 902 may include one or more modules to facilitate interaction between processing component 902 and other components. For example, processing component 902 may include a multimedia module to facilitate interaction between multimedia component 908 and processing component 902.

[0110] Memory 904 is configured to store various types of data to support the operation of device 900. Examples of this data include instructions for any application or method operating on device 900, contact data, phonebook data, messages, pictures, videos, etc. Memory 904 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0111] Power supply component 906 provides power to various components of device 900. Power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 900.

[0112] Multimedia component 908 includes a screen that provides an output interface between device 900 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 908 includes a front-facing camera and / or a rear-facing camera. When device 900 is in an operating mode, such as shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0113] Audio component 910 is configured to output and / or input audio signals. For example, audio component 910 includes a microphone (MIC) configured to receive external audio signals when device 900 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 904 or transmitted via communication component 916. In some embodiments, audio component 910 also includes a speaker for outputting audio signals.

[0114] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0115] Sensor assembly 914 includes one or more sensors for providing state assessments of various aspects of device 900. For example, sensor assembly 914 may detect the on / off state of device 900, the relative positioning of components such as the display and keypad of device 900, changes in position of device 900 or a component of device 900, the presence or absence of user contact with device 900, orientation or acceleration / deceleration of device 900, and temperature changes of device 900. Sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 914 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 914 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0116] Communication component 916 is configured to facilitate wired or wireless communication between device 900 and other devices. Device 900 can access wireless networks based on communication standards such as Wi-Fi, 3G, 4G, 5G, other communication standards, or combinations thereof. In some embodiments of this disclosure, communication component 916 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of this disclosure, communication component 916 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0117] In some embodiments of this disclosure, the apparatus 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0118] In some embodiments of this disclosure, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions that can be executed by a processor 920 of device 900 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0119] In some embodiments of this disclosure, a non-transitory computer-readable storage medium enables a terminal to perform an audio adjustment method when instructions in the storage medium are executed by a terminal's processor.

[0120] In some embodiments of this disclosure, a computer program product is also provided, including a computer program / instructions that, when executed by a processor, implement an audio adjustment method.

[0121] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.

[0122] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.

[0123] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0124] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An audio adjustment method, characterized in that, include: In response to receiving feedback from the user on the test audio, a user hearing feature model corresponding to the user is generated based on the feedback information; In response to receiving the head-related transfer function corresponding to the user, an audio adjustment strategy is generated based on the user's hearing feature model and the head-related transfer function; The audio is adjusted based on the aforementioned audio adjustment strategy to obtain the target audio.

2. The method according to claim 1, characterized in that, The feedback information includes feedback information regarding the volume of the test audio and feedback information regarding the frequency of the test audio. The step of generating a user hearing feature model corresponding to the user based on the feedback information includes: Based on feedback information regarding the volume of the test audio and feedback information regarding the frequency of the test audio, a binaural hearing threshold curve is generated for the user. The binaural hearing threshold curve is used to represent the hearing threshold of each ear of the user at different frequencies.

3. The method according to claim 1, characterized in that, The method further includes: Obtain the header-related transfer function corresponding to the user.

4. The method according to claim 3, characterized in that, The step of obtaining the header-related transmission function corresponding to the user includes: The head-related transfer function is determined based on the user's human geometric features.

5. The method according to claim 4, characterized in that, Determining the head-related transfer function based on the user's human geometric features includes: The user's human geometric features are input into a pre-trained intelligent model to obtain the head-related transfer function.

6. The method according to claim 4, characterized in that, Determining the head-related transfer function based on the user's human geometric features includes: The user's human geometric features are obtained using a depth camera and / or 3D scanning. The head-related transfer function is determined based on the user's human geometric features.

7. The method according to claim 1, characterized in that, The step of generating an audio adjustment strategy based on the user's hearing feature model and the head-related transfer function in response to receiving the user's head-related transfer function includes: An audio adjustment strategy is generated based on signal processing algorithms, user hearing feature models, and the head-related transfer function.

8. The method according to claim 1, characterized in that, The process of adjusting the audio based on the audio adjustment strategy to obtain the target audio includes: Based on the audio adjustment strategy, at least one of the following is adjusted: the position of the audio in the virtual space, the time of transmission to different ears of the user, the frequency of the audio transmitted to different ears of the user, and the level difference of the audio transmitted to different ears of the user, to obtain the target audio.

9. An audio adjustment device, characterized in that, include: The first generation module is used to generate a user hearing feature model corresponding to the user based on the feedback information received from the user on the test audio. The second generation module is used to generate an audio adjustment strategy based on the user's hearing feature model and the head-related transfer function in response to receiving the head-related transfer function corresponding to the user. The adjustment module is used to adjust the audio based on the audio adjustment strategy to obtain the target audio.

10. The apparatus according to claim 9, characterized in that, The first generation module includes: The first generation unit is used to generate the user's binaural hearing threshold curve based on feedback information regarding the volume of the test audio and feedback information regarding the frequency of the test audio. The binaural hearing threshold curve is used to represent the hearing threshold of each ear of the user at different frequencies.

11. The apparatus according to claim 9, characterized in that, The device further includes: The acquisition module is used to acquire the header-related transmission function corresponding to the user.

12. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the audio adjustment method according to any one of claims 1 to 8.

13. A non-transitory computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of a terminal, enable the terminal to perform the steps of an audio adjustment method according to any one of claims 1 to 8.

14. A computer program product, said computer program product comprising a computer program or computer instructions, characterized in that, The computer program or the computer instructions are loaded and executed by the processor to enable the computer to implement the steps of the audio adjustment method as described in any one of claims 1-8.