Audio processing method, head-mounted display device, and computer-readable storage medium

By collecting and recognizing external audio information in a head-mounted display device, and performing acoustic parameter compensation and visualization processing, the problem of users being unable to obtain key external information in a timely manner during immersive experiences is solved. This enables timely prompts and accurate positioning of key external information, thereby improving user experience and security.

CN116312620BActive Publication Date: 2026-05-15GOERTEK INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GOERTEK INC
Filing Date
2023-02-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing head-mounted display devices, while providing an immersive experience, fail to effectively alert users to key environmental information, causing users to be unable to promptly discern external environmental conditions, resulting in inconvenience and safety hazards.

Method used

By dynamically collecting ambient audio information, using a converged audio recognition neural network model to identify key audio information, and compensating and adjusting acoustic parameters, key audio enhancement information is output, including beam phase shift and volume enhancement. Combined with a visualized display of the spatial location of the sound source, this ensures that users can obtain key external information in a timely manner during an immersive experience.

Benefits of technology

It effectively enhances users' perception and location of key audio information from the outside world, improves user experience, reduces security risks, and ensures that users do not miss important sound prompts during immersive experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312620B_ABST
    Figure CN116312620B_ABST
Patent Text Reader

Abstract

The application discloses an audio processing method, a head-mounted display device and a computer readable storage medium. The audio processing method comprises the following steps: dynamically collecting environmental audio information of an external environment, and identifying whether preset key audio information exists in the environmental audio information through a convergent audio recognition neural network model; if the key audio information exists, performing compensation adjustment on acoustic parameters of the key audio information to obtain key audio enhancement information; and outputting the key audio enhancement information. In the process of using the head-mounted display device, the application can effectively prompt the user with key environmental information of the external environment, so that the user can timely distinguish the condition of the external environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of wearable device technology, and more particularly to an audio processing method, a head-mounted display device, and a computer-readable storage medium. Background Technology

[0002] VR (Virtual Reality) devices or AR (Augmented Reality) devices are head-mounted display devices that are currently developing and becoming increasingly widespread. As the immersive experience of VR / AR devices continues to improve, users often become less aware of their surroundings while wearing these devices for an immersive experience, such as being less sensitive to external sounds. However, in some application scenarios, users still want to clearly hear key external sounds while using head-mounted displays. Examples include alarms (such as fire alarms), knocking on the door, and calls from others in a home setting; announcements at bus stops; and car horns while walking. In other words, while enjoying an immersive experience, users still need to pay attention to their environment. In many applications, these environmental elements contain crucial information, even warnings of danger. However, current head-mounted displays that provide an immersive experience almost completely isolate the user's hearing, preventing them from effectively and promptly obtaining crucial environmental information, undoubtedly causing significant inconvenience.

[0003] Therefore, how to effectively prompt users with key environmental information when they are immersed in using head-mounted display devices, so as to avoid users being unable to distinguish the external environment in time, has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] The main objective of this application is to provide an audio processing method, a head-mounted display device, and a computer-readable storage medium, aiming to solve the technical problem that, during the use of a head-mounted display device, it is impossible to effectively prompt users with key environmental information, resulting in users being unable to promptly distinguish the external environmental conditions.

[0005] To achieve the above objectives, this application provides an audio processing method applied to a head-mounted display device, the method comprising:

[0006] The system dynamically collects ambient audio information and uses a converged audio recognition neural network model to identify whether preset key audio information exists in the ambient audio information.

[0007] If the key audio information exists, then the acoustic parameters of the key audio information are compensated and adjusted to obtain key audio enhancement information;

[0008] Output the key audio enhancement information.

[0009] Optionally, the step of compensating and adjusting the acoustic parameters of the key audio information includes:

[0010] Obtain the user's current focus level coefficient on the extended reality environment presented in the head-mounted display device;

[0011] Determine the current ambient audio loss level associated with the current focus coefficient, wherein the higher the current focus coefficient, the higher the associated current ambient audio loss level;

[0012] From the preset loss level mapping relationship, the source spatial position deviation value and / or audio intensity loss value of the current environment audio loss level mapping are obtained by querying.

[0013] Based on the mapped spatial position deviation value of the sound source and / or the audio intensity loss value, the acoustic parameters of the key audio information are compensated and adjusted.

[0014] Optionally, the step of compensating and adjusting the acoustic parameters of the key audio information based on the mapped spatial location deviation value of the sound source and / or the audio intensity loss value includes:

[0015] The beam phase shift of the key audio information is determined based on the mapped spatial position deviation value of the sound source, and / or the beam amplitude loss of the key audio information is determined based on the mapped audio intensity loss value.

[0016] Based on the beam phase shift and / or the beam amplitude loss, determine the audio parameter compensation information for the key audio information;

[0017] Based on the audio parameter compensation information, acoustic parameter compensation adjustments are made to the key audio information to compensate for beam phase shift and / or beam amplitude loss of the key audio information.

[0018] Optionally, before the step of determining the current environmental audio loss level associated with the current focus coefficient, the method further includes:

[0019] Play a preset virtual test audio, wherein the spatial location of the sound source of the virtual test audio is a preset spatial orientation;

[0020] Output a preset guidance interface to guide the user to determine the spatial location of the sound source of the virtual test audio;

[0021] Obtain the directional information input by the user in response to the preset guidance interface, compare the directional information with the preset spatial orientation, and determine the user's key audio resolution based on the comparison result;

[0022] Based on the key audio resolution, the user's perception sensitivity to key audio information is determined. Based on the magnitude of the perception sensitivity, a focus mapping gradient matching the perception sensitivity is selected from a preset mapping gradient database. The focus mapping gradient includes multiple focus coefficients and the environmental audio loss level associated with each focus coefficient.

[0023] The step of determining the current environmental audio loss level associated with the current focus coefficient includes:

[0024] Based on the matched attention mapping gradient, the current environmental audio loss level associated with the current attention coefficient is determined.

[0025] Optionally, the step of obtaining the user's current focus coefficient on the extended reality environment presented in the head-mounted display device includes:

[0026] The system detects current user physiological characteristics and device usage status information. The user physiological characteristics include at least one of pupil size, blink frequency, heart rate, respiratory rate, and body temperature. The device usage status information includes at least one of the following: duration of use of the head-mounted display device, movement status, power consumption rate, and currently running applications.

[0027] Based on the user's physiological characteristics and the device's usage status information, the user's current focus coefficient on the extended reality environment presented in the head-mounted display device is determined.

[0028] Optionally, the step of determining the user's current focus coefficient on the extended reality environment presented in the head-mounted display device based on the user's physiological characteristics information and the device usage status information includes:

[0029] Obtain a preset focus recognition neural network model;

[0030] The user's physiological characteristics and the device's usage status information are input into the attention recognition neural network model to predict the user's current attention coefficient to the extended reality environment presented in the head-mounted display device.

[0031] Optionally, the method further includes:

[0032] Obtain audio information for requirement recognition corresponding to at least one application scenario;

[0033] Multiple demand identification audio information items are associated with key audio tags to obtain a key audio sample set, and multiple environmental noise information items are associated with interference audio tags to obtain an interference audio sample set, wherein the environmental noise information does not include the demand identification audio information;

[0034] The preset neural network model is trained using the key audio sample set and the interference audio sample set to obtain a converged audio recognition neural network model.

[0035] Optionally, after the step of outputting the key audio enhancement information, the method further includes:

[0036] The spatial location of the sound source corresponding to the key audio information is displayed on the display interface of the head-mounted display device by marking it on a radar map or azimuth scale.

[0037] This application also provides a head-mounted display device, which is a physical device. The head-mounted display device includes: a memory, a processor, and a program of the audio processing method stored in the memory and executable on the processor. When the program of the audio processing method is executed by the processor, it can implement the steps of the audio processing method as described above.

[0038] This application also provides a computer-readable storage medium storing a program implementing an audio processing method, the program implementing the audio processing method being executed by a processor to implement the steps of the audio processing method as described above.

[0039] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the audio processing method described above.

[0040] The technical solution of this application involves dynamically acquiring ambient audio information and using a convergent audio recognition neural network model to identify whether preset key audio information exists within this ambient audio information. This allows for the identification of audio information that is important to the user in the current scenario, thus capturing this key audio information. If the key audio information is present, its acoustic parameters are compensated and adjusted to obtain enhanced key audio information, which is then output. This ensures that when key audio information is identified in the ambient audio information, acoustic compensation is applied to enhance its volume, making it easier to shift the user's attention to the highlighted key audio information. Alternatively, the spatial location of the sound source corresponding to the key audio information can be calibrated, allowing the user to more clearly and accurately perceive the key audio information and identify its corresponding sound source. This effectively alerts the user to key environmental information during immersive use of the head-mounted display device, preventing situations where the user cannot promptly distinguish external objects, thus improving the user experience while reducing safety hazards. Attached Figure Description

[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0042] To more clearly illustrate the technical solutions in this embodiment or the prior art, the accompanying drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a flowchart illustrating the first embodiment of the audio processing method of this application;

[0044] Figure 2 This is a flowchart illustrating the second embodiment of the audio processing method of this application;

[0045] Figure 3 This is a flowchart illustrating the third embodiment of the audio processing method of this application;

[0046] Figure 4 A schematic diagram illustrating a scenario where a user locates the source of key audio information while wearing a head-mounted display device;

[0047] Figure 5 This is a schematic diagram illustrating a scenario in which the spatial location of key audio information is visualized in an embodiment of this application.

[0048] Figure 6This is a preset guide interface in one embodiment of this application;

[0049] Figure 7 This is a schematic diagram of the hardware operating environment involved in the head-mounted display device in this embodiment.

[0050] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0051] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] In this embodiment, the head-mounted display device includes, but is not limited to, Mixed Reality (MR) devices (such as MR glasses or MR helmets), Augmented Reality (AR) devices (such as AR glasses or AR helmets), Virtual Reality (VR) devices (such as VR glasses or VR helmets), Extended Reality (XR) devices, or some combination thereof, etc.

[0053] In some application scenarios, users wearing head-mounted displays need to clearly hear key external sounds, such as alarms (e.g., fire alarms), knocking on doors, and calls from others in a home environment; announcements at bus stops; and car horns while walking. In other words, while enjoying an immersive experience, users also need to pay attention to their environment. This is because, in many applications, environmental elements contain crucial information, even warnings of danger. However, current head-mounted displays that provide an immersive experience almost completely block out external hearing, preventing users from effectively and promptly obtaining crucial environmental information, undoubtedly causing significant inconvenience.

[0054] Example 1

[0055] Based on this, please refer to Figure 1 This embodiment provides an audio processing method, the method comprising:

[0056] Step S10: Dynamically collect ambient audio information from the outside world, and identify whether there is preset key audio information in the ambient audio information through a converged audio recognition neural network model;

[0057] In this embodiment, ambient audio information refers to the sound information generated by the external environment when the user is wearing the head-mounted display device. That is, ambient audio information is distinct from the sound generated by the head-mounted display device itself.

[0058] In one embodiment, the ambient audio information can be acquired via a microphone on the head-mounted display device. In another embodiment, it can be acquired by receiving ambient audio information sent from other terminal devices (such as smartwatches, mobile phones, or smart speakers) that are communicatively connected to the head-mounted display device.

[0059] It should be noted that multiple operating modes can be set in the head-mounted display device system. These operating modes can include immersive mode and smart mode. Users can trigger the function button corresponding to the immersive mode to activate the head-mounted display device. After the head-mounted display device enters immersive mode, it will completely isolate external ambient audio information and will not execute step S10, as well as subsequent steps S20 and S30 of this embodiment, thereby improving the immersive experience of the user when using the head-mounted display device. Users can also trigger the function button corresponding to the smart mode to activate the head-mounted display device. After the head-mounted display device enters smart mode, during the user's use of the head-mounted display device, step S10 will be executed: dynamically collecting external ambient audio information and identifying whether preset key audio information exists in the ambient audio information through a converged audio recognition neural network model, and subsequent steps S20 and S30 will also be executed. This allows users to enter immersion mode when in scenarios with high safety requirements or low need for awareness of external environmental elements (e.g., in a relatively quiet VR experience venue or under supervision), by touching a specific button. This maximizes the user's immersion experience. Conversely, in scenarios with low safety requirements or high need for awareness of external environmental elements (e.g., fire alarms, knocking, and calls from others at home; announcements at bus stops; car horns while walking), users can enter smart mode by touching a specific button. This mode effectively alerts the user to key environmental information, allowing them to quickly assess their surroundings and further enhancing immersion. In other words, users can select the most suitable operating mode based on the specific application scenario, matching their actual needs and providing a better product experience.

[0060] In this embodiment, environmental audio information can be input into a converged audio recognition neural network model, thereby identifying whether preset key audio information exists in the environmental audio information. This key audio information refers to audio that is important to the user, such as alarm sounds (e.g., fire alarms), knocking sounds, and other people calling to the user in a home setting; station announcements in a bus setting; and car horns in a walking setting.

[0061] Step S20: If the key audio information exists, then the acoustic parameters of the key audio information are compensated and adjusted to obtain key audio enhancement information;

[0062] In this embodiment, because users' attention is focused on the augmented reality content of the head-mounted display device, even if key audio information in the ambient audio is prompted, users may not perceive it due to their immersion in the augmented reality content, thus ignoring the key audio information from the outside world (this is because the sensitivity of the human sensory system is affected by subjective focus). Therefore, this embodiment requires compensation and adjustment of the acoustic parameters of the key audio information to obtain enhanced key audio information. For example, the beam amplitude of the key audio information can be compensated to increase its volume, thereby making it easier to shift the user's attention to the prompted key audio information. Alternatively, the beam phase shift of the key audio information can be compensated to more accurately present the spatial location of the sound source corresponding to the key audio information, making it easier for the user to clearly and accurately distinguish the sound source location and avoid risks in time. Those skilled in the art will recognize that existing technologies such as beamforming, adaptive filtering, and adaptive volume can be used to compensate and adjust the acoustic parameters of the key audio information to ensure that users can receive prompts for the key audio information while focusing on the augmented reality content.

[0063] Step S30: Output the key audio enhancement information.

[0064] In this embodiment, the key audio enhancement information can be played through the built-in headphones corresponding to the left and right ears in the head-mounted display device to achieve the purpose of outputting the key audio enhancement information.

[0065] The technical solution of this embodiment is to dynamically collect ambient audio information and use a convergent audio recognition neural network model to identify whether there is preset key audio information in the ambient audio information. This identifies whether there is audio information that is important to the user in the current scene, thus capturing the key audio information. If the key audio information is present, the acoustic parameters of the key audio information are compensated and adjusted to obtain key audio enhancement information, which is then output. This allows for acoustic compensation of the key audio information when it is determined that it exists in the ambient audio information, thereby enhancing the volume of the key audio information and making it easier to shift the user's attention to the key audio information. Alternatively, the spatial position of the sound source corresponding to the key audio information can be calibrated, enabling the user to more clearly and accurately obtain the key audio information and distinguish the corresponding sound source. This allows the user to be effectively prompted with key environmental information while immersing themselves in using the head-mounted display device, preventing situations where the user cannot distinguish external objects in time, improving the user experience, and reducing safety hazards.

[0066] For example, after the step of outputting the key audio enhancement information, the method further includes:

[0067] Step A10: The spatial location of the sound source corresponding to the key audio information is displayed on the display interface of the head-mounted display device by marking it on a radar map or azimuth scale.

[0068] It's important to note that binaural localization ability, or the ability to distinguish sound sources, can be quantified as the precision of the azimuth angle of the sound source relative to the user's own position—the smallest azimuth angle that can be resolved. For example, if the 360 ​​degrees surrounding the user are divided into four equal parts, the user can only distinguish which of the four parts the azimuth angle of the sound source relative to their own position belongs to; they cannot subdivide it further. In this case, the smallest azimuth angle the user can resolve is 90 degrees. As another example, if the 360 ​​degrees surrounding the user are divided into eight equal parts, the user can only distinguish which of the eight parts the azimuth angle of the sound source relative to their own position belongs to. In this case, the smallest azimuth angle the user can resolve is 40 degrees (e.g., ...). Figure 6 (As shown).

[0069] However, due to the immersion and excessive focus users experience when using head-mounted displays, their auditory system may ignore external stimuli excessively. This affects sound source localization by increasing the minimum azimuth angle and reducing positioning accuracy. Consequently, users may misjudge the spatial location of sound sources corresponding to key audio information, causing unnecessary trouble and potential dangers. For example, ... Figure 4As shown, if the minimum azimuth angle is α1 under normal conditions (i.e., not focused on the extended reality content presented by the head-mounted display), and α2 under focused conditions (i.e., focused on the extended reality content presented by the head-mounted display), and α1 < α2, then if B is a hazard (e.g., the sound of a heavy object falling at a construction site, or the sound of a car horn in a walking scene), then it may be misidentified as B' under focused conditions.

[0070] Based on this, this embodiment displays the spatial location of the sound source corresponding to the key audio information on the display interface of the head-mounted display device by marking it on a radar chart or azimuth scale. This makes the sound source location visible, meaning that this embodiment enhances the user's auditory experience and visually prompts the user with the spatial location information of the sound source, making it easier for the user to receive. Figure 5 , Figure 5 This is a schematic diagram illustrating a scenario in which the spatial location of key audio information sources is visualized in an embodiment of this application.

[0071] Among them, the visual display interface is as follows: Figure 5 As shown in the right figure, the display interface can display graphical information such as radar charts or azimuth scales (which may include head movement angle scales and sound source location scales). These radar charts or azimuth scales are marked with the spatial locations of the sound sources corresponding to key audio information, providing users with intuitive visual cues. This allows users to more clearly and accurately identify the sound source location corresponding to key audio information, effectively alerting them to crucial environmental information during immersive use of the head-mounted display device and reducing unnecessary trouble and potential dangers.

[0072] In one possible implementation, please refer to Figure 2 The step of compensating and adjusting the acoustic parameters of the key audio information includes:

[0073] Step S21: Obtain the user's current focus coefficient on the extended reality environment presented in the head-mounted display device;

[0074] In this embodiment, the current focus coefficient is used to characterize the user's current level of focus on the extended reality environment presented in the head-mounted display device. The higher the current focus coefficient, the higher the level of focus.

[0075] In one embodiment, the step of obtaining the user's current focus coefficient on the extended reality environment presented in the head-mounted display device may specifically be: obtaining the focus coefficient input by the user, and using the focus coefficient as the user's current focus coefficient on the extended reality environment presented in the head-mounted display device.

[0076] In another embodiment, the step of obtaining the user's current focus coefficient on the extended reality environment presented in the head-mounted display device may further include: obtaining the target application currently running on the head-mounted display device, querying the focus coefficient mapped to the target application from a preset application mapping coefficient table, and using the focus coefficient mapped to the target application as the current focus coefficient. The application mapping coefficient table stores a one-to-one mapping relationship between each application and the focus coefficient. These applications include, but are not limited to, VR games, VR movies, VR shopping, photography, music playback, settings (setting functions for parameters such as sound or images), information notifications, weather status, and voice calls. Those skilled in the art will understand that different applications often map to different focus coefficients; for example, the focus coefficient mapped to VR games or VR movies is often higher than that mapped to music playback, settings (setting functions for parameters such as sound or images), or information notifications.

[0077] Step S22: Determine the current ambient audio loss level associated with the current focus coefficient, wherein the higher the current focus coefficient, the higher the associated current ambient audio loss level;

[0078] In this embodiment, the current ambient audio loss level is used to characterize the degree of loss of ambient audio information caused by the user's reduced perception of ambient audio due to focusing on the extended reality environment presented in the head-mounted display device. The higher the current ambient audio loss level, the greater the degree of loss of ambient audio information.

[0079] Those skilled in the art will understand that the more focused a user is on the extended reality content of a head-mounted display, the lower their sensitivity to ambient audio information, and the easier it is for them to ignore key audio information. Therefore, the higher the current focus coefficient, the higher the associated ambient audio loss level. As an example, the associated ambient audio loss level can be obtained by querying a preset focus coefficient mapping table. This focus coefficient mapping table includes multiple different focus coefficients and the ambient audio loss levels mapped to each focus coefficient.

[0080] Step S23: From the preset loss level mapping relationship, query the source spatial position deviation value and / or audio intensity loss value of the current environment audio loss level mapping.

[0081] In this embodiment, the information in the loss level mapping relationship includes multiple different environmental audio loss levels, as well as the source spatial location deviation value and / or audio intensity loss value mapped to each environmental audio loss level. It is readily understood that the higher the current environmental audio loss level, the higher the mapped source spatial location deviation value and / or audio intensity loss value.

[0082] Step S24: Based on the mapped spatial position deviation value of the sound source and / or the audio intensity loss value, the acoustic parameters of the key audio information are compensated and adjusted.

[0083] For example, the step of compensating and adjusting the acoustic parameters of the key audio information based on the mapped spatial location deviation value of the sound source and / or the audio intensity loss value includes:

[0084] Step B10: Determine the beam phase shift of the key audio information based on the mapped spatial position deviation value of the sound source, and / or determine the beam amplitude loss of the key audio information based on the mapped audio intensity loss value.

[0085] Step B20: Determine the audio parameter compensation information for the key audio information based on the beam phase shift and / or the beam amplitude loss;

[0086] Step B30: Based on the audio parameter compensation information, adjust the acoustic parameters of the key audio information to compensate for the beam phase shift and / or beam amplitude loss of the key audio information.

[0087] In this embodiment, a larger spatial position deviation value of the sound source indicates a larger beam phase shift of the key audio information. Correspondingly, a larger audio intensity loss value indicates a larger beam amplitude loss of the key audio information. It is easy to understand that this audio parameter compensation information includes phase shift compensation information to compensate for or correct the spatial position deviation of the sound source, and / or beam amplitude compensation information to compensate for or correct the audio intensity loss. This facilitates the adjustment of acoustic parameters for the key audio information based on the audio parameter compensation information, achieving the goal of accurately compensating for the beam phase shift and / or beam amplitude loss of the key audio information.

[0088] This embodiment obtains the user's current focus coefficient on the extended reality environment presented in the head-mounted display device, determines the current environmental audio loss level associated with the current focus coefficient, wherein the higher the current focus coefficient, the higher the associated current environmental audio loss level; from the preset loss level mapping relationship, it retrieves the sound source spatial position deviation value and / or audio intensity loss value mapped to the current environmental audio loss level; based on the mapped sound source spatial position deviation value and / or audio intensity loss value, it compensates and adjusts the acoustic parameters of key audio information, thereby increasing the volume of key audio information by compensating for the beam amplitude of key audio information, which is more conducive to shifting the user's attention to the key audio information, and / or, by compensating for the beam phase shift of key audio information, it more accurately presents the spatial position of the sound source corresponding to the key audio information, which is more conducive to the user clearly and accurately distinguishing the sound source position corresponding to the key audio information. Furthermore, it enables the user to be effectively prompted with key environmental information during the immersive use of the head-mounted display device, avoiding situations where the user cannot distinguish external environmental objects in time.

[0089] In one possible implementation, prior to the step of determining the current ambient audio loss level associated with the current focus coefficient, the method further includes:

[0090] Step C10: Play a preset virtual test audio, wherein the spatial location of the sound source of the virtual test audio is a preset spatial orientation;

[0091] In this embodiment, the virtual test audio has relative sound source location information (i.e., spatial location of the sound source) to the user. It should be noted that the virtual test audio is audio simulated by the head-mounted display device. That is, by controlling the built-in headphones corresponding to the left and right ears of the head-mounted display device to play audio with a frequency response time difference (using the different distances from the sound source to the left and right ears, so that the sound signal will also have a slight difference in the time it takes to travel to the two ears, such a time difference can help us understand the spatial location of the sound source. Based on the difference in time and loudness of the audio to a person's two ears, the spatial location of the sound source corresponding to the audio can be determined), so that the user feels that the audio is coming from a preset spatial location, thereby simulating stereo sound with sound source spatial location information.

[0092] Step C20: Output a preset guidance interface to guide the user to determine the spatial location of the sound source of the virtual test audio.

[0093] In one embodiment, the preset guidance interface can guide the user to directly input the spatial orientation range value of the virtual test audio corresponding to the sound source, for example, the spatial orientation range value is the orientation range value of 30 degrees to 60 degrees to the right front of the user.

[0094] In another embodiment, the preset guidance interface may include a directional selection ring, which is divided into multiple equal blocks, and the user's position is preset to be at the center of the directional selection ring. The user can select one of the equal blocks by touch to input the spatial orientation information of the corresponding sound location of the virtual test audio. Figure 6 As shown, Figure 6 As shown in one embodiment of this application, the preset guidance interface displays an azimuth selection ring, which is divided into eight equal parts, each part being 40 degrees (i.e., the smallest azimuth angle currently being distinguished is 40 degrees). By listening to the virtual test audio played by the head-mounted display device, the user identifies the direction of the sound source spatial location as the direction pointed to by the dark-colored part. Therefore, the user can input the spatial azimuth information corresponding to the sound source spatial location by touching the dark-colored part.

[0095] Step C30: Obtain the directional information input by the user in response to the preset guidance interface, compare the directional information with the preset spatial orientation, and determine the user's key audio resolution based on the comparison result;

[0096] It should be noted that this key audio resolution is used to characterize the ability to identify the spatial location of the sound source of key audio information.

[0097] In this embodiment, in the comparison results, the smaller the deviation between the location information and the preset spatial location, the higher the user's key audio resolution. Conversely, the larger the deviation between the location information and the preset spatial location, the lower the user's key audio resolution.

[0098] It should be noted that those skilled in the art can test the key audio resolution for different user vocal positions by sequentially playing multiple virtual test audios, each with a different spatial location of the sound source.

[0099] Step C40: Based on the key audio resolution, determine the user's perception sensitivity to key audio information; based on the magnitude of the perception sensitivity, select a focus mapping gradient matching the perception sensitivity from a preset mapping gradient database, wherein the focus mapping gradient includes multiple focus coefficients and the environmental audio loss level associated with each focus coefficient.

[0100] Specifically, the lower the perceptual sensitivity, the larger the focus mapping gradient. Conversely, the higher the perceptual sensitivity, the smaller the focus mapping gradient.

[0101] To aid in understanding the embodiments of this application, an example is provided. In this example, the attention mapping gradients stored in the mapping gradient database, from smallest to largest, are: the first mapping gradient and the second mapping gradient. In the first mapping gradient, when the attention coefficient ranges from [0.1, 0.35), the associated environmental audio loss level is low audio loss; when the attention coefficient ranges from [0.35, 0.65), the associated environmental audio loss level is mid audio loss; and when the attention coefficient ranges from [0.65, 0.9), the associated environmental audio loss level is mid-high audio loss. Similarly, in the second mapping gradient, when the attention coefficient ranges from [0.1, 0.35), the associated environmental audio loss level is mid audio loss; when the attention coefficient ranges from [0.35, 0.65), the associated environmental audio loss level is mid-high audio loss; and when the attention coefficient ranges from [0.65, 0.9), the associated environmental audio loss level is high audio loss. It should be noted that the above examples of mapping gradient databases and focus mapping gradients are only used to assist in understanding this application and are not intended to limit this application. Any simple modifications based on the technical concepts or principles of the embodiments of this application are all within the protection scope of this application.

[0102] The step of determining the current environmental audio loss level associated with the current focus coefficient includes:

[0103] Step C50: Determine the current environmental audio loss level associated with the current focus coefficient based on the matched focus mapping gradient.

[0104] In this embodiment, it is easy to understand that the larger the focus mapping gradient, the greater the level of current environmental audio loss associated with the current focus coefficient at the same time (because the level of current environmental audio loss associated with the same focus coefficient is greater).

[0105] This embodiment plays a preset virtual test audio, where the spatial location of the virtual test audio source is a preset spatial orientation. A preset guidance interface is output to guide the user in determining the spatial location of the virtual test audio source. The orientation information input by the user in response to the preset guidance interface is acquired, compared with the preset spatial orientation, and the user's key audio resolution is determined based on the comparison result. This accurately detects the user's ability to identify sound source localization. Based on the key audio resolution, the user's perceptual sensitivity to key audio information is determined. According to the magnitude of the perceptual sensitivity, a focus mapping gradient matching the perceptual sensitivity is selected from a preset mapping gradient database. Since the mapping gradient database has multiple preset focus mapping gradients that match one-to-one with perceptual sensitivity, different perceptual sensitivities are matched with different focus mapping gradients. Therefore, this embodiment can detect the user's perception of key audio information. The system accurately tests the sensitivity of the user's attention mapping gradient (which characterizes the correlation between each attention coefficient and the level of environmental audio loss) to better reflect the user's individual situation. Then, based on the matched attention mapping gradient, the system determines the current environmental audio loss level associated with the current attention coefficient. This allows the system to determine a more suitable attention mapping gradient based on the user's ability to locate sound sources in the environment, thus more accurately calibrating the current environmental audio loss level. This facilitates the subsequent personalized determination of audio parameter compensation information tailored to the user's individual needs, compensating for beam phase shift and / or beam amplitude loss in key audio information. Ultimately, this effectively alerts the user to key environmental information, enhancing the user's immersion in the head-mounted display device by enabling them to promptly identify environmental conditions.

[0106] In one implementable manner, the step of obtaining the user's current focus coefficient on the extended reality environment presented in the head-mounted display device includes:

[0107] Step D10: Detect current user physiological characteristics and device usage status information. The user physiological characteristics include at least one of pupil size, blink frequency, heart rate, respiratory rate, and body temperature. The device usage status information includes at least one of the following: duration of use of the head-mounted display device, movement status, power consumption rate, and currently running applications.

[0108] Step D20: Based on the user's physiological characteristics information and the device usage status information, determine the user's current focus coefficient on the extended reality environment presented in the head-mounted display device.

[0109] For example, further, in one possible implementation, the step of determining the user's current focus coefficient on the extended reality environment presented in the head-mounted display device based on the user's physiological characteristic information and the device usage status information includes:

[0110] Step E10: Obtain the preset attention recognition neural network model;

[0111] Step E20: Input the user's physiological characteristic information and the device usage status information into the attention recognition neural network model to predict the user's current attention coefficient to the extended reality environment presented in the head-mounted display device.

[0112] In this embodiment, those skilled in the art will understand that user physiological characteristics, such as pupil size and heart rate, can reflect a user's emotions. When a user's emotions fluctuate (e.g., experiencing excitement or tension while playing VR games or watching VR videos), it often reflects a relatively high level of focus (or immersion) on the extended reality environment presented in the head-mounted display device. Conversely, when a user's emotions are relatively calm, it often reflects a relatively low level of focus. Furthermore, user physiological characteristics, such as blink rate, heart rate, respiratory rate, and body temperature, can reflect the user's physical or psychological activity level. A higher level of activity often reflects a relatively high level of focus on the extended reality environment presented in the head-mounted display device, while a lower level of activity often reflects a relatively low level of focus.

[0113] Furthermore, in this embodiment, it is easy to understand that a longer duration of use of the head-mounted display device better reflects a higher level of user focus on the extended reality environment presented on the head-mounted display. Similarly, more active movement of the head-mounted display device indicates a higher level of user physical activity, which often reflects a relatively higher level of user focus on the extended reality environment presented on the head-mounted display. A higher power consumption rate of the head-mounted display device often indicates a more complex operating environment of the application running on the head-mounted display, which is more likely to attract the user's attention to the extended reality environment presented on the head-mounted display (for example, the power consumption rate of VR games or VR movie applications is higher than that of music playback, settings, or information notifications). Correspondingly, the level of user focus on the extended reality environment presented on the head-mounted display often varies depending on the currently running application; for example, the level of focus required for VR games or VR movies is often higher than that required for music playback, settings, or information notifications.

[0114] Therefore, this embodiment detects current user physiological characteristics and device usage status information. The user physiological characteristics include at least one of pupil size, blink frequency, heart rate, respiratory rate, and body temperature. The device usage status information includes at least one of the duration of use of the head-mounted display device, movement status, power consumption rate, and currently running applications. By comprehensively considering multiple factors of user physiological characteristics and device usage status information, the system evaluates the user's level of attention to the extended reality environment presented on the head-mounted display device, thereby accurately determining the user's current focus coefficient.

[0115] Example 2

[0116] Based on the first embodiment of this application, in another embodiment of this application, the same or similar content as in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 The method further includes:

[0117] Step S40: Obtain audio information for requirement recognition corresponding to at least one application scenario;

[0118] In this embodiment, the application scenario may include, but is not limited to, scenarios such as taking a bus, staying at home, and walking. Specifically, the audio information for demand recognition in a bus scenario may be a station announcement. The audio information for demand recognition in a home scenario may be a fire alarm, a knock on the door, or someone calling the user. The audio information for demand recognition in a walking scenario may be a car horn.

[0119] Step S50: Associate multiple demand identification audio information with key audio tags to obtain a key audio sample set, and associate multiple environmental noise information with interference audio tags to obtain an interference audio sample set, wherein the environmental noise information does not contain the demand identification audio information;

[0120] Step S60: Train the preset neural network model using the key audio sample set and the interference audio sample set to obtain a converged audio recognition neural network model.

[0121] This embodiment obtains demand identification audio information corresponding to at least one application scenario, associates multiple demand identification audio information with key audio tags to obtain a key audio sample set, and associates multiple environmental noise information with interference audio tags to obtain an interference audio sample set. Using the key audio sample set and the interference audio sample set, a preset neural network model is trained, thereby accurately and efficiently training a converged audio recognition neural network model. This facilitates the subsequent accurate identification of whether preset key audio information exists in environmental audio information using the converged audio recognition neural network model.

[0122] Example 3

[0123] This invention also provides an audio processing apparatus, the audio processing apparatus comprising:

[0124] The recognition module is used to dynamically collect ambient audio information and identify whether there is preset key audio information in the ambient audio information through a converged audio recognition neural network model.

[0125] The compensation module is used to compensate and adjust the acoustic parameters of the key audio information to obtain key audio enhancement information;

[0126] The output module is used to output the key audio enhancement information.

[0127] Optionally, the compensation module is further configured to:

[0128] Obtain the user's current focus level coefficient on the extended reality environment presented in the head-mounted display device;

[0129] Determine the current ambient audio loss level associated with the current focus coefficient, wherein the higher the current focus coefficient, the higher the associated current ambient audio loss level;

[0130] From the preset loss level mapping relationship, the source spatial position deviation value and / or audio intensity loss value of the current environment audio loss level mapping are obtained by querying.

[0131] Based on the mapped spatial position deviation value of the sound source and / or the audio intensity loss value, the acoustic parameters of the key audio information are compensated and adjusted.

[0132] Optionally, the compensation module is further configured to:

[0133] The beam phase shift of the key audio information is determined based on the mapped spatial position deviation value of the sound source, and / or the beam amplitude loss of the key audio information is determined based on the mapped audio intensity loss value.

[0134] Based on the beam phase shift and / or the beam amplitude loss, determine the audio parameter compensation information for the key audio information;

[0135] Based on the audio parameter compensation information, acoustic parameter compensation adjustments are made to the key audio information to compensate for beam phase shift and / or beam amplitude loss of the key audio information.

[0136] Optionally, the compensation module is further configured to:

[0137] Play a preset virtual test audio, wherein the spatial location of the sound source of the virtual test audio is a preset spatial orientation;

[0138] Output a preset guidance interface to guide the user to determine the spatial location of the sound source of the virtual test audio;

[0139] Obtain the directional information input by the user in response to the preset guidance interface, compare the directional information with the preset spatial orientation, and determine the user's key audio resolution based on the comparison result;

[0140] Based on the key audio resolution, the user's perception sensitivity to key audio information is determined. Based on the magnitude of the perception sensitivity, a focus mapping gradient matching the perception sensitivity is selected from a preset mapping gradient database. The focus mapping gradient includes multiple focus coefficients and the environmental audio loss level associated with each focus coefficient.

[0141] The step of determining the current environmental audio loss level associated with the current focus coefficient includes:

[0142] Based on the matched attention mapping gradient, the current environmental audio loss level associated with the current attention coefficient is determined.

[0143] Optionally, the compensation module is further configured to:

[0144] The system detects current user physiological characteristics and device usage status information. The user physiological characteristics include at least one of pupil size, blink frequency, heart rate, respiratory rate, and body temperature. The device usage status information includes at least one of the following: duration of use of the head-mounted display device, movement status, power consumption rate, and currently running applications.

[0145] Based on the user's physiological characteristics and the device's usage status information, the user's current focus coefficient on the extended reality environment presented in the head-mounted display device is determined.

[0146] Optionally, the compensation module is further configured to:

[0147] Obtain a preset focus recognition neural network model;

[0148] The user's physiological characteristics and the device's usage status information are input into the attention recognition neural network model to predict the user's current attention coefficient to the extended reality environment presented in the head-mounted display device.

[0149] Optionally, the audio processing device further includes a training module, the training module being used for:

[0150] Obtain audio information for requirement recognition corresponding to at least one application scenario;

[0151] Multiple demand identification audio information items are associated with key audio tags to obtain a key audio sample set, and multiple environmental noise information items are associated with interference audio tags to obtain an interference audio sample set, wherein the environmental noise information does not include the demand identification audio information;

[0152] The preset neural network model is trained using the key audio sample set and the interference audio sample set to obtain a converged audio recognition neural network model.

[0153] Optionally, the output module is further configured to:

[0154] The spatial location of the sound source corresponding to the key audio information is displayed on the display interface of the head-mounted display device by marking it on a radar map or azimuth scale.

[0155] The audio processing apparatus provided in this invention, employing the audio processing method described in Embodiment 1 or Embodiment 2, can solve the technical problem that, during the use of a head-mounted display device, it is impossible to effectively prompt the user with key environmental information, resulting in the user's inability to promptly distinguish the external environmental conditions. Compared with the prior art, the beneficial effects of the audio processing apparatus provided in this invention are the same as those of the audio processing method provided in the above embodiments, and other technical features in the audio processing apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0156] Example 4

[0157] This invention provides a head-mounted display device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the audio processing method described in Embodiment 1 above.

[0158] The following is for reference. Figure 7This document illustrates a structural schematic diagram suitable for implementing a head-mounted display device according to embodiments of the present disclosure. The head-mounted display device in these embodiments can be headphones or other head-mounted display devices. This head-mounted display device includes, but is not limited to, Mixed Reality (MR) devices (e.g., MR glasses or MR helmets), Augmented Reality (AR) devices (e.g., AR glasses or AR helmets), Virtual Reality (VR) devices (e.g., VR glasses or VR helmets), Extended Reality (XR) devices, or some combination thereof. Figure 7 The head-mounted display device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0159] like Figure 7 As shown, the head-mounted display device may include a processing unit 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM 1002) or a program loaded from a storage device into a random access memory (RAM 1004). The RAM 1004 also stores various programs and data required for the operation of the head-mounted display device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface is also connected to the bus 1005.

[0160] Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the head-mounted display to communicate wirelessly or wiredly with other devices to exchange data. Although head-mounted display devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.

[0161] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.

[0162] The head-mounted display device provided by this invention, employing the audio processing method in the above embodiments, can solve the technical problem that, during the use of the head-mounted display device, it is impossible to effectively prompt users with key environmental information, resulting in users being unable to promptly distinguish the external environmental conditions. Compared with the prior art, the beneficial effects of the head-mounted display device provided by this invention are the same as the beneficial effects of the audio processing method provided in the above embodiments, and other technical features in this head-mounted display device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0163] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0164] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0165] Example 5

[0166] This invention provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to execute the audio processing method described in the above embodiments.

[0167] The computer-readable storage medium provided in this embodiment of the invention may be, for example, a USB flash drive, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0168] The aforementioned computer-readable storage medium may be included in the head-mounted display device; or it may exist independently and not assembled into the head-mounted display device.

[0169] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by the head-mounted display device, the head-mounted display device: dynamically acquires ambient audio information from the outside world, and identifies whether preset key audio information exists in the ambient audio information through a converged audio recognition neural network model; if the key audio information exists, it compensates and adjusts the acoustic parameters of the key audio information to obtain key audio enhancement information; and outputs the key audio enhancement information.

[0170] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0171] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0172] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0173] The computer-readable storage medium provided by this invention stores computer-readable program instructions for executing the above-described audio processing method. This solves the technical problem that, during the use of a head-mounted display device, the user cannot effectively be prompted with key environmental information, leading to the user's inability to promptly discern the external environmental conditions. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this embodiment are the same as those of the audio processing method provided in Embodiment 1 or Embodiment 2, and will not be repeated here.

[0174] Example 6

[0175] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the audio processing method described above.

[0176] The computer program product provided in this application solves the technical problem that, during the use of head-mounted display devices, it is impossible to effectively prompt users with key environmental information, resulting in users being unable to promptly distinguish the external environmental conditions. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this invention are the same as the beneficial effects of the audio processing methods provided in Embodiment 1 or Embodiment 2 above, and will not be repeated here.

[0177] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

Claims

1. An audio processing method, characterized in that, The audio processing method is applied to a head-mounted display device, and the method includes: The system dynamically collects ambient audio information and uses a converged audio recognition neural network model to identify whether there is preset key audio information in the ambient audio information. If the key audio information exists, then the acoustic parameters of the key audio information are compensated and adjusted to obtain key audio enhancement information; Output the key audio enhancement information; The step of compensating and adjusting the acoustic parameters of the key audio information includes: Obtain the user's current focus level coefficient on the extended reality environment presented in the head-mounted display device; Determine the current ambient audio loss level associated with the current focus coefficient, wherein the higher the current focus coefficient, the higher the associated current ambient audio loss level; From the preset loss level mapping relationship, the source spatial position deviation value and / or audio intensity loss value of the current environment audio loss level mapping are obtained by querying. Based on the mapped spatial position deviation value of the sound source and / or the audio intensity loss value, the acoustic parameters of the key audio information are compensated and adjusted. Prior to the step of determining the current environmental audio loss level associated with the current focus coefficient, the method further includes: Play a preset virtual test audio, wherein the spatial location of the sound source of the virtual test audio is a preset spatial orientation; Output a preset guidance interface to guide the user to determine the spatial location of the sound source of the virtual test audio; Obtain the directional information input by the user in response to the preset guidance interface, compare the directional information with the preset spatial orientation, and determine the user's key audio resolution based on the comparison result; Based on the key audio resolution, the user's perception sensitivity to key audio information is determined. Based on the magnitude of the perception sensitivity, a focus mapping gradient matching the perception sensitivity is selected from a preset mapping gradient database. The focus mapping gradient includes multiple focus coefficients and the environmental audio loss level associated with each focus coefficient. The step of determining the current environmental audio loss level associated with the current focus coefficient includes: Based on the matched attention mapping gradient, the current environmental audio loss level associated with the current attention coefficient is determined.

2. The audio processing method as described in claim 1, characterized in that, The step of compensating and adjusting the acoustic parameters of the key audio information based on the mapped spatial position deviation value of the sound source and / or the audio intensity loss value includes: The beam phase shift of the key audio information is determined based on the mapped spatial position deviation value of the sound source, and / or the beam amplitude loss of the key audio information is determined based on the mapped audio intensity loss value. Based on the beam phase shift and / or the beam amplitude loss, determine the audio parameter compensation information for the key audio information; Based on the audio parameter compensation information, acoustic parameter compensation adjustments are made to the key audio information to compensate for beam phase shift and / or beam amplitude loss of the key audio information.

3. The audio processing method as described in claim 1, characterized in that, The step of obtaining the user's current focus coefficient on the extended reality environment presented in the head-mounted display device includes: The system detects current user physiological characteristics and device usage status information. The user physiological characteristics include at least one of pupil size, blink frequency, heart rate, respiratory rate, and body temperature. The device usage status information includes at least one of the following: duration of use of the head-mounted display device, movement status, power consumption rate, and currently running applications. Based on the user's physiological characteristics and the device's usage status information, the user's current focus coefficient on the extended reality environment presented in the head-mounted display device is determined.

4. The audio processing method as described in claim 3, characterized in that, The step of determining the user's current focus coefficient on the extended reality environment presented in the head-mounted display device based on the user's physiological characteristics information and the device usage status information includes: Obtain a preset focus recognition neural network model; The user's physiological characteristics and the device's usage status information are input into the attention recognition neural network model to predict the user's current attention coefficient to the extended reality environment presented in the head-mounted display device.

5. The audio processing method as described in claim 1, characterized in that, The method further includes: Obtain audio information for requirement recognition corresponding to at least one application scenario; Multiple demand identification audio information items are associated with key audio tags to obtain a key audio sample set, and multiple environmental noise information items are associated with interference audio tags to obtain an interference audio sample set, wherein the environmental noise information does not include the demand identification audio information; The preset neural network model is trained using the key audio sample set and the interference audio sample set to obtain a converged audio recognition neural network model.

6. The audio processing method according to any one of claims 1 to 5, characterized in that, After the step of outputting the key audio enhancement information, the method further includes: The spatial location of the sound source corresponding to the key audio information is displayed on the display interface of the head-mounted display device by marking it on a radar map or azimuth scale.

7. A head-mounted display device, characterized in that, The head-mounted display device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the steps of the audio processing method as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program implementing an audio processing method, which is executed by a processor to implement the steps of the audio processing method as claimed in any one of claims 1 to 6.