Audio playing method and system, head-mounted device, audio device, medium and product

By integrating the posture detection device in the headset, detecting and sending posture information to the audio device, and adjusting the optimal listening area of ​​the sound field, the problem of limited audio playback effects of existing headsets is solved, and a better audio experience and better sound immersion is achieved.

CN120166348APending Publication Date: 2025-06-17GOLDANA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510459280.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing headsets have certain restrictions on audio playback effects, which cannot bring users a further sound effect experience.

Method used

The position information of the device is detected by the position detection device in the headset, and the information is sent to the audio device with the audio device. The audio device adjusts the optimal listening area of ​​the sound field according to the position information, so that it overlaps with the position of the headset, performs near-ear sound field rendering processing and plays the audio.

Benefits of technology

It achieves a better audio experience for users, provides near-ear audio effects, enhances sound immersion, and ensures the consistency of sound effects when users move on their headsets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166348A_ABST
    Figure CN120166348A_ABST
Patent Text Reader

Abstract

The invention discloses an audio playing method and system, head-mounted equipment, audio equipment, a medium and a product, and relates to the technical field of head-mounted equipment. The audio playing method is applied to the head-mounted device, the head-mounted device establishes communication connection with a split audio device, and the audio playing method comprises the following steps: detecting pose information of the head-mounted device through a pose detection device in the head-mounted device; and sending the to-be-played audio and the pose information to the audio equipment, so that the audio equipment adjusts the optimal listening area of the sound field according to the pose information, the optimal listening area is overlapped with the position of the head-mounted equipment, near-ear sound field rendering processing is performed on the to-be-played audio, and the processed audio is played. According to the invention, when the user wears the head-mounted device, a better audio effect is provided for the user, and meanwhile, a better sound immersion feeling is brought to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of head-mounted devices, and in particular, to an audio playback method, a head-mounted device, an audio device, an audio playback system, a storage medium, and a computer program product. Background Art

[0002] With the development and maturity of head-mounted device technology, head-mounted devices such as MR (Mixed Reality) devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, and smart glasses have been widely used and loved by the public. At present, most head-mounted devices such as MR devices, VR devices, AR devices, and smart glasses are equipped with a micro speaker system to reproduce sound. Due to limitations such as hardware size and model, there are certain limitations in aspects such as timbre balance, linearity, and loudness, and it is impossible to bring a better sound effect experience to users.

[0003] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of the present application is to provide an audio playback method, a head-mounted device, an audio device, an audio playback system, a storage medium, and a computer program product, aiming to solve the problem that the audio playback effect of current head-mounted devices has certain limitations.

[0005] To achieve the above object, the present application proposes an audio playback method. The audio playback method is applied to a head-mounted device, and the head-mounted device establishes a communication connection with a separately arranged audio device. The audio playback method includes:

[0006] Detecting pose information of the head-mounted device through a pose detection device in the head-mounted device;

[0007] Sending the audio to be played and the pose information to the audio device, so that the audio device adjusts the optimal listening area of the sound field according to the pose information, overlaps the optimal listening area with the position of the head-mounted device, and performs near-ear sound field rendering processing on the audio to be played and plays the processed audio.

[0008] Optionally, after the step of detecting the pose information of the head-mounted device through the pose detection device in the head-mounted device, the method further includes:

[0009] Determining whether the head-mounted device is in a large movement state, where the large movement state refers to a state in which the change amount of the pose information detected within a continuous first preset duration compared to the pose information sent to the audio device last time is greater than a preset change amount.

[0010] The step of sending the pose information to the audio device includes:

[0011] If it is determined that the head-mounted device is in the large movement state, send the latest detected pose information to the audio device.

[0012] Optionally, the audio playback method further includes:

[0013] Obtain target information, where the target information includes first information about the current running scenario of the head-mounted device or second information about the current running game;

[0014] Dynamically update the detection frequency of the pose information according to the target information;

[0015] The step of detecting the pose information of the head-mounted device by the pose detection device in the head-mounted device includes:

[0016] Detect the pose information of the head-mounted device by the pose detection device in the head-mounted device according to the latest determined detection frequency.

[0017] Optionally, the first information includes information for indicating whether the current running scenario needs to interact with the user. In the case where the target information includes the first information, the step of dynamically updating the detection frequency of the pose information according to the target information includes:

[0018] Dynamically update the detection frequency of the pose information according to the first information, where the updated detection frequency in the first case is greater than the updated detection frequency in the second case. The first case refers to the case where the first information indicates that the current running scenario needs to interact with the user, and the second case refers to the case where the first information indicates that the current running scenario does not need to interact with the user.

[0019] Optionally, the second information includes interaction information of the current running game within a second preset duration in the future. In the case where the target information includes the second information, the step of dynamically updating the detection frequency of the pose information according to the target information includes:

[0020] Dynamically update the detection frequency of the pose information according to the interaction information, where the updated detection frequency in the third case is greater than the updated detection frequency in the fourth case. The third case refers to the case where the interaction information indicates that the user needs to change the head pose within the second preset duration in the future, and the fourth case refers to the case where the interaction information indicates that the user does not need to change the head pose within the second preset duration in the future.

[0021] Optionally, the pose detection device includes a UWB module and an IMU module. The step of detecting the pose information of the head-mounted device through the pose detection device in the head-mounted device includes:

[0022] When detecting the pose information for the first time after establishing a connection with the audio device, detecting the initial relative pose between the head-mounted device and the audio device through the UWB module as the pose information;

[0023] When detecting the pose information non-first time after establishing a connection with the audio device, detecting the amount of pose change of the head-mounted device through the IMU module, and using the amount of pose change as the pose information.

[0024] Optionally, when the head-mounted device includes speakers and microphones arranged at the positions of both ears, the audio playback method further includes:

[0025] When sending the audio to be played to the audio device for playback, collecting a sound signal through the microphone;

[0026] Performing sound passthrough processing on the sound signal, and playing the sound signal after sound passthrough processing through the speaker.

[0027] Optionally, the audio playback method further includes:

[0028] In response to a playback channel selection instruction, selecting a playback channel from the speaker of the head-mounted device and the audio device;

[0029] When selecting the speaker of the head-mounted device as the playback channel, playing the audio to be played through the speaker of the head-mounted device;

[0030] When selecting the audio device as the playback channel, performing the step of sending the audio to be played and the pose information to the audio device.

[0031] In addition, to achieve the above object, the present application further provides an audio playback method. The audio playback method is applied to an audio device. The audio device is communicatively connected to a head-mounted device, and the audio device is separately arranged from the head-mounted device. The audio playback method includes:

[0032] Receiving the audio to be played and the pose information sent by the head-mounted device, where the pose information is detected by a pose detection device in the head-mounted device;

[0033] Adjusting the optimal listening area of the sound field according to the pose information, overlapping the optimal listening area with the position of the head-mounted device, and performing near-ear sound field rendering processing on the audio to be played and playing the processed audio.

[0034] Optionally, the audio playback method further includes:

[0035] Sending feedback information to the head-mounted device.

[0036] In addition, to achieve the above object, the present application further provides a head-mounted device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the audio playback method applied to the head-mounted device as described above.

[0037] In addition, to achieve the above object, the present application further provides an audio device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the audio playback method applied to the audio device as described above.

[0038] In addition, to achieve the above object, the present application further provides an audio playback system, which includes the head-mounted device as described above and the audio device as described above.

[0039] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the audio playback method as described above are implemented.

[0040] In addition, to achieve the above object, the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the audio playback method as described above are implemented.

[0041] One or more technical solutions proposed by the present application have at least the following technical effects:

[0042] In this embodiment, the pose detection device in the head-mounted device detects the pose information of the head-mounted device, and sends the pose information and the audio to be played in the head-mounted device to the audio device. The audio device adjusts the optimal listening area of the sound field to overlap with the position of the head-mounted device according to the pose information, performs near-ear sound field rendering processing on the audio to be played, and plays the processed audio, providing an audio playback solution that brings a better audio experience to users in the head-mounted device usage scenario. On the one hand, the solution in this embodiment plays the audio to be played in the head-mounted device through an external audio device, enabling users to experience both the functions and services provided by the head-mounted device and the better audio effects provided by the audio device compared to the micro speakers in the head-mounted device when using the head-mounted device. On the other hand, since the user wears the head-mounted device, the solution in this embodiment dynamically detects the pose information through the head-mounted device and sends it to the audio device, allowing the audio device to adjust the optimal listening area of the sound field to overlap with the position of the head-mounted device according to the pose information, perform near-ear sound field rendering processing on the audio to be played, provide a near-ear audio effect for users, bring a better sense of sound immersion to users, and ensure that the sound effect remains consistent during the process of the user wearing the head-mounted device and moving. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0044] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a schematic flowchart provided for the first embodiment of the audio playback method of the present application;

[0046] Figure 2 It is a schematic diagram of the usage scenario involved in an embodiment of the audio playback method of the present application;

[0047] Figure 3 It is a system flowchart involved in an embodiment of the audio playback method of the present application;

[0048] Figure 4 It is a schematic diagram of the multi-audio device usage scenario involved in an embodiment of the audio playback method of the present application;

[0049] Figure 5 It is a schematic diagram of the usage scenario involved in an embodiment of the audio playback method of the present application for dynamically updating the detection frequency according to the running scenario;

[0050] Figure 6 This is a schematic diagram of a usage scenario for dynamically updating the detection frequency according to interaction information, which relates to an embodiment of the audio playback method of the present application.

[0051] The realization of the purpose, functional features, and advantages of the present application will be further described in conjunction with embodiments with reference to the accompanying drawings. Specific embodiments

[0052] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0053] To better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific embodiments.

[0054] The main solution of the embodiments of the present application is:

[0055] With the development and maturity of head-mounted device technology, head-mounted devices such as MR devices, VR devices, AR devices, and smart glasses have been widely used and loved by the public. Currently, most head-mounted devices such as MR devices, VR devices, AR devices, and smart glasses are equipped with a micro speaker system to reproduce sound. Due to limitations such as hardware size and model, there are certain limitations in aspects such as tone balance, linearity, and loudness, and it is impossible to bring a better sound effect experience to users.

[0056] The present application provides a solution. The pose detection device in the head-mounted device is used to detect the pose information of the head-mounted device, and the pose information and the audio to be played in the head-mounted device are sent to the audio device. The audio device adjusts the position of the optimal listening area of the sound field to overlap with the head-mounted device according to the pose information, performs near-ear sound field rendering processing on the audio to be played, and plays the processed audio, providing an audio playback solution that brings a better audio experience to users in the usage scenario of the head-mounted device. On the one hand, the solution of this embodiment plays the audio to be played by the head-mounted device through an external audio device, so that when the user uses the head-mounted device, the user can not only experience the functions and services provided by the head-mounted device, but also experience the better audio effect provided by the audio device compared with the micro speaker in the head-mounted device. On the other hand, since the user wears the head-mounted device, the solution of this embodiment dynamically detects the pose information through the head-mounted device and sends it to the audio device, so that the audio device adjusts the position of the optimal listening area of the sound field to overlap with the head-mounted device according to the pose information, performs near-ear sound field rendering processing on the audio to be played, provides a near-ear audio effect for the user, brings a better sense of sound immersion to the user, and ensures that the sound effect remains consistent during the process of the user wearing the head-mounted device and moving.

[0057] The first embodiment of the audio playback method of the present application is proposed below. Refer to Figure 1, Figure 1 This is a schematic flowchart of the first embodiment of the audio playback method of the present application. In this embodiment, the audio playback method is applied to a head-mounted device. Specifically, its usage scenarios may include head-mounted devices that play audio for users. It can be understood that there are many types of such head-mounted devices. For example, MR devices, VR devices, AR devices, or smart glasses, etc. The types of head-mounted devices to which the audio playback method is applied are not limited in the embodiments of the present application. A pose detection device is provided in the head-mounted device to detect the relative pose between the head-mounted device and the audio device. For example, it can detect the position and orientation of the head-mounted device relative to the audio device. There are many ways to implement the pose detection device. For example, considering the protection of user personal privacy, non-camera methods such as UWB (Ultra-Wideband) modules and IMU (Inertial Measurement Unit) modules can be used to implement pose detection. In this embodiment, the implementation method of the pose detection device is not limited. The audio playback method includes steps S10 to S20:

[0058] Step S10: Detect the pose information of the head-mounted device through the pose detection device in the head-mounted device.

[0059] It should be noted that the audio device and the head-mounted device described in the embodiments of the present application are separate, referring to an external audio device, not the audio device built into the head-mounted device.

[0060] The head-mounted device can establish a communication connection with the audio device through wired or wireless means. The audio device can be an audio device configured with a near-ear sound field rendering algorithm, which can perform near-ear sound field rendering processing on the audio, so that the played audio can achieve a near-ear effect. That is, when the user wears the head-mounted device, the audio played from the audio device feels like it is playing in the ear, that is, it can simulate the direct sound effect of near-ear listening on the audio device (non-headphone device). The type of the audio device is not limited in this embodiment. For example, the audio device can be a soundbar, a speaker array, a stereo system, a surround sound system, a panoramic sound system, etc. While experiencing the visual effect, it can bring a better audio experience to the user.

[0061] The head-mounted device can be configured to automatically establish a communication connection with the audio device, or can be configured to respond to a control instruction and establish a communication connection with the audio device after detecting the control instruction. For example, the user can trigger a control instruction to establish a communication connection with the audio device through a physical button in the head-mounted device or a virtual button displayed on the display screen. When the head-mounted device detects the control instruction, it searches for connectable audio devices and, in the presence of connectable audio devices, establishes a communication connection with the audio device.

[0062] When the head-mounted device establishes a communication connection with the audio device, the pose information of the head-mounted device can be detected by the pose detection device. Pose refers to position and orientation. The pose information can be information representing the position and orientation of the head-mounted device relative to the audio device, or information that can be used to calculate the position and orientation of the head-mounted device relative to the audio device. In this embodiment, the specific data form of the pose information is not limited. It can be understood that when the user wears the head-mounted device, the pose of the head-mounted device can represent the pose of the user's head.

[0063] There are many specific implementation methods for detecting the pose information of the head-mounted device by the pose detection device, which are not limited in this embodiment. For example, the pose detection device can include a UWB module and an IMU module. The position of the head-mounted device relative to the audio device is detected by the UWB module, and the orientation of the head-mounted device relative to the audio device is detected by the IMU module. The detected position and orientation are used as the pose information of the head-mounted device.

[0064] It should be noted that the UWB module realizes the precise positioning function by emitting extremely short pulses (nanosecond level) and occupying an extremely wide spectrum bandwidth (usually exceeding 500 MHz). Therefore, the pose information of the head-mounted device relative to the audio device can be detected by the cooperation of the UWB module in the head-mounted device and the UWB module in the audio device. The IMU module measures the motion state and spatial orientation of an object through built-in sensors (accelerometer, gyroscope), and detects physical quantities such as acceleration and angular velocity to track the motion trajectory and direction change of the object in real time. Therefore, it can be used to detect the pose information of the head-mounted device.

[0065] In a feasible implementation manner, the pose detection device includes a UWB module and an IMU module. The step S10 includes: when detecting the pose information for the first time after establishing a connection with the audio device, detecting the initial relative pose between the head-mounted device and the audio device through the UWB module as the pose information; when detecting the pose information not for the first time after establishing a connection with the audio device, detecting the pose change amount of the head-mounted device through the IMU module, and using the pose change amount as the pose information. Among them, the UWB module may include at least two UWB units, and the attitude of the head-mounted device relative to the audio device can be calculated based on the distances between the head-mounted device and the audio device detected by the two UWB units. Considering that in the usage scenarios of head-mounted devices such as VR devices and AR devices, the IMU module is generally in an on state, the pose change amount detected by the IMU module can be reused and sent to the audio device; and considering that the IMU module can only detect the pose change amount of the head-mounted device itself and cannot directly detect the pose of the head-mounted device relative to the audio device, the UWB module is used to detect the pose of the head-mounted device relative to the audio device during the first detection. Through the pose detection method of this implementation manner, the original IMU module of the head-mounted device is reused, avoiding additional power consumption.

[0066] In another feasible implementation, to improve the accuracy of pose detection, the UWB module and the IMU module can be used in cooperation for pose detection. Specifically, the step S10 includes: when detecting the pose information for the first time after establishing a connection with the audio device, detecting the first relative pose (i.e., the initial relative pose) between the head-mounted device and the audio device through the UWB module as the pose information; when detecting the pose information non-first time after establishing a connection with the audio device, detecting the amount of pose change of the head-mounted device through the IMU module, detecting the second relative pose between the head-mounted device and the audio device through the UWB module (which may have changed relative to the initial relative pose), and calculating the third relative pose between the head-mounted device and the audio device according to the amount of pose change; if the difference between the second relative pose and the third relative pose with the same timestamp is less than a preset threshold, performing weighted averaging on the second relative pose and the third relative pose with the same timestamp to obtain a fourth relative pose, and using the fourth relative pose as the pose information; if the difference between the second relative pose and the third relative pose with the same timestamp is greater than or equal to the preset threshold, outputting a fault prompt. In a specific implementation, the preset threshold can be set in advance according to needs. When the difference between the second relative pose and the third relative pose is greater than or equal to the preset threshold, it indicates that the detection results of the two modules differ greatly, and the detection result of at least one module may be incorrect. Since it is impossible to determine which detection result of the two modules is incorrect, a fault prompt can be output to prompt the user to detect and repair the fault. The weights for weighted averaging can be set in advance according to needs. For example, they can be determined according to the pose detection accuracy of the UWB module and the IMU module. The higher the accuracy, the greater the corresponding weight.

[0067] The head-mounted device can detect the pose information at a certain frequency, which can be a preset fixed frequency or dynamically adjusted according to a preset dynamic adjustment strategy. In this embodiment, there is no limitation on the detection frequency.

[0068] Step S20: Sending the audio to be played and the pose information to the audio device, so that the audio device can adjust the optimal listening area of the sound field according to the pose information, overlap the optimal listening area with the position of the head-mounted device, and perform near-ear sound field rendering processing on the audio to be played and play the processed audio.

[0069] The audio to be played refers to the audio that needs to be played in the current running scenario. For example, in a music playing scenario, the audio to be played is the music to be played; in a movie playing scenario, the audio to be played is the movie audio to be played; in a game scenario, the audio to be played is the background audio or game sound effects of the game.

[0070] It should be noted that the frequency at which the head-mounted device sends the audio to be played to the audio device and the frequency at which it sends the pose information can be the same or different, that is, the two types of data can be sent independently. The audio device can perform near-ear sound field rendering processing on the latest received audio to be played according to the latest received pose information.

[0071] The frequency at which the head-mounted device detects the pose information and the frequency at which it sends the pose information to the audio device can also be the same or different, and there is no limitation in this embodiment.

[0072] The audio device can be configured with a near-ear sound field rendering algorithm, and use the near-ear sound field rendering algorithm to perform near-ear sound field rendering processing on the audio to be played. During the near-ear sound field rendering processing, the optimal listening area of the sound field is adjusted according to the pose information, so that the nearest listening area overlaps with the position of the head-mounted device. There are many types of near-ear sound field rendering algorithms, and there is no limitation on which specific algorithm is used in this embodiment. The optimal listening area (also known as the sweet spot area) refers to the range of listening positions with the best sound field spatialization effect. In this area, the listener can obtain the most accurate sound source localization, frequency response balance, and immersive spatial audio experience. By adjusting the optimal listening area of the sound field to overlap with the position of the head-mounted device according to the pose information, the aim is to make the optimal listening area overlap with the positions of the user's ears, so that the user can experience a high-fidelity and personalized spatial audio experience. The audio device can determine the pose of the head-mounted device according to the pose information, and input this pose into the near-ear sound field rendering algorithm, so as to achieve the overlap of the optimal listening area of the adjusted sound field and the position of the head-mounted device.

[0073] It can be understood that when there is no audio to be played in the current running scenario, the pose detection operation or the pose information sending operation can be set not to be performed.

[0074] In this embodiment, there are many usage scenarios for the solution of the head-mounted device and the audio device to cooperate in near-ear audio rendering. For example, the relevant algorithms of this embodiment can be configured in the head-mounted device and audio device in the cinema for users to experience in the cinema, or users can purchase the head-mounted device and audio device configured with the relevant algorithms of this embodiment for use in the home scenario. There are also other usage scenarios, such as game rooms, VR / AR game cockpits, etc., and there is no limitation in this embodiment. As Figure 2 shown, it is a schematic diagram of the scenario where the user uses the head-mounted device and the audio device to experience near-ear sound effects. The user wears the head-mounted device in scenarios such as rooms, cinemas, game rooms, and game cockpits. The head-mounted device detects the pose information and sends it to the audio device. The user moves from the position at time T to the position at time T+1, and the pose information at time T+1 is different from the pose information at time T. It should be noted that the +1 in T+1 means that T+1 is the moment after time T and does not represent a unit of time.

[0075] As Figure 3 shown, in a feasible implementation, the IMU module or UWB module in the head-mounted device detects the pose information, and sends the pose information and the audio to be played to the audio device. The audio device renders the audio to be played using the near-ear sound field rendering algorithm according to the pose information, and then plays the rendered audio through the speaker.

[0076] In this embodiment, the pose detection device in the head-mounted device detects the pose information of the head-mounted device, and sends the pose information and the audio to be played in the head-mounted device to the audio device. The audio device adjusts the best listening area of the sound field according to the pose information to overlap with the position of the head-mounted device, performs near-ear sound field rendering processing on the audio to be played, and plays the processed audio, providing an audio playback solution that brings a better audio experience to users in the head-mounted device usage scenario. On the one hand, the solution in this embodiment plays the audio to be played in the head-mounted device through an external audio device, so that when the user uses the head-mounted device, the user can not only experience the functions and services provided by the head-mounted device, but also experience the better audio effect provided by the audio device compared with the micro speaker in the head-mounted device; on the other hand, since the user wears the head-mounted device, the solution in this embodiment dynamically detects the pose information by the head-mounted device and sends it to the audio device for the audio device to adjust the best listening area of the sound field to overlap with the position of the head-mounted device according to the pose information, performs near-ear sound field rendering processing on the audio to be played, provides a near-ear audio effect for the user, brings a better sense of sound immersion to the user, and ensures that the sound effect remains consistent during the process of the user wearing the head-mounted device and moving.

[0077] In a specific implementation, the head-mounted device can be communicatively connected to one audio device to implement the audio playback method in this embodiment, or can be communicatively connected to multiple audio devices to implement the audio playback method in this embodiment. In the scenario where the head-mounted device is communicatively connected to multiple audio devices, the head-mounted device can send the audio to be played and the pose information of the head-mounted device to each of the multiple audio devices respectively. The audio to be played sent to each audio device can be the same or different. The pose information sent to each audio device is respectively used to determine the relative pose between the head-mounted device and each audio device. Each audio device adjusts the best listening area of its own sound field according to the received pose information to make the best listening area overlap with the position of the head-mounted device, and each audio device respectively performs near-ear sound field rendering processing on the received audio to be played and plays the processed audio. The multiple audio devices can be distributed at different positions. Exemplarily, Figure 4The scenario where the head-mounted device is communicatively connected to two audio devices is shown. At time T, the head-mounted device sends pose information 1 to audio device 1 and sends pose information 2 to audio device 2. Based on pose information 1, the relative pose between the head-mounted device and audio device 1 can be determined, and based on pose information 2, the relative pose between the head-mounted device and audio device 2 can be determined. It should be noted that Figure 4 the number and distribution positions of the audio devices in this are just an example.

[0078] In a specific implementation, the head-mounted device may or may not be provided with a speaker, and this embodiment does not limit this. If a head-mounted device with a speaker is used, then, in a specific implementation, when a communication connection is established between the head-mounted device and the audio device, it can be set to preferentially play the audio to be played through the audio device, or play the audio to be played through the audio device in response to a control instruction triggered by the user; when no communication connection is established with the audio device, it can be set to play the audio to be played through the speaker provided in the head-mounted device.

[0079] In a feasible implementation, the audio playback method may further include: in response to a playback channel selection instruction, select a playback channel from the speaker of the head-mounted device and the audio device; when selecting the speaker of the head-mounted device as the playback channel, play the audio to be played through the speaker of the head-mounted device; when selecting the audio device as the playback channel, perform the step of sending the audio to be played and the pose information to the audio device. Among them, the channel selection instruction is used to indicate the playback channel, which can be triggered by the user, such as triggered by a physical button in the head-mounted device or triggered by voice interaction, or can be triggered according to a preset trigger condition. For example, it can be set that when a communication connection is established with the audio device, a channel selection instruction indicating the audio device as the playback channel is automatically triggered, and when no audio device is connected, a channel selection instruction indicating the speaker of the head-mounted device as the playback channel is automatically triggered.

[0080] In a feasible implementation manner, when the head-mounted device includes speakers and microphones disposed at the binaural positions, the speakers disposed at the binaural positions may block the audio played by the audio device to a certain extent, such that the near-ear audio effect felt by the user is affected to a certain extent. In response to this, the speakers and the microphones can cooperate to operate in a sound passthrough mode. The sound passthrough mode (also known as the transparency mode or the ambient sound enhancement mode) is a function that transmits the external ambient sound to the wearer's ears through technical means. In this implementation manner, this function is configured in the head-mounted device to enable, when the user uses the head-mounted device with speakers disposed at the binaural positions, while feeling the near-ear audio effect rendered by the external audio device, to avoid the near-ear audio effect being blocked by the speakers, and further improve the user experience. Specifically, the audio playback method further includes: when sending the audio to be played to the audio device for playback, collecting a sound signal through the microphone; performing sound passthrough processing on the sound signal, and playing the sound signal after the sound passthrough processing through the speaker.

[0081] Based on the above first embodiment, a second embodiment of the audio playback method of the present application is proposed. In this embodiment, for the content that is the same as or similar to the above first embodiment, reference can be made to the above introduction and will not be elaborated hereinafter. In a specific application scenario, when the user wears the head-mounted device, there may be a situation where the user slightly moves or turns the head (hereinafter referred to as a small movement) and then quickly returns to the original position. In this case, the head-mounted device sends the pose information after the small movement to the audio device, and the audio device processes and outputs the audio according to the pose information. However, if the hardware computing power of the audio device is limited or the communication speed is limited, during the process of the audio device processing and outputting the audio, before the processed audio is output, the user's head may have returned to the original position, that is, the user is actually not in the optimal listening area, which results in a deviation of the near-ear sound field effect felt by the user. If the user frequently makes small movements and then returns to the original position, there will be problems of sound field stuttering and instability. In this embodiment, considering this possible problem, it is proposed that after step S10, step S30 is further included, and step S20 includes step S201.

[0082] Step S30, determining whether the head-mounted device is in a large movement state, where the large movement state refers to a state in which the change amount of the pose information detected within a continuous first preset duration compared to the pose information sent to the audio device last time is greater than the preset change amount.

[0083] Step S201, if it is determined that the head-mounted device is in the large movement state, sending the latest detected pose information to the audio device.

[0084] After sending the pose information to the audio device once, the head-mounted device can compare all the subsequently detected pose information with the pose information sent this time, calculate the change amount compared to the pose information sent this time, and compare the change amount with a preset change amount. The preset change amount can be set as needed. When the change amount of the subsequently detected pose information compared to the pose information sent this time is less than the preset change amount, it indicates that the user has made small movements compared to the moment when the pose information was sent last time. It should be noted that when the pose information includes both position and attitude information, the preset change amount includes two change amount thresholds corresponding to the position and attitude respectively.

[0085] It should be noted that after establishing a communication connection with the audio device, the pose information detected for the first time by the head-mounted device can be directly sent to the audio device.

[0086] In this embodiment, a large movement state is defined. When it is determined that the head-mounted device is in the large movement state, the latest detected pose information is sent to the audio device so that the audio device can adjust the position of the optimal listening area based on the newly received pose information. The large movement state refers to a state where the change amount of the pose information detected within a continuous first preset duration compared to the pose information sent to the audio device last time is greater than the preset change amount. That is, if the actions that occur within the continuous first preset duration are not small actions, it can be determined that the device is in the large movement state. The first preset duration can be set as needed and is not limited in this embodiment. For example, it can be set that when the rotation angle is greater than 5°, the movement distance is greater than 5 cm, and the duration is 20 ms, it is determined that the device is in the large movement state.

[0087] In this embodiment, by judging the detected pose information, when it is determined that the head-mounted device is in the large movement state, the pose information is sent to the audio device, avoiding the problems of sound field stuttering and instability caused by sound field rendering delay when quickly returning to the original position after small movements.

[0088] Based on the above first and / or second embodiments, the third embodiment of the audio playback method of this application is proposed. In this embodiment, the same or similar content as in the above first and second embodiments can be referred to the above introduction and will not be elaborated hereinafter. In this embodiment, the audio playback method further includes steps S40 to S50, and step S20 includes step S202:

[0089] Step S40, obtaining target information, where the target information includes the first information of the current running scenario of the head-mounted device or the second information of the current running game.

[0090] The first piece of information is information related to the current running scenario, which is used to determine the detection frequency of the pose information. That is, in this embodiment, the detection frequency of the pose information can be dynamically determined according to the information related to the current running scenario, and the detection frequencies may be different in different running scenarios. In this embodiment, there is no limitation on what the first piece of information specifically is. For example, the first piece of information can be the detection frequency adapted to the current running scenario, which can be set in advance according to experience or user-defined; or, the first piece of information can be the type of the current running scenario, such as a music-playing scenario, a movie-playing scenario, a game scenario, etc.; or, the first piece of information can be information indicating whether the current running scenario needs to interact with the user. For example, 1 indicates that interaction is required, and 0 indicates that interaction is not required.

[0091] The second piece of information is information related to the current running game, which is used to determine the detection frequency of the pose information. That is, in this embodiment, the detection frequency of the pose information can be dynamically determined according to the information related to the current running game, and the detection frequencies may be different in different games or different scenarios in the game. In this embodiment, there is no limitation on what the second piece of information specifically is. For example, the second piece of information can be the detection frequency adapted to the type, progress stage, scenario, etc. of the current running game, which can be set in advance according to experience or user-defined; or, the second piece of information can be the type of the current running game, such as a shooting game, a music appreciation game, a map-running game, etc.; or, the second piece of information can be information indicating whether the current running game needs to interact with the user. For example, 1 indicates that interaction is required, and 0 indicates that interaction is not required.

[0092] In the specific implementation manner, when the current running scenario is a game, the target information can be the first piece of information or the second piece of information. When the current running scenario is other scenarios other than a game, the target information is the first piece of information.

[0093] Step S50: Dynamically update the detection frequency of the pose information according to the target information.

[0094] When the target information is the detection frequency itself, the target information can be used as the latest detection frequency. When the target information is other types of information, a mapping relationship between different information and the detection frequency can be set in advance. After obtaining the target information, the detection frequency corresponding to the target information is determined according to the mapping relationship as the latest detection frequency. The mapping relationship can be represented in the form of a table, a calculation formula, etc., and there is no limitation in this embodiment.

[0095] Step S202: Detect the pose information of the head-mounted device by the pose detection device in the head-mounted device according to the latest determined detection frequency.

[0096] In this embodiment, considering that the real-time requirements for the rendering of the near-ear sound field by the user are different in different operating scenarios of the head-mounted device or different scenarios of the game, the detection frequency of the pose information is dynamically determined according to the relevant information of the current operating scenario or the relevant information of the current running game, so that in the application scenario of the head-mounted device, the rendering effect of the near-ear sound field is more flexible. In some scenarios where the real-time requirements for the rendering of the near-ear sound field are not high, a lower detection frequency can be adopted, thereby reducing the power consumption of the head-mounted device and increasing the battery life of the head-mounted device.

[0097] In a feasible implementation manner, the first information includes information for indicating whether the current operating scenario needs to interact with the user. For example, the first information may be the type of the current operating scenario, such as a music playing scenario, a movie playing scenario, a game scenario. Some scenarios need to interact with the user, and some scenarios do not need to interact with the user. Scenarios that need to interact with the user and scenarios that do not need to interact with the user can be pre-configured. For example, the music playing scenario and the movie playing scenario can be configured as scenarios that do not require interaction, and the game scenario can be configured as a scenario that requires interaction; the head-mounted device obtains the type of the current operating scenario and can determine whether the current operating scenario needs to interact with the user according to the pre-configured information. For another example, the first information may be information such as 1 or 0 that directly indicates whether interaction with the user is required. The first information corresponding to different operating scenarios can be set in advance as needed. When the target information includes the first information, step S50 includes:

[0098] Step S501, dynamically update the detection frequency of the pose information according to the first information.

[0099] In this embodiment, there is no limitation on the basis for dynamically updating the detection frequency of the pose information according to the first information. A calculation formula, a mapping table, etc. can be set to represent the mapping relationship between the first information and the detection frequency. No matter what basis is set, the following relationship can be satisfied: the updated detection frequency in the first case is greater than the updated detection frequency in the second case. The first case refers to the case where the first information indicates that the current operating scenario needs to interact with the user, and the second case refers to the case where the first information indicates that the current operating scenario does not need to interact with the user.

[0100] In this embodiment, when the current running scenario requires interaction with the user, a higher detection frequency is adopted. When the user's head moves or rotates frequently during the interaction, the head-mounted device can detect the movement or rotation of the user's head more timely, and thus can send it to the audio device more timely, enabling the audio device to perform near-ear sound field rendering according to the latest pose of the user more timely, thereby avoiding the problems of sound field jitter and instability caused by sound field rendering delay, and improving the experience of the near-ear sound field effect when the user wears the head-mounted device for interaction. When the current running scenario does not require interaction with the user, a lower detection frequency is adopted. On the one hand, it avoids the problems of sound field jitter and instability caused by sound field rendering delay when the position quickly returns to the original position after a small movement. On the other hand, it can also avoid unnecessary power consumption by reducing the detection frequency, thereby increasing the battery life of the head-mounted device.

[0101] As Figure 5 shown, at time T, the running scenario of the head-mounted device is a game scenario, and the determined detection frequency is f1. At time T+1, the running scenario of the head-mounted device is a movie playback scenario, and the determined detection frequency is f2. f1 can be set to be greater than f2.

[0102] In a feasible embodiment, the second information includes the interaction information of the currently running game within a second preset duration in the future. The second preset duration can be set as needed and is not limited here. The interaction information can be information indicating whether the user needs to change the head pose within the second preset duration of the currently running game. For example, it can be 1 indicating that the head pose needs to be changed, and 0 indicating that the head pose does not need to be changed. The second information can be pre-set in the game configuration information, that is, the interaction information of each running stage of the game is pre-configured, and the interaction information of the currently running game within the second preset duration in the future is extracted according to the configuration information. Alternatively, the second information can also be determined by analyzing the in-game events. Specifically, by analyzing the in-game events, it is determined whether it meets a specific pre-set trigger condition. When the trigger condition is met, it is determined that interaction is required within the second preset duration in the future (thus obtaining the second information). For example, the trigger condition can be set to include triggering a running map task. When the target information includes the second information, step S50 includes:

[0103] Step S502, dynamically updating the detection frequency of the pose information according to the interaction information.

[0104] In this embodiment, there is no limitation on the basis for dynamically updating the detection frequency of the pose information according to the interaction information. A calculation formula, a mapping table, etc. can be set to represent the mapping relationship between the interaction information and the detection frequency. No matter what basis is set, the following relationship can be satisfied: the detection frequency updated in the third case is greater than the detection frequency updated in the fourth case. The third case refers to the situation where the interaction information indicates that the user needs to change the head pose within the next second preset duration, and the fourth case refers to the situation where the interaction information indicates that the user does not need to change the head pose within the next second preset duration.

[0105] In this embodiment, when the currently running game needs to interact with the user within the next second preset duration, the detection frequency is updated in advance to a larger value, so that when the user's head moves or rotates frequently during the game interaction, the head-mounted device can detect the movement or rotation of the user's head more timely, and thus can send it to the audio device more timely, enabling the audio device to perform near-ear sound field rendering according to the user's latest pose more timely, thereby avoiding the problems of sound field stuttering and instability caused by sound field rendering delay, and improving the experience of the near-ear sound field effect when the user wears the head-mounted device to play games. When the currently running game does not need to interact with the user within the next second preset duration, the detection frequency is updated in advance to a smaller value. On the one hand, it avoids the problems of sound field stuttering and instability caused by sound field rendering delay in the case of quickly returning to the original position after a small movement, and on the other hand, it can also avoid unnecessary power consumption by reducing the detection frequency, thereby increasing the battery life of the head-mounted device.

[0106] As Figure 6 shown, at time T, the game currently running on the head-mounted device needs to interact within the next preset second preset duration, and the determined detection frequency is f1. At time T + 1, the game currently running on the head-mounted device does not need to interact within the next preset second preset duration, and the determined detection frequency is f2, and f1 can be set to be greater than f2.

[0107] Based on the above first, second, and / or third embodiments, a fourth embodiment of the audio playback method of the present application is proposed. In this embodiment, the same or similar content as the above first, second, and third embodiments can be referred to the above introduction and will not be repeated hereinafter. In this embodiment, the audio playback method is applied to an audio device. The audio device is communicatively connected to the head-mounted device, and the audio device and the head-mounted device are separately arranged. The audio playback method includes:

[0108] Step A10, receiving the audio to be played and the pose information sent by the head-mounted device, where the pose information is detected by a pose detection device in the head-mounted device.

[0109] Step A20: Adjust the optimal listening area of the sound field according to the pose information, overlap the optimal listening area with the position of the head-mounted device, and perform near-ear sound field rendering processing on the audio to be played and play the processed audio.

[0110] In this embodiment, the specific implementation manners of steps A10 to A20 can refer to the specific implementation manners of the audio playback method applied to the head-mounted device in the above embodiments, and will not be elaborated here.

[0111] In a feasible implementation manner, the audio playback method further includes:

[0112] Step A30: Send feedback information to the head-mounted device.

[0113] In addition to receiving the audio to be played and the pose information sent by the head-mounted device, the audio device can also perform other information interactions with the head-mounted device, including sending feedback information to the head-mounted device. The specific content of the feedback information is not limited in this implementation manner and can be set according to the interaction requirements between the head-mounted device and the audio device. For example, the feedback information can include the playback volume, the device information of the audio device, the reception status of the audio to be played and the pose information, etc.

[0114] In the specific implementation manner, after receiving the feedback information sent by the audio device, the head-mounted device can perform corresponding processing on the feedback information, such as storage, display, output to other devices (such as the user's mobile phone), etc. The processing operations taken by the head-mounted device on the feedback information can be specifically set according to the interaction requirements between the head-mounted device and the audio device.

[0115] An embodiment of the present application provides a head-mounted device. A pose detection device is provided on the head-mounted device. The head-mounted device further includes: a memory, a processor, and an audio playback program stored on the memory and executable on the processor. When the audio playback program is executed by the processor, it implements the steps of the audio playback method applied to the head-mounted device in the above embodiments.

[0116] Compared with the prior art, the beneficial effects of the head-mounted device provided by the embodiment of the present application are the same as those of the audio playback method applied to the head-mounted device provided by the above embodiment, and other technical features in the head-mounted device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0117] An embodiment of the present application provides an audio device. The audio device includes: a memory, a processor, and an audio playback program stored on the memory and executable on the processor. When the audio playback program is executed by the processor, it implements the steps of the audio playback method applied to the audio device in the above embodiments.

[0118] Compared with the prior art, the beneficial effects of the head-mounted device provided in the embodiments of the present application are the same as those of the audio playback method applied to the audio device provided in the above embodiments, and other technical features in the head-mounted device are the same as the features disclosed in the method of the above embodiments, which will not be elaborated herein.

[0119] The embodiments of the present application provide an audio playback system, including an audio device and a head-mounted device as in the above embodiments. Among them, the audio device is used to execute the steps of the audio playback method applied to the audio device in the above embodiments, and the head-mounted device is used to execute the steps of the audio playback method applied to the head-mounted device in the above embodiments.

[0120] Compared with the prior art, the beneficial effects of the audio playback system provided in the embodiments of the present application are the same as those of the audio playback method provided in the above embodiments, and other technical features in the audio playback system are the same as the features disclosed in the method of the above embodiments, which will not be elaborated herein.

[0121] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0122] The embodiments of the present application provide a computer-readable storage medium, having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the audio playback method in the above embodiments.

[0123] The computer-readable storage medium provided by the embodiments of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0124] The above computer-readable storage medium may be included in the head-mounted device; or it may exist separately without being assembled into the head-mounted device.

[0125] The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed by the head-mounted device, the head-mounted device is caused to perform the above functions defined in the method of the disclosed embodiments of the present application.

[0126] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0128] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0129] The readable storage medium provided by the embodiments of the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above audio playback method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the embodiments of the present application are the same as those of the audio playback method provided by the above embodiments, and will not be elaborated here.

[0130] The embodiments of the present application also provide a computer program product, including a computer program, and the steps of the above audio playback method are implemented when the computer program is executed by a processor.

[0131] Compared with the prior art, the beneficial effects of the computer program product provided by the embodiments of the present application are the same as those of the audio playback method provided by the above embodiments, and will not be elaborated here.

[0132] The above are only some embodiments of the present application, and do not limit the patent scope of the present application accordingly. All equivalent structural transformations made under the technical concept of the present application by using the content of the specification and drawings of the present application, or directly / indirectly applied in other related technical fields, are included in the patent protection scope of the present application.

Claims

1. An audio playback method, characterized in that: The audio playing method is applied to a head mounted device, the head mounted device establishes a communication connection with a separately arranged audio device, and the audio playing method comprises: Detecting the posture information of the head-mounted device by a posture detection device in the head-mounted device; The audio to be played and the posture information are sent to the audio device, so that the audio device adjusts the best listening area of ​​the sound field according to the posture information, so that the best listening area overlaps with the position of the head-mounted device, performs near-ear sound field rendering processing on the audio to be played, and plays the processed audio.

2. The audio playback method according to claim 1, wherein: After the step of detecting the posture information of the head mounted device by the posture detection device in the head mounted device, the method further includes: Determine whether the head mounted device is in a significant movement state, where the significant movement state refers to a state in which a change amount between the posture information detected within a first preset time period and the posture information last sent to the audio device is greater than a preset change amount; The step of sending the posture information to the audio device comprises: If it is determined that the head mounted device is in the large movement state, the latest detected posture information is sent to the audio device.

3. The audio playback method according to claim 1, wherein: The audio playing method further comprises: Acquire target information, where the target information includes first information of a scene currently running on the head mounted device or second information of a game currently running on the head mounted device; Dynamically update the detection frequency of the posture information according to the target information; The step of detecting the posture information of the head mounted device by the posture detection device in the head mounted device comprises: The posture information of the head mounted device is detected by a posture detection device in the head mounted device according to a latest determined detection frequency.

4. The audio playback method according to claim 3, characterized in that: The first information includes information for indicating whether the current running scene needs to interact with the user. When the target information includes the first information, the step of dynamically updating the detection frequency of the posture information according to the target information includes: The detection frequency of the posture information is dynamically updated according to the first information, wherein the updated detection frequency in the first case is greater than the updated detection frequency in the second case, the first case refers to the case where the first information indicates that the current running scene requires interaction with the user, and the second case refers to the case where the first information indicates that the current running scene does not require interaction with the user.

5. The audio playback method according to claim 3, characterized in that: The second information includes interactive information of the currently running game within a second preset time period in the future. When the target information includes the second information, the step of dynamically updating the detection frequency of the posture information according to the target information includes: The detection frequency of the posture information is dynamically updated according to the interaction information, wherein the updated detection frequency in the third case is greater than the updated detection frequency in the fourth case, the third case refers to the case where the interaction information indicates that the user needs to change the head posture within the second preset time period in the future, and the fourth case refers to the case where the interaction information indicates that the user does not need to change the head posture within the second preset time period in the future.

6. The audio playback method according to claim 1, wherein: The posture detection device includes a UWB module and an IMU module, and the step of detecting the posture information of the head-mounted device by the posture detection device in the head-mounted device includes: When detecting the posture information for the first time after establishing a connection with the audio device, detecting the initial relative posture between the head mounted device and the audio device as the posture information through the UWB module; When it is not the first time to detect the posture information after establishing a connection with the audio device, the posture change of the head-mounted device is detected by the IMU module, and the posture change is used as the posture information.

7. The audio playback method according to any one of claims 1 to 6, characterized in that: In the case where the head mounted device includes a speaker and a microphone disposed at binaural positions, the audio playback method further includes: When the audio to be played is sent to the audio device for playing, collecting a sound signal through the microphone; The sound signal is subjected to sound transparent transmission processing, and the sound signal after the sound transparent transmission processing is played through the speaker.

8. The audio playback method according to any one of claims 1 to 6, characterized in that: The audio playing method further comprises: In response to a playback channel selection instruction, selecting a playback channel from a speaker of the head mounted device and the audio device; When the speaker of the head mounted device is selected as the playback channel, the audio to be played is played through the speaker of the head mounted device; When the audio device is selected as the playback channel, the step of sending the audio to be played and the posture information to the audio device is performed.

9. An audio playing method, characterized in that: The audio playback method is applied to an audio device, the audio device establishes a communication connection with a head-mounted device, the audio device and the head-mounted device are separately arranged, and the audio playback method includes: Receiving the audio to be played and the posture information sent by the head mounted device, wherein the posture information is detected by a posture detection device in the head mounted device; The best listening area of ​​the sound field is adjusted according to the posture information so that the best listening area overlaps with the position of the head mounted device, and near-ear sound field rendering processing is performed on the audio to be played and the processed audio is played.

10. The audio playing method according to claim 9, characterized in that: The audio playing method further comprises: Send feedback information to the head mounted device.

11. A head mounted device, characterized in that: The head mounted device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the audio playback method according to any one of claims 1 to 8.

12. An audio device, characterized in that: The audio device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the audio playing method according to any one of claims 9 and 10.

13. An audio playback system, characterized in that: The audio playback system comprises the head mounted device as claimed in claim 11 and the audio device as claimed in claim 12.

14. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the audio playback method according to any one of claims 1 to 10 are implemented.

15. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the audio playing method according to any one of claims 1 to 10 are implemented.