Audio playing method, device and equipment and computer readable storage medium

By mapping the real physical environment into the virtual reality environment and adjusting the frequency and rendering parameters of the audio signal based on translation speed and position, the problem of poor audio playback in virtual reality is solved, improving the realism of spatial sound effects and user experience.

CN121645127APending Publication Date: 2026-03-10MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies suffer from poor audio playback quality in virtual reality content, low differentiation between sound and the real environment, lack of consideration for personalized needs, and an imperfect virtual sound source position adjustment mechanism, resulting in a reduced user experience.

Method used

By mapping the real physical environment to the virtual reality environment based on the mapping relationship, the user's translation speed and position in the virtual reality environment are determined. The audio signal is frequency transformed based on the translation speed, and the audio signal is rendered using rendering parameters including distance, azimuth, pitch angle and loudness, thereby enhancing the realism of spatial sound effects.

Benefits of technology

It improves audio playback in virtual reality environments, enhances the realism of spatial sound effects, meets users' personalized audio experience needs, and improves auditory immersion and positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121645127A_ABST
    Figure CN121645127A_ABST
Patent Text Reader

Abstract

The invention discloses an audio playing method, device and equipment and a computer readable storage medium, and is applied to the technical field of computers, and the method comprises the steps: mapping a real physical environment to a virtual reality environment based on a mapping relation; determining the translation speed and position of the user in the virtual reality environment; performing frequency conversion on the audio signal based on the translation speed to obtain an audio signal after frequency conversion, and determining a rendering parameter based on the position; wherein the rendering parameters at least comprise a distance parameter, an azimuth angle, a pitch angle and loudness; and rendering the audio signal after frequency conversion based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal. According to the translation speed of the head of the user in the virtual reality environment, frequency conversion is carried out according to the Doppler effect, the sense of reality of the spatial sound effect is enhanced, and the sense of reality of the spatial sound effect is enhanced according to the rendering parameters such as the distance parameter, the azimuth angle, the pitch angle and the loudness of the rendering algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to an audio playing method, device, equipment and computer readable storage medium. BACKGROUND

[0002] The sound environment of the current virtual reality content sound source is not effectively restored, and the played sound is not realistic enough. There are problems of low distinction between sound in virtual reality environment and sound in real environment, poor audio rendering quality, lack of personalized demand consideration, and imperfect virtual sound source position adjustment mechanism.

[0003] It can be seen that how to improve the audio playing effect is a technical problem that technicians in the field urgently need to solve. SUMMARY

[0004] Therefore, the purpose of the present application is to provide an audio playing method, device, equipment and computer readable storage medium, which solves the technical problem of poor audio playing effect in the prior art.

[0005] To solve the above technical problems, the present application provides an audio playing method, comprising:

[0006] mapping a real physical environment to a virtual reality environment based on a mapping relationship;

[0007] determining the translation speed and position of a user in the virtual reality environment;

[0008] frequency transforming an audio signal based on the translation speed to obtain a frequency transformed audio signal, and determining a rendering parameter based on the position; wherein the rendering parameter at least includes a distance parameter, an azimuth angle, an elevation angle and a loudness;

[0009] rendering the frequency transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal.

[0010] Optionally, before mapping the real physical environment to the virtual reality environment based on the mapping relationship, it further comprises:

[0011] obtaining real-time pose information of a user's head in a physical space, and determining an offset vector of the real-time pose information relative to the HOA sphere center;

[0012] converting the offset vector into an offset angle in a spherical coordinate system, and updating the spherical coordinates of a loudspeaker group based on the offset angle;

[0013] determining an updated spherical coordinate basis based on the updated spherical coordinates of the loudspeaker group;

[0014] Correspondingly, the frequency-transformed audio signal is rendered based on the rendering parameter to obtain a rendered audio signal, and the rendered audio signal is played, including:

[0015] The frequency-transformed audio signal is rendered based on the updated spherical coordinate basis and the rendering parameter to obtain the rendered audio signal.

[0016] Optionally, real-time pose information of a user's head in a physical space is acquired, and an offset vector of the real-time pose information relative to a HOA sphere center is determined, including:

[0017] Real-time physical coordinates of a listener's head are acquired in real time through an inertial measurement unit acquisition device of a head-mounted device in combination with an environment identifier calculated by a physical space;

[0018] Real-time pose information of a user's head is analyzed in real time through a camera in a physical environment based on the physical coordinates of the listener's head;

[0019] Historical pose data of a user's head is determined;

[0020] The real-time pose information of the user's head is normalized to obtain normalized real-time pose information.

[0021] The historical pose data is filtered to obtain filtered historical pose data.

[0022] The normalized real-time pose information and the filtered historical pose data are fused to obtain comprehensive real-time pose information of the head;

[0023] An offset vector of the comprehensive real-time pose information of the head relative to a HOA sphere center is determined.

[0024] Optionally, a real physical environment is mapped to a virtual reality environment based on a mapping relationship, including:

[0025] It is determined whether a playback environment is a virtual reality environment;

[0026] When the playback environment is a virtual reality environment, the real physical environment is mapped to the virtual reality environment based on a mapping relationship between the real physical environment and the virtual reality environment determined by using normalized spherical harmonics.

[0027] Optionally, the audio signal is frequency-transformed based on the translation speed to obtain a frequency-transformed audio signal, and the rendering parameter is determined based on the position, including:

[0028] The distance of a user's head relative to a sound source is determined;

[0029] The distance parameter of a rendering algorithm is adjusted according to the distance;

[0030] adjust the azimuth and the elevation of the rendering algorithm according to the ear orientation;

[0031] adjust the loudness of the rendering parameter according to the position and the ear orientation.

[0032] Optionally, before rendering the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal, the method further comprises:

[0033] obtaining a corresponding reverberation time, reverberation intensity and ambient brightness according to a category of the current virtual reality environment;

[0034] Correspondingly, rendering the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal, comprises:

[0035] adjusting a reverberation effect of the frequency-transformed audio signal according to the reverberation time and the reverberation intensity, and adjusting a brightness of the frequency-transformed audio signal according to the ambient brightness to obtain a reverberation-adjusted audio signal;

[0036] rendering the reverberation-adjusted audio signal based on the rendering parameter to obtain the rendered audio signal.

[0037] Optionally, frequency-transforming the audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determining the rendering parameter based on the position, comprises:

[0038] detecting a translation speed of a user's head;

[0039] determining a Doppler frequency transformation parameter according to the translation speed.

[0040] frequency-transforming the audio signal based on the Doppler frequency transformation parameter to obtain the frequency-transformed audio signal.

[0041] The present application also provides an audio playing device, comprising:

[0042] a mapping module configured to map a real physical environment to a virtual reality environment based on a mapping relationship;

[0043] a translation speed and position determination module configured to determine a translation speed and a position of a user in the virtual reality environment;

[0044] a rendering parameter determination module configured to frequency-transform an audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determine a rendering parameter based on the position; wherein the rendering parameter at least comprises a distance parameter, an azimuth, an elevation and a loudness.

[0045] render the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and play the rendered audio signal.

[0046] The application further provides an audio playing device, comprising:

[0047] a memory for storing a computer program;

[0048] a processor for executing the computer program to implement the steps of the audio playing method.

[0049] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the audio playing method.

[0050] The application further provides a computer program product, comprising a computer program / instruction, and the computer program / instruction is executed by a processor to implement the steps of the audio playing method.

[0051] It can be seen that the application maps a real physical environment to a virtual reality environment based on a mapping relationship, determines a translation speed and a position of a user in the virtual reality environment, frequency-transforms an audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determines a rendering parameter based on the position, wherein the rendering parameter at least comprises a distance parameter, an azimuth angle, an elevation angle and a loudness; the frequency-transformed audio signal is rendered based on the rendering parameter to obtain a rendered audio signal, and the rendered audio signal is played. The application firstly maps a real physical space of a reconstructed three-dimensional sound field to a virtual reality environment, determines a translation speed and a position of a user in the virtual reality environment, and then frequency-transforms the three-dimensional sound field according to the translation speed of the head of the user in the virtual reality environment to enhance the reality of the spatial sound effect; according to the position of the head of the user relative to a sound source in the virtual reality environment, the distance, the loudness, the azimuth angle and the elevation angle of the rendering algorithm and other parameters are controlled to enhance the reality of the spatial sound effect.

[0052] In addition, the application further provides an audio playing device, an audio playing apparatus and a computer readable storage medium, which also have the beneficial effects. BRIEF DESCRIPTION OF DRAWINGS

[0053] In order to make the technical solutions of the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the accompanying drawings in the following description only aim to explain the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative work on the basis of the provided drawings.

[0054] Figure 1 A flowchart of an audio playing method provided by an embodiment of the present application;

[0055] Figure 2 A flowchart of an audio playing method provided by an embodiment of the present application;

[0056] Figure 3 A structural schematic diagram of an audio playing device provided by an embodiment of the present application;

[0057] Figure 4 A structural schematic diagram of an audio playing device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the technical solutions of the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the accompanying drawings in the following description only aim to explain the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative work on the basis of the provided drawings.

[0059] Please refer to Figure 1 , Figure 1 A flowchart of an audio playing method provided by an embodiment of the present application. The method can include:

[0060] S101, mapping a real physical environment to a virtual reality environment based on a mapping relationship.

[0061] Each step in the embodiment can be executed by a specified electronic device, which can be a server, a portable terminal or other forms. The mapping relationship in the embodiment can mean that based on normalized spherical harmonics, the spatial relationship between the sound environment of a sound source and the played sound environment is corresponded. That is, according to the mapping relationship, the position and direction in the physical space are mapped to the virtual reality environment; for example, wearing a VR (virtual display) head-mounted display to watch a moving bee, the user's head will instinctively dodge, and the sound direction heard thereby needs to be consistent with the vision, so as to avoid the phenomenon of dizziness sound.

[0062] It needs to be further explained that based on any of the above embodiments, before mapping the real physical environment to the virtual reality environment based on the mapping relationship, the following can also be included:

[0063] Step 1, obtaining real-time pose information of the user's head in the physical space, determining the offset vector of the real-time pose information relative to the HOA sphere center.

[0064] This embodiment obtains real-time pose information of the user's head in the physical space, and determines the offset of the real-time pose information of the head relative to the HOA (Higher Order Ambisonics) sphere center; the HOA sphere center corresponds to the spherical center position in the played physical environment. For example, when playing, the default orientation at the sphere center. Sitting on a chair, the physical is distributed on the sphere, and the offset of the loudspeaker on the sphere is adjusted, and the virtual adjustment is adjusted, and the several simulated virtual loudspeakers are repositioned on the sphere.

[0065] Step 2, convert the offset vector to an offset angle in the spherical coordinate system, and update the spherical coordinates of the loudspeaker group based on the offset angle.

[0066] This embodiment applies the offset to the spherical coordinate calculation of the loudspeaker group in the rendering algorithm in real time, adds the original spherical coordinates and the offset vector, and simply calculates it. This embodiment converts the offset vector to an offset angle in the spherical coordinate system; when the user's head deviates from the HOA sphere center, the HOA loudspeaker distribution has deviated from the sphere and is in a special-shaped distribution. To maintain the balance and stability of hearing, it is necessary to update the spherical coordinate position and angle of the loudspeaker group in the rendering algorithm. Update the spherical coordinates of the loudspeaker group according to the offset angle.

[0067] Step 3, determining the updated spherical coordinate basis based on the updated spherical coordinates of the loudspeaker group.

[0068] This embodiment can use an interpolation algorithm to calculate the updated spherical coordinate basis. When the loudspeaker group is in a special-shaped distribution, the loudspeaker group of the rendering algorithm is not one-to-one corresponding to the physical loudspeaker group, and needs to be fitted to the state of zero offset vector through interpolation calculation of the surrounding several (1-3) loudspeakers. Correspondingly, rendering the frequency-transformed audio signal based on the updated spherical coordinate basis and the rendering parameters to obtain the rendered audio signal, playing the rendered audio signal can include: rendering the frequency-transformed audio signal based on the updated spherical coordinate basis and the rendering parameters to obtain the rendered audio signal. This embodiment renders the frequency-transformed audio signal based on the updated spherical coordinate basis and the rendering parameters, which improves the accuracy of rendering.

[0069] It needs to be further explained that based on any of the above embodiments, the above obtaining real-time pose information of the user's head in the physical space, determining the offset vector of the real-time pose information relative to the HOA sphere center, can include:

[0070] Step 1: Real-time acquisition of the physical coordinates of the listener's head through the inertial measurement unit acquisition device of the head-mounted device combined with the environmental identifier calculated by the physical space.

[0071] This embodiment acquires the physical coordinates of the listener's head in real time through the IMU (Inertial Measurement Unit) acquisition device of the head-mounted device combined with the environmental identifier calculated by the physical space.

[0072] Step 2: Real-time analysis of the real-time pose information of the user's head based on the physical coordinates of the listener's head through the camera in the physical environment.

[0073] This embodiment analyzes the pose information of the user's head in real time through the camera (such as a TOF (Time of Flight) camera) in the physical environment.

[0074] Step 3: Determine the historical pose data of the user's head.

[0075] This embodiment needs to save the historical pose data of the user's head for calculating the motion change of the user's head pose.

[0076] Step 4: Normalize the real-time pose information of the user's head to obtain normalized real-time pose information.

[0077] This embodiment normalizes the acquired physical coordinates and converts them to a normalized coordinate system, which is referred to as a normalized coordinate system.

[0078] Step 5: Filter the historical pose data to obtain filtered historical pose data.

[0079] This embodiment filters the historical pose data to extract the stable part; it is equivalent to low-frequency filtering, ignoring the interference of transient jitter data.

[0080] Step 6: Fuse the normalized real-time pose information and the filtered historical pose data to obtain comprehensive real-time head pose information.

[0081] This embodiment fuses the normalized physical coordinates and the filtered historical pose data to obtain comprehensive real-time head pose information

[0082] Step 7: Determine the offset vector of the comprehensive real-time head pose information relative to the HOA sphere center.

[0083] The real-time pose information in this embodiment is real-time information determined based on normalized real-time pose information and filtered historical pose data, thereby improving the accuracy of the determination of the head real-time pose information, and thus the accuracy of the offset vector. When determining the offset vector, the linear distance of the user's head relative to the center of the HOA sphere can be calculated; the angle of the head offset from the center of the HOA sphere can be calculated; and the distance and the angle are combined to form the offset vector.

[0084] It should be further explained that, based on any of the above embodiments, the above mapping of the real physical environment to the virtual reality environment based on the mapping relationship can include: determining whether the playback environment is a virtual reality environment; and when the playback environment is a virtual reality environment, mapping the real physical environment to the virtual reality environment based on the mapping relationship between the real physical environment and the virtual reality environment determined using the normalized spherical harmonic function. This embodiment selects a corresponding three-dimensional sound field reconstruction method according to the type of the playback environment, which specifically refers to distinguishing whether the sound environment of the virtual reality content or the sound environment of the non-virtual reality content (such as 2D movies or music) is to be played back. It is determined whether virtual reality environment mapping is needed, which is determined by the type of the program played (whether it is virtual reality content), and the type of the program played is a default known input. In this embodiment, mapping is only performed when the playback environment is a virtual reality environment, because the audio playback of the virtual reality environment cannot be directly played using the audio in the real environment.

[0085] S102, determining the translation speed and position of the user in the virtual reality environment.

[0086] This embodiment determines the translation speed and relative position of the user in the virtual environment. For example, the linear motion speed of the head is calculated through IMU data processing. This embodiment determines the relative position of the user's head relative to the sound source in the virtual reality environment.

[0087] S103, frequency transforming the audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determining rendering parameters based on the position; wherein the rendering parameters at least include a distance parameter, an azimuth angle, an elevation angle, and a loudness.

[0088] This embodiment performs frequency transformation according to the translation speed of the head, which can refer to the Doppler effect formula for frequency transformation. This embodiment can adjust the rendering parameters according to the position of the head relative to the sound source and the orientation of the ear.

[0089] It should be further explained that, based on any of the above embodiments, the above frequency transformation of the audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determining the rendering parameters based on the position can include:

[0090] S1031, determining the distance of the user's head relative to the sound source;

[0091] S1032, adjusting a distance parameter of the rendering algorithm according to the distance;

[0092] S1033, adjusting an azimuth angle and an elevation angle of the rendering algorithm according to the ear orientation;

[0093] S1034, adjusting a loudness of the rendering parameter according to the position and the ear orientation.

[0094] The embodiment determines the distance of the user's head relative to the sound source; adjusts the distance parameter of the rendering algorithm according to the distance, adjusts the azimuth angle and the elevation angle of the rendering algorithm according to the ear orientation, and adjusts the loudness of the audio rendering according to the relative position and the orientation. It can be understood that the movement of the listener's head and the change of the ear orientation can affect the virtual reality environment and change the characteristics of the sound source, such as distance, azimuth angle, elevation angle, and loudness. The embodiment provides a specific method for determining each rendering parameter, which improves the accuracy of determining the rendering parameter.

[0095] It needs to be further explained that, based on any of the above embodiments, the above frequency transformation of the audio signal based on the translation speed, obtaining the frequency-transformed audio signal, and determining the rendering parameter based on the position, includes: detecting the translation speed of the user's head; determining the Doppler frequency transformation parameter according to the translation speed. The audio signal is frequency-transformed based on the Doppler frequency transformation parameter to obtain the frequency-transformed audio signal. The embodiment detects the translation speed of the user's head; calculates the Doppler frequency transformation parameter according to the translation speed, and performs Doppler transformation processing on the audio signal. For ease of understanding, please refer to Table 1, which is a frequency transformation table provided by an embodiment of the present application.

[0096] Table 1: A frequency transformation table

[0097]

[0098] S104, rendering the frequency-transformed audio signal based on the rendering parameter to obtain the rendered audio signal, and playing the rendered audio signal.

[0099] The execution subject of the embodiment is a playback device, which can play the rendered audio signal based on the playback device. The spherical coordinate basis is updated to render the spatial audio data; the rendering data is encoded; the encoded data is transmitted to the playback device; and the audio is decoded and output according to the playback environment type.

[0100] It needs to be further explained that, based on any of the above embodiments, before the frequency-transformed audio signal is rendered based on the rendering parameter to obtain a rendered audio signal, and the rendered audio signal is played, it can further include: obtaining the corresponding reverberation time, reverberation intensity and environment brightness according to the category of the current virtual reality environment; correspondingly, rendering the frequency-transformed audio signal based on the rendering parameter to obtain the rendered audio signal, and playing the rendered audio signal can include: adjusting the reverberation effect of the frequency-transformed audio signal according to the reverberation time and the reverberation intensity, and adjusting the brightness of the frequency-transformed audio signal according to the environment brightness to obtain a reverberation-adjusted audio signal; rendering the reverberation-adjusted audio signal according to the rendering parameter to obtain the rendered audio signal. This embodiment distinguishes whether the sound environment of the virtual reality content or the sound environment of the non-virtual reality content (such as 2D movie or music) to be played back according to the category of the playback environment, and adjusts the reverberation parameter; the reverberation is one of the key parameters of the sound environment, and has a great influence on hearing. Therefore, the application increases the adjustment of the reverberation parameter when reconstructing the sound field, so as to restore the reverberation sound effect of the sound source as much as possible. This embodiment can obtain the corresponding reverberation time and reverberation intensity according to the category of the virtual reality environment, for example, during program production, a sound mixer can export the environment reverberation information of the corresponding scene in the virtual reality and fill it into the metadata. If the metadata lacks description of the reverberation parameter, the player can also automatically analyze through the picture sound combination and set several gears as the sound effect control input parameters of rendering. The reverberation effect of the audio rendering is adjusted according to the environment type, for example, the reverberation parameters of the playback sound environment measured in advance and the reverberation parameters in the sound environment of the sound source are compared to determine whether the reverberation sound effect of the rendering playback needs to be enhanced or weakened, the purpose is to approach the original reverberation sound effect. The brightness of the audio rendering is adjusted according to the environment brightness, which considers that the frequency response of the original sound source in different environments will also affect the change of the hearing sound effect, so active compensation can be performed, and the purpose is to approach the performance of the original sound source. For example: treble equalizer: the higher the environment, the higher the gain of the high frequency band; the darker the environment, the higher the gain of the high frequency band; filter cutoff frequency: low-pass filter. The higher the environment, the higher the cutoff frequency of the filter, so that more high frequencies pass through; the darker the environment, the lower the cutoff frequency, so that more high frequencies are filtered out, making the sound dull.

[0101] The audio playing method provided by the embodiment of the present application can comprise: S101, mapping a real physical environment to a virtual reality environment based on a mapping relationship; S102, determining a translation speed and a position of a user in the virtual reality environment; S103, performing frequency transformation on an audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determining a rendering parameter based on the position; wherein the rendering parameter at least comprises a distance parameter, an azimuth angle, an elevation angle and a loudness; S104, rendering the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal. For the spatial audio in the virtual reality environment, the present application firstly performs mapping processing on the real physical space for reconstructing a three-dimensional sound field and the virtual reality environment, determines the translation speed and the position of the user in the virtual reality environment, and then reconstructs the three-dimensional sound field: according to the translation speed of the head of the user in the virtual reality environment, performs frequency transformation to enhance the reality of the spatial audio effect; according to the position of the head of the user in the virtual reality environment relative to the sound source, controls the distance, the loudness, the azimuth angle and the elevation angle and other parameters of the rendering algorithm to enhance the reality of the spatial audio effect.

[0102] The prior art has the following disadvantages: (1) lack of effective distinction between sound in the virtual reality environment and sound in the real environment, resulting in reduced user experience; the reason is that the original sound environment information is not transmitted to the playing end. (2) The virtual reality device has poor quality in audio rendering, especially in complex scenes, the positioning accuracy in visual and dark light scenes is difficult to guarantee; poor playing environment will affect the tracking effect of the IMU device. (3) The existing spatial audio technology lacks consideration of the individual needs of different users, and cannot meet the individual audio experience needs of users; (4) in the virtual reality environment, the integration of real-time pose information of the user and the spatial audio rendering system is not effective enough, and it is difficult to fully improve the auditory immersion and reality of the user. (4) The existing virtual sound source position adjustment mechanism is not perfect, and cannot flexibly adjust the position of the virtual sound source according to the specific needs of the user and the change of the environment, thereby affecting the playing effect of the spatial audio.

[0103] In order to make the present application more convenient to understand, please refer to Figure 2 , Figure 2 The flowchart of the audio playing method provided by the embodiment of the present application can specifically comprise:

[0104] S201, obtaining real-time pose information of the head of the user in the physical space.

[0105] The embodiment can acquire the physical coordinates of the listener's head in real time through the IMU acquisition device of the head-mounted device in combination with the environment identifier calculated by the physical space; the embodiment can analyze the real-time pose information of the user's head in real time based on the physical coordinates of the listener's head through the camera (such as a TOF camera) in the physical environment; and the historical pose data of the user's head is saved.

[0106] In S202, the real-time pose information of the user's head is normalized to obtain normalized real-time pose information, and the normalized real-time pose information is fused with the filtered historical pose data to obtain comprehensive real-time pose information of the head.

[0107] In the embodiment, the real-time pose information of the head is normalized to obtain normalized real-time pose information. The historical pose data is filtered to obtain filtered historical pose data. The normalized real-time pose information is fused with the filtered historical pose data to obtain comprehensive real-time pose information of the head.

[0108] In S203, the offset of the head relative to the center of the HOA sphere is determined based on the comprehensive real-time pose information of the head.

[0109] In the embodiment, the linear distance of the user's head relative to the center of the HOA sphere is calculated; the included angle of the head deviating from the center of the HOA sphere is calculated; and the distance and the included angle are combined to form the offset.

[0110] In S204, the offset is applied to the spherical coordinate calculation of the loudspeaker group in the rendering algorithm in real time to determine an updated spherical coordinate base.

[0111] In the embodiment, the offset vector is converted into an offset angle in the spherical coordinate system; when the user's head deviates from the center of the HOA sphere, it means that the distribution of the loudspeakers of the HOA has deviated from the spherical surface and is in a special-shaped distribution. To maintain the balance and stability of the hearing, it is necessary to update the spherical coordinate position and angle of the loudspeaker group in the rendering algorithm. The spherical coordinates of the loudspeaker group are updated according to the offset angle; and the updated spherical coordinate base is calculated by using an interpolation algorithm.

[0112] In S205, the playback environment type is determined.

[0113] In the embodiment, if it is determined that the playback environment is a virtual reality environment, S206 is performed, and if it is the current physical environment, S208 is performed. In the embodiment, it is determined whether the playback environment is a virtual reality environment.

[0114] In S206, it is determined whether the virtual reality environment mapping is needed based on the playback environment type, and if yes, S207 is performed, and if no, S208 is performed.

[0115] In the embodiment, the playback program type (whether it is virtual reality content) determines the playback program type, which is a default known input.

[0116] S207, mapping the physical space with the mapping system of the virtual reality environment, and determining the translation speed and relative position of the user in the virtual reality environment.

[0117] The embodiment sets the maximum value of the actual safety distance to 1, establishes the mapping relationship between the virtual reality environment and the physical space; and based on the normalized spherical harmonic function, corresponds the spatial relationship between the sound environment of the sound source and the played sound environment. According to the mapping relationship, the position and direction in the physical space are mapped into the virtual reality environment.

[0118] S208, adjusting the reverberation parameters according to the playback environment category.

[0119] The embodiment obtains the corresponding reverberation time, reverberation intensity and environment brightness according to the virtual reality environment category; adjusts the reverberation effect of the audio rendering according to the reverberation time and the reverberation intensity in the virtual reality environment category; and adjusts the brightness of the audio rendering according to the environment brightness.

[0120] S209, frequency transforming the audio according to the head translation speed to obtain the frequency transformed audio.

[0121] The embodiment detects the translation speed of the user's head; calculates the Doppler frequency transformation parameter according to the translation speed. Based on the Doppler frequency transformation parameter, the audio signal is processed by Doppler transformation to obtain the Doppler transformed audio signal.

[0122] S210, adjusting the rendering parameters according to the position of the head relative to the sound source and the ear orientation.

[0123] The embodiment calculates the distance of the user's head relative to the sound source; adjusts the distance parameter of the rendering algorithm according to the distance; this is the characteristic of the user's 6DOF audio, that is, the movement of the listener's head and the change of the ear orientation can act on the virtual reality environment to change the characteristics of the sound source: distance, azimuth angle and elevation angle, and loudness. Adjust the azimuth angle and elevation angle of the rendering algorithm according to the ear orientation; adjust the loudness of the audio rendering according to the relative position and orientation.

[0124] S211, rendering and playing the frequency transformed audio according to the updated spherical coordinate basis, rendering parameters and reverberation parameters.

[0125] The embodiment can render the spatial audio data according to the updated spherical coordinate basis and rendering parameters to obtain the rendered spatial audio data; encode the rendering data; transmit the encoded data to the playback device; and perform audio decoding and output according to the playback environment type.

[0126] For ease of understanding, please refer to Embodiment One:

[0127] The application provides a spatial audio HOA playing method based on head real-time pose information, and the specific implementation steps are as follows:

[0128] Step 1, obtaining the user head real-time pose information of the physical space.

[0129] Step 101, obtaining the physical coordinates in real time through the IMU acquisition device of the head-mounted device combined with the environment identifier calculated by the physical space. Specifically, the IMU acquisition device includes an inertial measurement unit (IMU) and an ambient light sensor. By detecting the acceleration and angular velocity data of the user's head and combining the light intensity data of the ambient light sensor, the physical coordinates of the user's head are calculated.

[0130] Step 102, analyzing the pose information of the user's head in real time through the camera (such as TOF camera) in the physical environment. The TOF camera adopts the time difference method principle, calculates the depth information of the user's head by emitting and receiving laser pulse trains, and obtains the three-dimensional model of the user's head.

[0131] Step 103, receiving the historical pose data of the user's head. The historical pose data is issued by the cloud server and contains the head pose information of the user in the last 30 days, which is updated once a day.

[0132] Step 2, normalizing the obtained head real-time pose information.

[0133] Step 201, normalizing the obtained physical coordinates and converting them to a standardized coordinate system. Specifically, the physical coordinates are converted to a world coordinate system, wherein the X-axis points to the left side of the user's body, the Y-axis points to the top of the user's head, and the Z-axis points to the back of the user's body.

[0134] Step 202, filtering the historical pose data and extracting the stable part. A filtering algorithm based on median and standard deviation is used to evaluate each pose data point, and data within the median ± standard deviation range is retained.

[0135] Step 203, fusing the normalized physical coordinates and the filtered historical pose data to obtain comprehensive head real-time pose information. The new and old pose data are fused by weighted average, and the weights are 0.7 and 0.3 respectively.

[0136] Step 3, calculating the offset of the head relative to the HOA sphere center.

[0137] Step 301, calculating the straight-line distance of the user's head relative to the HOA sphere center. The HOA sphere center is defined as the geometric center of the loudspeaker group, and a spherical model is established by using 3D modeling technology, with a radius of 1 meter.

[0138] Step 302, calculate the angle between the head and the center of the HOA sphere. Through trigonometric calculation, the angle between the line connecting the head and the center of the HOA sphere and the horizontal plane is obtained , and the angle with the vertical plane .

[0139] Step 303, combine the distance and the angle to form an offset vector. The offset vector is recorded in the form of .

[0140] Step 4, apply the offset to the rendering algorithm in real time to calculate the spherical coordinates of the speaker group.

[0141] Step 401, convert the offset to an offset angle in the spherical coordinate system. According to the offset vector, the corresponding pitch angle and azimuth angle are calculated.

[0142] Step 402, update the spherical coordinates of the speaker group according to the offset angle. Using the spherical coordinate transformation formula, the new spherical coordinates are calculated.

[0143] Step 403, calculate the updated spherical coordinate base using an interpolation algorithm. Using a linear interpolation algorithm, the change rate of the spherical coordinates is controlled according to the time variation.

[0144] Step 5, select the corresponding reconstruction method of three-dimensional sound field according to the type of playback environment.

[0145] Step 501, determine whether the playback environment is a virtual reality environment. If it is a virtual reality environment, execute step 6; if it is the current physical environment, execute step 8.

[0146] Step 6, normalize the physical space of the reconstructed three-dimensional sound field and the mapping system of the virtual reality environment.

[0147] Step 701, establish the mapping relationship between the virtual reality environment and the physical space. Through the spatial correspondence relationship database, the mapping relationship between the characteristic points in the environment is established.

[0148] Step 702, map the position and direction in the physical space to the virtual reality environment according to the mapping relationship. Using a mapping algorithm based on characteristic points, the characteristic points in the physical space are mapped to the virtual environment.

[0149] Step 703, calculate the translation speed and relative position of the user in the virtual environment. Through optical sensor and IMU data fusion, the six-degree-of-freedom motion parameters of the user are calculated.

[0150] Step 8, adjust the reverberation parameters according to the playback environment category.

[0151] Step 801, according to the virtual reality environment category, obtain the corresponding reverberation time and reverberation intensity. The reverberation parameters of different environment categories are preset, such as conference room, cinema, etc.

[0152] Step 802, adjust the reverberation effect of audio rendering according to the environment type. Convolution algorithm is adopted to process audio signals according to the reverberation characteristics of the environment.

[0153] Step 803, adjust the brightness of audio rendering according to the brightness of the environment. The brightness of the environment is detected by light sensor, and the brightness of the audio is dynamically adjusted.

[0154] Step 9, frequency transformation according to head translation speed.

[0155] Step 901, detect the translation speed of the user's head. Through IMU data processing, the linear motion speed of the head is calculated.

[0156] Step 902, calculate the Doppler frequency transformation parameter according to the translation speed. The Doppler factor is calculated according to the speed and environmental parameters.

[0157] Step 903, Doppler transformation processing of audio signal. Time-frequency transformation method is adopted to realize the Doppler transformation of audio signal.

[0158] Step 10, adjust the rendering parameters according to the position of the head relative to the sound source and the orientation of the ear.

[0159] Step 1001, calculate the distance between the user's head and the sound source. Through three-dimensional reconstruction algorithm, the Euclidean distance between the head and the sound source is calculated.

[0160] Step 1002, adjust the distance parameter of the rendering algorithm according to the distance. Exponential decay function is adopted to change the volume size according to the distance.

[0161] Step 1003, adjust the azimuth and elevation angle of the rendering algorithm according to the orientation of the ear. Through head pose estimation algorithm, the spatial orientation of the ear is calculated.

[0162] Step 1004, adjust the loudness of audio rendering according to the relative position and orientation. The loudness calculation method based on distance and angle is adopted to adjust the loudness of audio.

[0163] Step 11, output the reconstructed three-dimensional sound field data.

[0164] Step 1101, render spatial audio data according to the updated spherical coordinate basis. The sound field rendering algorithm based on spherical coordinates is adopted to calculate three-dimensional audio data.

[0165] Step 1102, encode the rendering data. Efficient audio encoding standard is adopted to compress the audio data.

[0166] Step 1103, transmit the encoded data to the playback device. Real-time transmission of data is achieved through network transmission protocols such as UDP (User Datagram Protocol).

[0167] Step 1104, audio decoding and output according to the playback environment type. According to the type of playback device, such as VR headset or smart speaker, select the corresponding decoding method.

[0168] For convenience, please refer to Example Two:

[0169] The present application provides a kind of spatial audio HOA playing method based on head real-time pose information, specific implementation steps are as follows:

[0170] Step 1, obtain the real-time pose information of user's head in physical space.

[0171] Step 101, through the IMU acquisition device of head-mounted device in combination with the environment mark calculated by physical space, real-time physical coordinates are obtained. Specifically, the IMU acquisition device includes three-axis gyroscope and three-axis accelerometer, by detecting the angular velocity and acceleration of user's head, in combination with the illumination intensity data of ambient light sensor, the physical coordinates of user's head are calculated.

[0172] Step 102, real-time analyze the pose information of user's head through camera (such as TOF camera) in physical environment. TOF camera adopts time-of-flight principle, measures depth information by emitting and receiving infrared light pulse, so as to obtain three-dimensional model of user's head.

[0173] Step 103, receive the historical posture data of user's head. Historical posture data is issued through cloud server, contains user's head posture information in recent 90 days, updated twice a day.

[0174] Step 2, normalize the obtained real-time pose information of head.

[0175] Step 201, normalize the obtained physical coordinates, and convert them to standardized coordinate system. Specifically, convert physical coordinates to earth coordinate system (WGS84), wherein X axis points to left side of user's body, Y axis points to top of user's head, and Z axis points to back of user's body.

[0176] Step 202, filter the historical posture data, and extract the stable part. Adopt filtering algorithm based on median and standard deviation, evaluate each posture data point, and retain data within median ± standard deviation range.

[0177] Step 203, fuse the normalized physical coordinates with the filtered historical pose data to obtain comprehensive head real-time pose information. The new and old pose data are fused by weighted average, and the weights are 0.6 and 0.4 respectively.

[0178] Step 3, calculate the offset of the head relative to the HOA sphere center.

[0179] Step 301, calculate the straight-line distance of the user's head relative to the HOA sphere center. The HOA sphere center is defined as the geometric center of the loudspeaker group, and a spherical model is established by 3D modeling technology, with a radius of 1.2 meters.

[0180] Step 302, calculate the angle of the head deviating from the HOA sphere center. Through trigonometric function calculation, the angle of the head and the HOA sphere center connecting line with the horizontal plane is obtained , and the angle with the vertical plane is .

[0181] Step 303, combine the distance and angle to form the offset vector. The offset vector is recorded in the form of .

[0182] Step 4, apply the offset to the spherical coordinate calculation of the loudspeaker group in the rendering algorithm in real time.

[0183] Step 401, convert the offset to the offset angle in the spherical coordinate system. According to the offset vector, the corresponding pitch angle and azimuth angle are calculated.

[0184] Step 402, update the spherical coordinates of the loudspeaker group according to the offset angle. The new spherical coordinates are calculated by using the spherical coordinate transformation formula.

[0185] Step 403, calculate the updated spherical coordinate base by using the interpolation algorithm. The linear interpolation algorithm is used to control the change speed of the spherical coordinates according to the time change rate.

[0186] Step 5, select the corresponding reconstruction of three-dimensional sound field mode according to the playback environment type.

[0187] Step 501, judge whether the playback environment is a virtual reality environment. If it is a virtual reality environment, step 6 is executed; if it is the current physical environment, step 8 is executed.

[0188] Step 6, judge whether virtual reality environment mapping is needed.

[0189] Step 7, normalize the physical space of the reconstructed three-dimensional sound field and the map system of the virtual reality environment.

[0190] Step 701, establish the mapping relationship between the virtual reality environment and the physical space. Through the spatial correspondence relationship database, the mapping relationship between the environment feature points is established.

[0191] Step 702, map the position and direction in the physical space to the virtual reality environment according to the mapping relationship. The feature point-based mapping algorithm is used to map the feature points in the physical space to the virtual environment.

[0192] Step 703, calculate the translation speed and relative position of the user in the virtual environment. The six-degree-of-freedom motion parameters of the user are calculated through optical sensor and IMU data fusion.

[0193] Step 8, adjust the reverberation parameters according to the playback environment category.

[0194] Step 801, according to the virtual reality environment category, get the corresponding reverberation time and reverberation intensity. The reverberation parameters of different environment categories are preset, such as conference room, cinema, etc.

[0195] Step 802, adjust the reverberation effect of audio rendering according to the environment type. Convolution algorithm is used to process audio signals according to the reverberation characteristics of the environment.

[0196] Step 803, adjust the brightness of audio rendering according to the brightness of the environment. The brightness of the audio is dynamically adjusted by detecting the brightness of the environment through the light sensor.

[0197] Step 9, frequency transformation according to the head translation speed.

[0198] Step 901, detect the translation speed of the user's head. The linear motion speed of the head is calculated through IMU data processing.

[0199] Step 902, calculate the Doppler frequency transformation parameters according to the translation speed. The Doppler factor is calculated according to the speed and environmental parameters.

[0200] Step 903, Doppler transformation processing of audio signal. Time-frequency transformation method is used to realize the Doppler transformation of audio signal.

[0201] Step 10, adjust the rendering parameters according to the position of the head relative to the sound source and the orientation of the ear.

[0202] Step 1001, calculate the distance between the user's head and the sound source. The Euclidean distance between the head and the sound source is calculated through the three-dimensional reconstruction algorithm.

[0203] Step 1002, adjust the distance parameter of the rendering algorithm according to the distance. Exponential decay function is used to change the volume size according to the distance.

[0204] Step 1003, adjust the azimuth and elevation angles of the rendering algorithm according to the orientation of the ear. The spatial orientation of the ear is calculated through the head pose estimation algorithm.

[0205] Step 1004, adjusting the loudness of the audio rendering according to the relative position and orientation. Using a distance and angle based loudness calculation method, the loudness of the audio is adjusted.

[0206] Step 11, outputting the reconstructed three-dimensional sound field data.

[0207] Step 1101, rendering spatial audio data according to the updated spherical coordinate basis. Using a spherical coordinate based sound field rendering algorithm, three-dimensional audio data is calculated.

[0208] Step 1102, encoding the rendering data. Using an efficient audio encoding standard such as AAC or FLAC, the audio data is compressed.

[0209] Step 1103, transmitting the encoded data to the playback device. Through a network transmission protocol such as UDP, real-time transmission of data is achieved.

[0210] Step 1104, audio decoding and output according to the playback environment type. According to the type of playback device, such as VR headset or smart speaker, the corresponding decoding method is selected.

[0211] Compared with the prior art, the present application has the following beneficial effects:

[0212] 1. By normalizing the real-time pose information of the user's head in the physical space and applying the offset of the head relative to the HOA spherical center to the spherical coordinate calculation of the loudspeaker group in the rendering algorithm, the problem of distinguishing between virtual reality environment sound and real environment sound is effectively solved, and the user experience is improved;

[0213] 2. For spatial audio in a virtual reality environment, by normalizing the physical space of the reconstructed three-dimensional sound field with the map system of the virtual reality environment, the two are mapped 1:1, the important parameters such as the translation speed of the user's head in the map system of the virtual reality environment relative to the sound source, the relative position and the ear orientation can be accurately calculated, and the quality of the audio rendering of the virtual reality device is effectively improved;

[0214] 3. According to the category of the virtual reality environment that the user is in, the reverberation time and reverberation intensity specific to this environment are adjusted to enhance the immersive feeling of spatial sound effects, effectively solving the problem of lacking consideration of different user personalized needs in the prior art;

[0215] 4. According to the translation speed of the user's head in the virtual reality environment, the frequency is transformed according to the Doppler effect to enhance the realism of spatial sound effects, overcoming the imperfection of the virtual sound source position adjustment mechanism in the prior art;

[0216] 5. According to the relative position of the user's head to the sound source in the virtual reality environment, and combined with the ear orientation, the distance, loudness, azimuth and elevation angle and other parameters of the rendering algorithm are controlled to enhance the realism of the spatial sound effect, effectively integrating the real-time pose information of the user and the spatial audio rendering system, and improving the auditory immersion and realism of the user.

[0217] The audio playback device provided by the embodiments of the present application is described below. The audio playback device described below can be referred to in correspondence with the audio playback method described above.

[0218] For details, please refer to Figure 3 , Figure 3 The structural schematic diagram of the audio playback device provided by the embodiments of the present application can include:

[0219] The mapping module 100 is configured to map a real physical environment to a virtual reality environment based on a mapping relationship.

[0220] The translation speed and position determination module 200 is configured to determine the translation speed and position of the user in the virtual reality environment.

[0221] The rendering parameter determination module 300 is configured to perform frequency transformation on the audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determine a rendering parameter based on the position. The rendering parameter at least includes a distance parameter, an azimuth, an elevation angle and a loudness.

[0222] The audio playback module 400 is configured to render the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and play the rendered audio signal.

[0223] Further, based on any of the above embodiments, the audio playback device can further include:

[0224] The offset vector determination module is configured to obtain real-time pose information of the user's head in the physical space, and determine an offset vector of the real-time pose information relative to the HOA sphere center.

[0225] The updated spherical coordinate determination module of the loudspeaker group is configured to convert the offset vector into an offset angle in the spherical coordinate system, and update the spherical coordinates of the loudspeaker group based on the offset angle.

[0226] The updated spherical coordinate basis determination module is configured to determine an updated spherical coordinate basis based on the updated spherical coordinates of the loudspeaker group.

[0227] Correspondingly, the audio playback module 400 can include:

[0228] An audio playing unit is configured to render the frequency-transformed audio signal based on the updated spherical coordinate basis and the rendering parameter to obtain a rendered audio signal.

[0229] Further, based on any of the above embodiments, the offset vector determination module can comprise:

[0230] A physical coordinate determination unit of the listener's head is configured to acquire the physical coordinate of the listener's head in real time by means of an inertial measurement unit of a head-mounted device and an environmental identifier calculated in a physical space;

[0231] A real-time pose information determination unit of the user's head is configured to analyze the real-time pose information of the user's head in real time based on the physical coordinate of the listener's head by means of a camera in a physical environment;

[0232] A historical pose data determination unit is configured to determine the historical pose data of the user's head;

[0233] A normalized real-time pose information determination unit is configured to normalize the real-time pose information of the user's head to obtain normalized real-time pose information.

[0234] A filtering unit is configured to filter the historical pose data to obtain filtered historical pose data.

[0235] A fusion unit is configured to fuse the normalized real-time pose information and the filtered historical pose data to obtain comprehensive real-time pose information of the head;

[0236] An offset vector determination unit is configured to determine an offset vector of the comprehensive real-time pose information of the head relative to the HOA sphere center.

[0237] Further, based on any of the above embodiments, the mapping module 100 can comprise:

[0238] A determination unit is configured to determine whether a playback environment is a virtual reality environment;

[0239] A mapping unit is configured to map a real physical environment to a virtual reality environment based on a mapping relationship between the real physical environment and the virtual reality environment determined by using normalized spherical harmonics when the playback environment is a virtual reality environment.

[0240] Further, based on any of the above embodiments, the rendering parameter determination module 300 can comprise:

[0241] A distance determination unit is configured to determine a distance of the user's head relative to a sound source;

[0242] A distance parameter determination unit is configured to adjust the distance parameter of a rendering algorithm according to the distance.

[0243] an azimuth and elevation angle determining unit configured to adjust the azimuth and the elevation angle of a rendering algorithm according to the ear orientation;

[0244] a loudness determining unit configured to adjust the loudness of a rendering parameter according to the position and the ear orientation.

[0245] Further, based on any of the above embodiments, the audio playing device can further include:

[0246] a reverb parameter obtaining module configured to obtain corresponding reverb time, reverb intensity and environment brightness according to the category of the current virtual reality environment;

[0247] Correspondingly, the audio playing module 400 can include:

[0248] a reverb parameter based audio playing signal determining unit configured to adjust the reverb effect of the frequency transformed audio signal according to the reverb time and the reverb intensity, and adjust the brightness of the frequency transformed audio signal according to the environment brightness, to obtain the reverb adjusted audio signal;

[0249] a rendered audio signal determining unit configured to render the reverb adjusted audio signal according to the rendering parameter, to obtain the rendered audio signal.

[0250] Further, based on any of the above embodiments, the rendering parameter determining module 300 can include:

[0251] a translation speed determining unit configured to detect the translation speed of the user's head;

[0252] a transformation parameter determining unit configured to determine the Doppler frequency transformation parameter according to the translation speed.

[0253] a transformation unit configured to frequency transform the audio signal based on the Doppler frequency transformation parameter, to obtain the frequency transformed audio signal.

[0254] It should be noted that the order of the above-mentioned modules and units in the audio playing device can be changed without affecting the logic.

[0255] The audio playing device provided by the embodiment of the present application can comprise: a mapping module 100, configured to map a real physical environment to a virtual reality environment based on a mapping relationship; a translation speed and position determination module 200, configured to determine a translation speed and position of a user in the virtual reality environment; a rendering parameter determination module 300, configured to perform frequency transformation on an audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determine a rendering parameter based on the position; wherein the rendering parameter at least comprises a distance parameter, an azimuth angle, an elevation angle and a loudness; and an audio playing module 400, configured to render the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and play the rendered audio signal. For the spatial audio in the virtual reality environment, the present application firstly maps the real physical space of the reconstructed three-dimensional sound field to the virtual reality environment, determines the translation speed and position of the user in the virtual reality environment, and then reconstructs the three-dimensional sound field: according to the translation speed of the user's head in the virtual reality environment, performs frequency transformation to enhance the reality of the spatial audio effect; according to the position of the user's head relative to the sound source in the virtual reality environment, controls the distance, loudness, azimuth angle and elevation angle and other parameters of the rendering algorithm to enhance the reality of the spatial audio effect.

[0256] The audio playing device provided by the embodiment of the present application will be introduced below, and the audio playing device described below can be correspondingly referred to the audio playing method described above.

[0257] Please refer to Figure 4 , Figure 4 The structural schematic diagram of the audio playing device provided by the embodiment of the present application can comprise:

[0258] The memory 10 is configured to store a computer program;

[0259] The processor 20 is configured to execute the computer program to realize the audio playing method described above.

[0260] The memory 10, the processor 20 and the communication interface 30 can complete the communication among each other through the communication bus 40.

[0261] In the embodiment of the present application, the memory 10 is configured to store one or more programs, and the program can comprise program code, and the program code comprises computer operation instructions. In the embodiment of the present application, the memory 10 can store programs for realizing the following functions:

[0262] Map a real physical environment to a virtual reality environment based on a mapping relationship;

[0263] Determine a translation speed and position of a user in the virtual reality environment;

[0264] The frequency of the audio signal is transformed based on the translation speed to obtain a frequency-transformed audio signal, and a position is determined to obtain a rendering parameter; wherein the rendering parameter at least includes a distance parameter, an azimuth angle, an elevation angle and a loudness;

[0265] The frequency-transformed audio signal is rendered based on the rendering parameter to obtain a rendered audio signal, and the rendered audio signal is played.

[0266] In a possible implementation, the memory 10 can include a program storage area and a data storage area, wherein the program storage area can store an operating system, and application programs required by at least one function, etc.; and the data storage area can store data created in a use process.

[0267] In addition, the memory 10 can include a read-only memory and a random access memory, and provide instructions and data for the processor. A part of the memory can also include an NVRAM. The memory stores an operating system and operation instructions, executable modules or data structures, or a subset thereof, or an extended set thereof, wherein the operation instructions can include various operation instructions for implementing various operations. The operating system can include various system programs for implementing various basic tasks and processing hardware-based tasks.

[0268] The processor 20 can be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field programmable gate array or other programmable logic device, and the processor 20 can be a microprocessor or any conventional processor, etc. The processor 20 can invoke a program stored in the memory 10.

[0269] The communication interface 30 can be an interface of a communication module, and is used to connect with other devices or systems.

[0270] Of course, it should be noted that, Figure 4 The structures shown do not constitute a limitation on the audio playing device in the embodiments of the present application, and in actual applications, the audio playing device can include more or fewer components than Figure 4 those shown, or combine certain components.

[0271] The computer readable storage medium provided by the embodiments of the present application is introduced below, and the computer readable storage medium described below can be mutually referred to the audio playing method described above.

[0272] The present application also provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the audio playing method described above.

[0273] The computer readable storage medium can include a U disk, a mobile hard disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, and various media capable of storing program codes.

[0274] The various embodiments are described in a progressive manner in the specification, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the relevant part can be referred to the method part.

[0275] The skilled person can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0276] Finally, it should be noted that, in this document, relationships such as first and second are intended to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variant are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device.

[0277] The above describes in detail the audio playing method, device, equipment and computer readable storage medium provided by the present application. The principles and implementation modes of the present application are described by applying specific examples in this document. The above description of the embodiments is only to help understand the method and core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed; in view of the above, the content of the specification should not be understood as a limitation of the present application.

Claims

1. An audio playback method, characterized by, The method comprises: mapping a real physical environment to a virtual reality environment based on a mapping relationship; determining a translation speed and a position of a user in the virtual reality environment; frequency transforming an audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determining a rendering parameter based on the position; wherein the rendering parameter at least includes a distance parameter, an azimuth angle, an elevation angle, and a loudness; rendering the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal.

2. The audio playback method of claim 1, wherein, Before mapping the real physical environment to the virtual reality environment based on the mapping relationship, the method further comprises: obtaining real-time pose information of a user's head in a physical space, and determining an offset vector of the real-time pose information relative to a HOA sphere center; converting the offset vector into an offset angle in a spherical coordinate system, and updating a spherical coordinate of a loudspeaker group based on the offset angle; determining an updated spherical coordinate basis based on the updated spherical coordinate of the loudspeaker group; correspondingly, rendering the frequency-transformed audio signal based on the rendering parameter to obtain the rendered audio signal, and playing the rendered audio signal, comprises: rendering the frequency-transformed audio signal based on the updated spherical coordinate basis and the rendering parameter to obtain the rendered audio signal.

3. The audio playback method of claim 2, wherein, Obtaining real-time pose information of a user's head in a physical space, and determining an offset vector of the real-time pose information relative to a HOA sphere center, comprises: real-time obtaining physical coordinates of a listener's head through an inertial measurement unit of a head-mounted device combined with an environment identifier calculated in a physical space; real-time analyzing real-time pose information of a user's head based on the physical coordinates of the listener's head through a camera in a physical environment; determining historical pose data of the user's head; normalizing the real-time pose information of the user's head to obtain normalized real-time pose information; filtering the historical pose data to obtain filtered historical pose data; fusing the normalized real-time pose information and the filtered historical pose data to obtain comprehensive real-time pose information of the head; determining the offset vector of the comprehensive real-time pose information of the head relative to the HOA sphere center.

4. The audio playing method of any one of claims 1 to 3, characterized in that, Mapping a real physical environment to a virtual reality environment based on a mapping relationship comprises: determining whether a playback environment is a virtual reality environment; when the playback environment is a virtual reality environment, mapping the real physical environment to the virtual reality environment based on a mapping relationship between the real physical environment and the virtual reality environment determined by using normalized spherical harmonics.

5. The audio playback method of claim 1, wherein, Frequency transforming an audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determining a rendering parameter based on the position, comprises: determining a distance of a user's head relative to a sound source; adjusting the distance parameter of a rendering algorithm according to the distance; adjusting the azimuth angle and the elevation angle of the rendering algorithm according to an ear orientation; adjusting the loudness of the rendering parameter according to the position and the ear orientation.

6. The audio playback method of claim 1, wherein, Before rendering the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal, the method further comprises: According to the category of the current virtual reality environment, obtaining the corresponding reverberation time, reverberation intensity and ambient brightness; Accordingly, rendering the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal, comprises: Adjusting the reverberation effect of the frequency-transformed audio signal according to the reverberation time and the reverberation intensity, and adjusting the brightness of the frequency-transformed audio signal according to the ambient brightness to obtain a reverberation-adjusted audio signal; Rendering the reverberation-adjusted audio signal based on the rendering parameter to obtain the rendered audio signal.

7. The audio playback method of claim 1, wherein, Frequency-transforming the audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determining the rendering parameter based on the position, comprises: Detecting the translation speed of the user's head; Determining the Doppler frequency transformation parameter according to the translation speed; Frequency-transforming the audio signal based on the Doppler frequency transformation parameter to obtain the frequency-transformed audio signal.

8. An audio playback device, characterized by Comprise: A mapping module for mapping a real physical environment to a virtual reality environment based on a mapping relationship; A translation speed and position determination module for determining the translation speed and position of the user in the virtual reality environment; A rendering parameter determination module for frequency-transforming the audio signal based on the translation speed to obtain a frequency-transformed audio signal, and determining the rendering parameter based on the position; wherein the rendering parameter at least includes distance parameter, azimuth angle, elevation angle and loudness; An audio playing module for rendering the frequency-transformed audio signal based on the rendering parameter to obtain a rendered audio signal, and playing the rendered audio signal.

9. An audio playback device, comprising: Comprise: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the audio playing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the audio playing method according to any one of claims 1 to 7.