Audio playing method and device, electronic equipment and computer readable storage medium
By collecting and predicting the user's head pose using a head-mounted display device, a user coordinate system basis vector is constructed to realize sound source location mapping and speaker selection, solving the problem of insufficient sound source localization in VR devices and improving the user's auditory experience.
Patent Information
- Application Number
- CN202511317129.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Existing VR devices have weak sound source localization capabilities during audio rendering and lack real-time sound direction redirection, which affects the user's immersive experience.
By collecting the user's head pose data through a head-mounted display device, extrapolating and predicting, constructing the basis vector of the user coordinate system, mapping the sound source position, and using the target speaker to play audio, the sound source can be accurately located.
It improves the accuracy of sound source localization during the audio rendering process, enhancing the user's auditory experience.
Smart Images

Figure CN120832118B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtual reality technology, and in particular to an audio playback method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] With the increasing popularity of virtual reality (VR) and augmented reality (AR) devices, users are demanding higher levels of spatial awareness and realism in audio. However, most current VR video content only supports stereo or simple surround sound, which has weak sound source localization capabilities and lacks real-time sound direction redirection when the user moves or turns their head, severely impacting the user's immersive experience.
[0003] Therefore, how to achieve more accurate and effective sound source localization during audio rendering and playback, and further improve the user's auditory experience, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] The purpose of this application is to provide an audio playback method that can achieve more accurate and effective sound source localization during audio rendering and playback, thereby further enhancing the user's auditory experience. Another purpose of this application is to provide an audio playback device, an electronic device, and a computer-readable storage medium, all of which have the aforementioned beneficial effects.
[0005] In a first aspect, this application discloses an audio playback method, including:
[0006] The head pose data of the target user is collected using a head-mounted display device, and extrapolation prediction is performed based on the head pose data to obtain the predicted head pose.
[0007] Construct the basis vectors of the user coordinate system based on the predicted head pose;
[0008] The relative position between the sound source and the target user is obtained in the world coordinate system, and the relative position is mapped to the user coordinate system using the basis vector to obtain the sound source position;
[0009] Determine the virtual listening orientation of the target user based on the location of the sound source;
[0010] The target speaker is determined based on the virtual listening orientation, and audio is played using the target speaker.
[0011] Optionally, the head-mounted display device is used to collect head pose data of the target user, and extrapolation prediction is performed based on the head pose data to obtain the predicted head pose, including:
[0012] The head-mounted display device is used to collect the head position and head orientation of the target user;
[0013] Based on the head position, extrapolation is performed to predict the head position;
[0014] Based on the head orientation, extrapolation is performed to predict the head orientation;
[0015] The head prediction pose is generated based on the head prediction position and the head prediction orientation.
[0016] Optionally, extrapolation prediction is performed based on the head position to obtain the predicted head position, including:
[0017] The head-mounted display device is used to obtain the head movement speed of the target user;
[0018] The head position and the head movement speed are smoothed to obtain the target head position and the target head movement speed.
[0019] The audio-visual time difference is obtained as the extrapolation duration, and the target head position is extrapolated and predicted using the target head movement speed and the extrapolation duration to obtain the predicted head position.
[0020] Optionally, the head orientation includes the head orientation of a preset number of historical audio frames; extrapolation prediction is performed based on the head orientation to obtain the predicted head orientation, including:
[0021] The head orientation of the preset number of historical audio frames is smoothed to obtain the average quaternion of each group of adjacent historical audio frames.
[0022] Spherical linear interpolation is performed on each of the average quaternions using a preset interpolation factor to obtain the target quaternion of the future audio frame corresponding to the preset interpolation factor, and the target quaternion is converted into the head prediction orientation.
[0023] Optionally, determining a target speaker based on the virtual listening orientation and using the target speaker for audio playback includes:
[0024] Obtain all speaker combinations from all speakers; wherein each speaker combination comprises three speakers;
[0025] The virtual sound source synthesis rules are determined based on the virtual listening orientation, and the target speaker combination is obtained by filtering from all the speaker combinations according to the virtual sound source synthesis rules.
[0026] Each of the speakers in the target speaker assembly is used as the target speaker, and audio is played using the target speakers.
[0027] Optionally, after obtaining all speaker combinations from all speakers, the following is also included:
[0028] If the three speakers in the speaker assembly cannot be combined into a triangular area based on their positions, the speaker assembly will be removed.
[0029] Optionally, the audio playback method further includes:
[0030] The head pose data and audio data are bound to a master clock, and the audio and video time difference is used to synchronize the audio and video data, so as to achieve timestamp alignment of the head pose data, the audio data, and the video data.
[0031] Secondly, this application discloses an audio playback device, comprising:
[0032] The prediction module is used to collect head pose data of the target user using the head display device, and perform extrapolation prediction based on the head pose data to obtain the predicted head pose.
[0033] A construction module is used to construct the basis vectors of the user coordinate system based on the predicted head pose;
[0034] The mapping module is used to obtain the relative position between the sound source and the target user in the world coordinate system, and to map the relative position to the user coordinate system using the basis vector to obtain the sound source position;
[0035] The determining module is used to determine the virtual listening orientation of the target user based on the location of the sound source;
[0036] The playback module is used to determine the target speaker based on the virtual listening direction and to play audio using the target speaker.
[0037] Thirdly, this application discloses an electronic device, including:
[0038] Memory, used to store computer programs;
[0039] A processor, configured to implement any of the audio playback methods described above when executing the computer program.
[0040] Fourthly, this application discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the audio playback methods described above.
[0041] This application provides an audio playback method, comprising: acquiring head pose data of a target user using a head-mounted display device, and performing extrapolation prediction based on the head pose data to obtain a predicted head pose; constructing a basis vector of a user coordinate system based on the predicted head pose; obtaining the relative position between a sound source and the target user in a world coordinate system, and mapping the relative position to the user coordinate system using the basis vector to obtain the sound source position; determining the virtual listening orientation of the target user based on the sound source position; determining a target speaker based on the virtual listening orientation, and playing audio using the target speaker.
[0042] By applying the technical solution provided in this application, the head-mounted display device can first collect the head pose data of the target user in real time. Then, based on this head pose data, extrapolation prediction can be performed to obtain the head pose at a future time, i.e., the predicted head pose. Next, based on this predicted head pose, the sound source position can be transformed from the world coordinate system to the user coordinate system, obtaining the accurate sound source position relative to the target user, thereby determining the target user's virtual listening orientation. Thus, based on this virtual listening orientation, a target speaker can be selected from multiple speakers. Clearly, this target speaker is the one that can provide the best auditory effect for the target user among all speakers. Therefore, this technical solution can achieve more accurate and effective sound source localization during audio rendering and playback, further enhancing the user's auditory experience.
[0043] The audio playback device, electronic device, and computer-readable storage medium provided in this application also have the above-mentioned technical effects, and will not be described in detail here. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the prior art and the embodiments of this application, the accompanying drawings used in the description of the prior art and the embodiments of this application will be briefly introduced below. Of course, the accompanying drawings described below with respect to the embodiments of this application are only a part of the embodiments in this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort, and such other drawings also fall within the protection scope of this application.
[0045] Figure 1 This is a flowchart illustrating an audio playback method provided in an embodiment of this application.
[0046] Figure 2 A schematic diagram illustrating the principle of head orientation extrapolation prediction provided in this application embodiment;
[0047] Figure 3 A schematic diagram illustrating a time alignment mechanism provided in an embodiment of this application;
[0048] Figure 4 This is a schematic diagram illustrating the implementation process of an audio playback method provided in an embodiment of this application.
[0049] Figure 5 This is a schematic diagram of the structure of an audio playback device provided in an embodiment of this application;
[0050] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0051] The core of this application is to provide an audio playback method that can achieve more accurate and effective sound source localization during audio rendering and playback, thereby further enhancing the user's auditory experience. Another core aspect of this application is to provide an audio playback device, electronic device, and computer-readable storage medium, all of which have the aforementioned beneficial effects.
[0052] To provide a clearer and more complete description of the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0053] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an audio playback method provided in an embodiment of this application. The audio playback method may include, but is not limited to, the following S101 to S105.
[0054] S101: Collect head pose data of the target user using a head-mounted display device, and extrapolate and predict the head pose based on the head pose data to obtain the predicted head pose.
[0055] This step aims to acquire and predict the head pose data of the target user. Specifically, the head-mounted display (HUD) can be equipped with a pose sensor and related API (Application Programming Interface). Once the target user wears and activates the HUD, the pose sensor can acquire the user's head pose data in real time and send it to the main controller via the API. The main controller then extrapolates and predicts the head pose based on this data. In essence, the head pose data represents the target user's actual head pose at the current and historical moments, while the predicted head pose represents the predicted head pose at a future moment. This enables timely and effective head pose perception and prediction, facilitating more accurate sound source redirection in subsequent processes and ensuring higher-quality audio playback.
[0056] In one embodiment of this application, the head-mounted display device is used to collect head pose data of a target user, and extrapolation prediction is performed based on the head pose data to obtain a predicted head pose. This may include: collecting the head position and head orientation of the target user using the head-mounted display device; performing extrapolation prediction based on the head position to obtain a predicted head position; performing extrapolation prediction based on the head orientation to obtain a predicted head orientation; and generating a predicted head pose based on the predicted head position and the predicted head orientation.
[0057] Specifically, head pose data mainly includes the target user's head position and head orientation. Head position refers to the location of the target user's head center point, which can be the coordinates of the head center point in the user's coordinate system. Head orientation mainly includes the target user's head facing forward, backward, left, right, up (looking up), and down (looking down), corresponding to the gaze direction. Correspondingly, extrapolating the head position yields the predicted head position; extrapolating the head orientation yields the predicted head orientation, and combining the two yields the predicted head pose.
[0058] The process of extrapolating and predicting the head position based on the head position can include: using a head-mounted display device to obtain the target user's head movement speed; smoothing the head position and head movement speed to obtain the target head position and target head movement speed; obtaining the audio-visual time difference as the extrapolation duration; and using the target head movement speed and extrapolation duration to extrapolate and predict the target head position to obtain the predicted head position.
[0059] Specifically, the head-mounted display device can also be equipped with a speed sensor to sense the head movement speed of the target user; at the same time, it can obtain the time difference between audio data and video data as the extrapolation duration, that is, the future prediction duration (for example, if the audio is 40 milliseconds slower than the video, the extrapolation duration is 40 milliseconds); thus, the head movement speed and extrapolation duration can be used to extrapolate and predict the head position of the target user, and obtain the predicted head position.
[0060] First, a first-order exponential smoothing method is used to smooth the head position and head movement speed. The formula for first-order exponential smoothing is as follows:
[0061] ;
[0062] in, This represents the input at time t. This represents the output after smoothing the input at time t. This represents the output after smoothing the input at time t-1. This is a smoothing coefficient; the smaller the value, the smoother the output, but the slower the response.
[0063] Therefore, by substituting the head position into the exponential smoothing formula in the previous section, we can obtain the smoothed head position, which is the target head position; by substituting the head movement speed into the exponential smoothing formula in the previous section, we can obtain the smoothed head movement speed, which is the target head movement speed.
[0064] Furthermore, a position extrapolation prediction formula is used to extrapolate and predict the head position:
[0065] ;
[0066] in, Indicates the predicted head position. This indicates the target head position obtained after smoothing. This represents the target head movement speed obtained after smoothing. Indicates the extrapolation duration.
[0067] The head orientation includes the head orientation of a preset number of historical audio frames; extrapolation prediction is performed based on the head orientation to obtain the predicted head orientation, including: smoothing the head orientation of the preset number of historical audio frames to obtain the average quaternion of each group of adjacent historical audio frames; using a preset interpolation factor to perform spherical linear interpolation on each average quaternion to obtain the target quaternion of the future audio frame corresponding to the preset interpolation factor, and converting the target quaternion into the predicted head orientation.
[0068] Specifically, the prediction of head orientation can also be achieved by first smoothing and then extrapolating. The smoothing can be achieved by quaternion weighted average calculation, and the extrapolation can be achieved by spherical linear interpolation.
[0069] First, the head orientation (forward / up, where forward represents the head's orientation towards the front, back, left, or right, and up represents the head's orientation towards the top (head up) or bottom (head down)) includes the head orientation of a preset number of historical audio frames. Each head orientation is then converted into a quaternion Q1, Q2, ..., Q... N N represents the preset quantity. Given N quaternions Q1, Q2, ..., Q... N and the corresponding weights w1, w2, ..., w N The weighted calculation formula is as follows:
[0070] ;
[0071] in, It is the outer matrix of quaternions, with weights w i Control the impact of each quaternion on the final result.
[0072] Normalization is performed on matrix M to ensure that the matrix... The numerical stability, calculated using the normalized formula, is as follows:
[0073] ;
[0074] Through calculation The principal eigenvector, i.e., the average quaternion q, can be obtained by finding the eigenvector corresponding to the largest eigenvalue.
[0075] Assuming N is 3, we obtain the head orientation quaternion Q1 for the penultimate frame (current frame), Q2 for the penultimate frame, and Q3 for the penultimate frame. Then, by using the weighted average calculation formula and normalization calculation formula to calculate each adjacent quaternion, we can obtain the average quaternion q1 of Q2 and its corresponding weight w2 for the penultimate frame, and the average quaternion q2 of Q3 and its corresponding weight w3 for the penultimate frame, as well as the average quaternion q2 of Q1 and its corresponding weight w1 for the penultimate frame, and the average quaternion q2 of Q2 and its corresponding weight w2 for the penultimate frame.
[0076] Furthermore, by using spherical linear interpolation (SLERP) to extrapolate and estimate the head orientation, we can obtain the future head orientation information, i.e., the predicted head orientation. For details, please refer to... Figure 2 , Figure 2The head orientation extrapolation prediction principle diagram provided in this application embodiment interpolates the average quaternions q1 and q2 of the two most recent frames along the shortest path between two points on a unit four-dimensional sphere. The spherical linear interpolation formula is as follows:
[0077] ;
[0078] Where, q t Let q1 be the target quaternion, q2 be the initial average quaternion, q2 be the final average quaternion, and t be the interpolation factor. This represents the angle between two average quaternions (unit quaternion dot product); when t=0, we get q1; when t=1, we get q2; when t>1, it represents extrapolation to the direction of q2. For example, t=1.5 represents extrapolation to the head orientation corresponding to half a frame delay in the current frame, and t=2 represents extrapolation to the head orientation corresponding to the next frame delay. Finally, the target quaternion q... t Convert to forward / up to obtain the predicted head orientation.
[0079] Therefore, in this embodiment of the application, by smoothing the head position and head orientation respectively, and then using the change trend of the most recent frame for short-term linear extrapolation, the state after a certain delay time (such as 30ms) can be predicted, so as to obtain the future listening direction and complete the spatial audio rendering.
[0080] S102: Construct the basis vectors of the user coordinate system based on the predicted head pose.
[0081] This step aims to construct the user coordinate system basis vectors based on the predicted head pose, facilitating the transformation of the sound source location from the world coordinate system to the user coordinate system, and obtaining the accurate sound source location relative to the target user. Specifically, the construction method is as follows:
[0082] right = cross(up, forward); / / User's right-side direction
[0083] forward' = normalize(forward); / / Ensure normalization again
[0084] up' = normalize(up);
[0085] right' = normalize(right);
[0086] The above vectors can be used to form a local rotation matrix R=[right',up',forward'], which yields the basis vectors of the user coordinate system. Here, forward represents the head's orientation (front, back, left, right), and up represents the head's orientation (up or down).
[0087] S103: Obtain the relative position of the sound source and the target user in the world coordinate system, and use the basis vectors to map the relative position to the user coordinate system to obtain the sound source position.
[0088] This step aims to transform the sound source location from the world coordinate system to the user coordinate system, in order to obtain the accurate sound source location (in the user coordinate system) relative to the target user. First, obtain the relative position vector between the sound source and the target user in the world coordinate system: LtoS = S - L, where S represents the sound source and L represents the target user; further, project the relative position vector LtoS onto the user coordinate system:
[0089] x_rel = dot(LtoS, right'); / / Horizontal direction (right side is positive)
[0090] y_rel = dot(LtoS, up'); / / Vertical direction (up is positive)
[0091] z_rel = dot(LtoS, forward'); / / Forward distance (positive for forward)
[0092] relativePos = (x_rel, y_rel, z_rel);
[0093] Thus, the sound source position (x_rel, y_rel, z_rel) in the user coordinate system can be obtained.
[0094] S104: Determine the virtual listening orientation of the target user based on the location of the sound source.
[0095] This step aims to determine the virtual listening orientation of the target user. In the implementation process, the virtual listening orientation can be obtained by directly normalizing the sound source position: virtualDir = normalize(relativePos).
[0096] S105: Determine the target speaker based on the virtual listening direction, and use the target speaker to play audio.
[0097] This step aims to identify the target speaker and play audio based on it. It is understood that the target speaker is the one among all speakers that can provide the best listening experience for the target user; there may be multiple such speakers.
[0098] In one embodiment of this application, determining a target speaker based on a virtual listening orientation and using the target speaker for audio playback may include: obtaining all speaker combinations among all speakers; wherein each speaker combination includes three speakers; determining a virtual sound source synthesis rule (i.e., the virtual sound source synthesis formula below) based on the virtual listening orientation, and selecting a target speaker combination from all speaker combinations based on the virtual sound source synthesis rule; using each speaker in the target speaker combination as the target speaker, and using the target speaker for audio playback.
[0099] After obtaining all speaker combinations from all speakers, the process may further include: removing a speaker combination if the three speakers in a combination cannot be combined into a triangular region based on their positions.
[0100] Specifically, in the implementation process, all loudspeakers in the current spatial environment can be grouped into triangular regions of three loudspeakers each, with the triangular regions not overlapping. For any given loudspeaker combination, the unit direction vector of each loudspeaker relative to the target user can be named... , , Direct the virtual listening direction toward virtualDir using vectors Therefore, the formula for synthesizing a virtual sound source using this speaker combination is:
[0101] ;
[0102] Based on the above virtual sound source synthesis formula, all speaker combinations are iterated to find those that satisfy the requirements. and The target speaker combination, in which three speakers are used to render the virtual sound source, wherein the coefficients... , , These represent the weights of sound energy allocated to each target speaker. Finally, the rendered audio is captured and played through a multi-channel audio capture card.
[0103] In addition, the audio playback method may also include: binding head pose data and audio data to a master clock, and using the audio-video time difference to synchronize the audio data and video data, so as to achieve timestamp alignment of head pose data, audio data, and video data.
[0104] Understandably, in real-world applications, due to frequent and rapid head movements by users, and the fact that VR systems contain multiple asynchronous modules (audio, video, pose), a lack of a unified low-latency synchronization mechanism can lead to problems such as audio-visual asynchrony (sound and image misalignment, breaking immersion), sound misalignment (audio response is slow, creating "auditory delay"), and directional distortion (e.g., the head has turned but the sound still comes from the "stationary" location). Therefore, this application proposes an audio rendering mechanism that rapidly senses head movements and drives a multi-channel physical speaker system.
[0105] For details, please refer to Figure 3 , Figure 3 The schematic diagram illustrates a time alignment mechanism provided in this application. Since the head-mounted display (HMD) pose data clock originates from within the HMD, the audio playback clock from within the sound card, and the video rendering clock from within the GPU, and these clocks are different, a unified system clock (master clock) is needed to ensure the consistency of "pose data, audio rendering, and video synchronization" when performing high-precision synchronization tasks such as spatial audio rendering on a Windows system. Therefore, the Windows high-precision system time can be sampled, with an accuracy down to the microsecond level. The pseudocode is as follows:
[0106] LARGE_INTEGER frequency;
[0107] QueryPerformanceFrequency(&frequency); / / Retrieves all frequencies at once
[0108] LARGE_INTEGER counter;
[0109] QueryPerformanceCounter(&counter);
[0110] double time_seconds = counter.QuadPart / (double)frequency.QuadPart;
[0111] The above code is encapsulated into the GetMasterClockTime() function, which serves as the system's global clock. Based on this:
[0112] (1) The head display is bound to the master clock. When the head pose data is received, the function GetMasterClockTime() is called immediately and the frame is bound to the master clock.
[0113] (2) Audio is bound to the master clock. The GetMasterClockTime() function is called when each frame of audio data is written to the audio buffer, and the frame is bound to the master clock.
[0114] (3) Video data synchronization: Based on the timestamps carried by the audio and video, the video frames are synchronized with the audio frames.
[0115] This effectively ensures the consistency of timestamps for "pose data, audio rendering, and video synchronization".
[0116] It should be noted that the time alignment mechanism can be activated after the head pose data has been acquired, for example, please refer to... Figure 4 , Figure 4 The schematic diagram of the implementation process of an audio playback method provided in this application mainly includes the operation process of head pose acquisition, timestamp alignment, delay compensation and prediction (i.e. extrapolation prediction), pose-sound source mapping, audio rendering, and multi-channel playback.
[0117] As can be seen, the audio playback method provided in this application first utilizes a head-mounted display device to collect the head pose data of the target user in real time. Then, based on this head pose data, extrapolation prediction is performed to obtain the head pose at a future time, i.e., the predicted head pose. Next, based on this predicted head pose, the sound source position can be transformed from the world coordinate system to the user coordinate system, obtaining the accurate sound source position relative to the target user, thereby determining the target user's virtual listening orientation. Thus, based on this virtual listening orientation, a target speaker can be selected from multiple speakers. Clearly, this target speaker is the one that can provide the best auditory effect to the target user among all speakers. Therefore, this technical solution can achieve more accurate and effective sound source localization during audio rendering and playback, further enhancing the user's auditory experience.
[0118] This application provides an audio playback device.
[0119] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of an audio playback device provided in an embodiment of this application. The audio playback device may include:
[0120] Prediction module 1 is used to collect head pose data of the target user using the head display device, and extrapolate and predict the head pose based on the head pose data.
[0121] Module 2 is used to construct the basis vectors of the user coordinate system based on the predicted head pose;
[0122] Mapping module 3 is used to obtain the relative position between the sound source and the target user in the world coordinate system, and to map the relative position to the user coordinate system using basis vectors to obtain the sound source position;
[0123] Module 4 is used to determine the virtual listening orientation of the target user based on the location of the sound source;
[0124] Playback module 5 is used to determine the target speaker based on the virtual listening direction and to play audio using the target speaker.
[0125] As can be seen, the audio playback device provided in this application firstly utilizes a head-mounted display device to collect the head pose data of the target user in real time, and then extrapolates and predicts the head pose at a future time based on this head pose data, i.e., the predicted head pose. Then, based on this predicted head pose, the sound source position can be transformed from the world coordinate system to the user coordinate system, obtaining the accurate sound source position relative to the target user, thereby determining the virtual listening orientation of the target user. Thus, based on this virtual listening orientation, a target speaker can be selected from multiple speakers. Clearly, this target speaker is the one that can provide the best auditory effect for the target user among all speakers. Therefore, this technical solution can achieve more accurate and effective sound source localization during audio rendering and playback, further enhancing the user's auditory experience.
[0126] In one embodiment of this application, the prediction module 1 may include:
[0127] The acquisition unit is used to acquire the head position and head orientation of the target user using the head-mounted display device;
[0128] The first extrapolation unit is used to extrapolate and predict the head position based on the head position;
[0129] The second extrapolation unit is used to extrapolate and predict the head orientation based on the head orientation.
[0130] The generation unit is used to generate the head predicted pose based on the head predicted position and the head predicted orientation.
[0131] In one embodiment of this application, the first extrapolation unit can be specifically used to obtain the head movement speed of the target user using a head-mounted display device; to smooth the head position and head movement speed to obtain the target head position and target head movement speed; to obtain the audio-visual time difference as the extrapolation duration; and to extrapolate and predict the target head position using the target head movement speed and the extrapolation duration to obtain the predicted head position.
[0132] In one embodiment of this application, the second extrapolation unit can be specifically used to smooth the head orientation of a preset number of historical audio frames to obtain the average quaternion of each group of adjacent historical audio frames; to perform spherical linear interpolation on each average quaternion using a preset interpolation factor to obtain the target quaternion of the future audio frame corresponding to the preset interpolation factor, and to convert the target quaternion into the predicted head orientation.
[0133] In one embodiment of this application, the playback module 5 can be specifically used to obtain all speaker combinations among all speakers; wherein, each speaker combination includes three speakers; determine the virtual sound source synthesis rule according to the virtual listening direction, and filter out the target speaker combination from all speaker combinations according to the virtual sound source synthesis rule; use each speaker in the target speaker combination as the target speaker, and use the target speaker for audio playback.
[0134] In one embodiment of this application, the playback module 5 can also be used to remove a speaker combination when, after obtaining all speaker combinations among all speakers, the three speakers in the speaker combination cannot be combined into a triangular area according to the speaker positions.
[0135] In one embodiment of this application, the audio playback device may further include an alignment module for binding head pose data and audio data to a master clock, and using the audio-video time difference to synchronize the audio data and video data, so as to achieve timestamp alignment of head pose data, audio data, and video data.
[0136] For a description of the apparatus provided in the embodiments of this application, please refer to the above method embodiments; further details will not be repeated here.
[0137] This application also provides an electronic device, please refer to... Figure 6 , Figure 6 This application provides a schematic diagram of the structure of an electronic device, which may include:
[0138] Memory 11 is used to store computer programs;
[0139] The processor 10 is configured to execute computer programs to implement the steps of any of the audio playback methods described above.
[0140] like Figure 6 The diagram shows the structural composition of an electronic device, which may include a processor 10, a memory 11, a communication interface 12, and a communication bus 13. The processor 10, memory 11, and communication interface 12 all communicate with each other through the communication bus 13.
[0141] In this embodiment, the processor 10 may be a central processing unit (CPU), an application-specific integrated circuit, a digital signal processor, a field-programmable gate array, or other programmable logic devices.
[0142] The processor 10 can call the program stored in the memory 11. Specifically, the processor 10 can execute the operations in the embodiment of the audio playback method.
[0143] The memory 11 is used to store one or more programs. The programs may include program code, which includes computer operation instructions. In this embodiment, the memory 11 stores at least a program for implementing the following functions:
[0144] The system uses a head-mounted display to collect head pose data from the target user and extrapolates and predicts the head pose based on this data. It then constructs the basis vectors of the user coordinate system based on the predicted head pose. The system obtains the relative position between the sound source and the target user in the world coordinate system and maps this relative position to the user coordinate system using the basis vectors to obtain the sound source position. The system determines the virtual listening orientation of the target user based on the sound source position. Finally, it determines the target speaker based on the virtual listening orientation and uses the target speaker to play audio.
[0145] In one possible implementation, the memory 11 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; and the data storage area may store data created during use.
[0146] In addition, memory 11 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device.
[0147] Communication interface 12 can be an interface for the communication module, used to connect with other devices or systems.
[0148] Of course, it should be noted that, Figure 6 The structure shown does not constitute a limitation on the electronic device in the embodiments of this application. In practical applications, the electronic device may include more than Figure 6 More or fewer components as shown, or combinations of certain components.
[0149] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of any of the audio playback methods described above.
[0150] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0151] For a description of the computer-readable storage medium provided in this application, please refer to the above method embodiments; further details will not be repeated here.
[0152] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0153] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0154] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0155] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of this application. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. An audio playback method, characterized by, The method comprises the following steps: Collecting head pose data of a target user by using a head-mounted device, and performing extrapolation prediction according to the head pose data to obtain a head prediction pose; Constructing a base vector of a user coordinate system according to the head prediction pose; Obtaining the relative position of a sound source and the target user in a world coordinate system, and mapping the relative position to the user coordinate system by using the base vector to obtain a sound source position; Determining a virtual listening orientation of the target user according to the sound source position; Determining a target loudspeaker according to the virtual listening orientation, and playing audio by using the target loudspeaker; The method of collecting head pose data of a target user by using a head-mounted device, and performing extrapolation prediction according to the head pose data to obtain a head prediction pose comprises the following steps: collecting the head position and head orientation of the target user by using the head-mounted device; performing extrapolation prediction according to the head position to obtain a head prediction position; performing extrapolation prediction according to the head orientation to obtain a head prediction orientation; and generating the head prediction pose according to the head prediction position and the head prediction orientation. The method of performing extrapolation prediction according to the head position to obtain a head prediction position comprises the following steps: obtaining the head movement speed of the target user by using the head-mounted device; performing smoothing processing on the head position and the head movement speed to obtain a target head position and a target head movement speed; obtaining an audio-video time difference as an extrapolation time length, and performing extrapolation prediction on the target head position by using the target head movement speed and the extrapolation time length to obtain the head prediction position.
2. The audio playback method of claim 1, wherein, The head orientation comprises the head orientations of a preset number of historical audio frames. The method of performing extrapolation prediction according to the head orientation to obtain a head prediction orientation comprises the following steps: Performing smoothing processing on the head orientations of the preset number of historical audio frames to obtain the average quaternion of each group of adjacent historical audio frames; Performing spherical linear interpolation processing on each average quaternion by using a preset interpolation factor to obtain a target quaternion of a future audio frame corresponding to the preset interpolation factor, and converting the target quaternion into the head prediction orientation.
3. The audio playback method of claim 1, wherein, The method of determining a target loudspeaker according to the virtual listening orientation, and playing audio by using the target loudspeaker comprises the following steps: Obtaining all loudspeaker combinations in all loudspeakers; wherein each loudspeaker combination comprises three loudspeakers; Determining a virtual sound source synthesis rule according to the virtual listening orientation, and screening a target loudspeaker combination from all the loudspeaker combinations according to the virtual sound source synthesis rule; Taking each loudspeaker in the target loudspeaker combination as the target loudspeaker, and playing audio by using the target loudspeaker.
4. The audio playback method of claim 3, wherein, After obtaining all loudspeaker combinations in all loudspeakers, the method further comprises the following steps: When the three loudspeakers in the loudspeaker combination cannot form a triangular area according to the loudspeaker position combination, the loudspeaker combination is excluded.
5. The audio playing method of any one of claims 1 to 4, characterized in that, The method further comprises the following steps: The head pose data and the audio data are bound to a master clock, and the audio data and the video data are synchronized by using an audio-video time difference, so as to realize timestamp alignment of the head pose data, the audio data and the video data.
6. An audio playback device, characterized by Comprise: A prediction module is configured to collect head pose data of a target user by using a head-mounted device, and to perform extrapolation prediction according to the head pose data to obtain a head predicted pose; A construction module is configured to construct a base vector of a user coordinate system according to the head predicted pose; A mapping module is configured to obtain a relative position of a sound source and the target user in a world coordinate system, and to map the relative position to the user coordinate system by using the base vector to obtain a sound source position; A determination module is configured to determine a virtual listening orientation of the target user according to the sound source position; A playing module is configured to determine a target loudspeaker according to the virtual listening orientation, and to perform audio playing by using the target loudspeaker; The prediction module comprises: An acquisition unit is configured to collect a head position and a head orientation of the target user by using the head-mounted device; A first extrapolation unit is configured to perform extrapolation prediction according to the head position to obtain a head predicted position; A second extrapolation unit is configured to perform extrapolation prediction according to the head orientation to obtain a head predicted orientation; A generation unit is configured to generate the head predicted pose according to the head predicted position and the head predicted orientation; The first extrapolation unit is specifically configured to obtain a head movement speed of the target user by using the head-mounted device; to perform smoothing processing on the head position and the head movement speed to obtain a target head position and a target head movement speed; to obtain an audio-video time difference as an extrapolation time length, and to perform extrapolation prediction on the target head position by using the target head movement speed and the extrapolation time length to obtain the head predicted position.
7. An electronic device, comprising: Comprise: A memory is configured to store a computer program; A processor is configured to execute the computer program to implement steps of the audio playing method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium, and the computer program is executed by the processor to implement steps of the audio playing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Three-dimensional audio signal generation method and system for non-spherical speaker array
CN105392102A
Information processing device, information processing system, information processing method, and program
WO2024247561A1