Spatial audio rendering method and apparatus

EP4648441A4Pending Publication Date: 2026-05-20HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2023-11-22
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

Spatial audio rendering at audio content provider ends, such as mobile devices, experiences long delays due to high computing power requirements, disrupting the immersive auditory experience, especially during head movements.

Method used

A method and apparatus that predict a user's head movement position to synchronize audio rendering, reducing delays by sending predicted head movement positions to the audio content provider end, allowing for real-time alignment with user movements.

Benefits of technology

Ensures that rendered audio frames accurately match the user's head position, providing a real and immersive auditory experience by offsetting delays in the rendering process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGAF001_ABST
    Figure IMGAF001_ABST
Patent Text Reader

Abstract

This application provides a spatial audio rendering method and apparatus. The spatial audio rendering method in this application includes: obtaining delay information of an audio content provider end; obtaining a predicted head movement position based on the delay information; and sending first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position. In this application, a delay can be eliminated, so that rendering effect exactly matches a head movement position of a user, to achieve real and immersive auditory experience.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202310108164.5, filed with the China National Intellectual Property Administration on January 30, 2023 and entitled "SPATIAL AUDIO RENDERING METHOD AND APPARATUS", which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] This application relates to audio processing technologies, and in particular, to a spatial audio rendering method and apparatus.BACKGROUND

[0003] In real three-dimensional space, sound may have two characteristics: a sense of orientation and a sense of space. A spatial audio technology is usually a technology that simulates immersive audio experience with the foregoing two characteristics on a head-mounted play device such as a headset. People's judgment on a sense of orientation for sound is mainly affected by a time difference, a sound level difference, human body filter effect learned during growth, head shaking, and other factors. The time difference, the sound level difference, and the human body filter effect may be comprehensively expressed as a head-related transfer function (Head-Related Transfer Function, HRTF). The head shaking in all directions is of great help for determining a position of a sound source. An indoor sound field may include direct sound, early reflected sound (Early Reflection, ER), and late reverberant sound (Late Reverb, LR). People's sense of space for sound is built based on the ER and the LR.

[0004] Sound sources in different formats, such as stereo, multi-channel surround sound, and a specially produced multi-object sound source, may be rendered into a binaural signal by using the spatial audio technology, so that real and immersive auditory experience can be enjoyed through a headset. Currently, head tracking (Head Tracking) may be performed through a gyroscope or another sensor in a headset, and then a corresponding rotation or displacement change is incorporated in spatial rendering of a sound source.

[0005] However, spatial audio rendering usually needs to be supported by high computing power, and a low delay is needed for head movement tracking. If spatial audio rendering is performed at an audio content provider end (for example, a mobile phone, a tablet computer, or a PC), although a computing power requirement can be met, a long delay occurs, and overall auditory experience is affected, especially in a head movement scenario.SUMMARY

[0006] This application provides a spatial audio rendering method and apparatus, to eliminate a delay, so that rendering effect exactly matches a head movement position of a user, to achieve real and immersive auditory experience.

[0007] According to a first aspect, this application provides a spatial audio rendering method, including: obtaining delay information of an audio content provider end; obtaining a predicted head movement position based on the delay information; and sending first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position.

[0008] In this application, a head movement position of a user is predicted to obtain a predicted head movement position, and after the predicted head movement position is sent to the audio content provider end, the audio content provider end may render an audio frame based on the predicted head movement position. In this way, an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position. In this case, when the rendered audio frame is sent to an audio signal play end, predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.

[0009] The delay information indicates a delay status of the audio content provider end, and may include the following cases.

[0010] In a first case, the delay information includes a delay sent by the audio content provider end.

[0011] After determining the delay caused by a transmission link, rendering, and the like of the audio content provider end, the audio content provider end may package the delay into first information to be sent to the audio signal play end, and send the first information to the audio signal play end. In this way, the audio signal play end can directly extract the delay of the audio content provider end from the first information that comes from the audio content provider end.

[0012] In a second case, the delay information includes a second historical head movement position sent by the audio content provider end.

[0013] The audio signal play end (for example, a headset) may periodically detect, through a sensor (for example, a gyroscope or a gravity sensor), a current position (referred to as a measured head movement position in this specification) of a head of a user wearing the headset. In this way, the audio signal play end can send the measured head movement position obtained through measurement to the audio content provider end.

[0014] After receiving the measured head movement position, the audio content provider end does not process the measured head movement position, but still performs related processing on an audio frame according to a specified process. When first information needs to be sent to the audio signal play end after the processing is completed, the measured head movement position is carried in the first information. In this case, because a period of time has elapsed, the measured head movement position becomes a historical head movement position (referred to as the second historical head movement position in this specification). It should be understood that, in a mathematical sense, the measured head movement position is the same as the second historical head movement position.

[0015] Based on this, the audio signal play end may obtain, based on time at which the second historical head movement position is received and time at which the measured head movement position is sent, a time difference between the time at which the second historical head movement position is received and the time at which the measured head movement position is sent (that is, calculate a time difference between receiving time and sending time of a same head movement position), to indirectly obtain the delay of the audio content provider end.

[0016] In a third case, the delay information includes a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end.

[0017] For the second historical head movement position, refer to the foregoing descriptions of the second case. Details are not described herein again.

[0018] The first uncertainty coefficient is an important parameter in a Kalman filter, and helps improve accuracy of a prediction result.

[0019] To enable a user to enjoy real and immersive auditory experience through the audio signal play end (for example, a headset), sound sources in different formats, such as stereo, multi-channel surround sound, and a specially produced multi-object sound source, may be rendered into a binaural signal by using a spatial audio technology. In the real world, when a head of a user rotates or moves, an absolute position of a sound source does not change, but a relative direction between the sound source and the head changes. For example, a guitar is being played in front of a user, and if the user turns to the right, sound of the guitar is correspondingly shifted to the left of the user. For another example, there is a guitar on a left side of a stage, and there is a saxophone on a right side. When a user moves to a side of the stage, sound of the guitar and sound of the saxophone overlap, and come from a same direction. It can be learned that the so-called real and immersive auditory experience is the auditory experience with the guitar or the stage performance described above. With a change in a position of a head of a user (including but not limited to head displacements in up, down, left, and right directions, head rotation, and the like), sound from a same audio source is rendered into different auditory effect in ears of the user. Such rendering is a rendering method with head movement effect. To be specific, head tracking (Head Tracking) is performed through a sensor in a headset, and then a corresponding rotation or displacement change is incorporated in spatial rendering of a sound source.

[0020] In addition, although the audio content provider end may provide high computing power to implement the foregoing rendering algorithm with head movement effect, due to characteristics of the audio content provider end, a long delay occurs, causing disorder of important spatial information such as an orientation. For example, an acoustic image of an object is designed to be directly in front of space. If a head of a user rotates, the acoustic image first moves to be directly in front of a face of the user, and then is restored to be directly in front of space. A clear delay and "sense of damping" occur, affecting auditory experience.

[0021] To resolve the foregoing problem caused by the delay, the audio signal play end may predict a head movement position of a user, and then send the predicted head movement position to the audio content provider end, so that the audio content provider end can render an audio frame based on the predicted head movement position. In this way, an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position. In this case, when the rendered audio frame is sent to the audio signal play end, predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.

[0022] In this application, the audio signal play end may first obtain a first historical head movement position corresponding to the delay information, and then obtain a measured head movement position through the sensor, and then obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.

[0023] Based on the foregoing three cases of the delay information, the audio signal play end may obtain the first historical head movement position by using three methods.

[0024] 1. When the delay information includes the delay sent by the audio content provider end (the first case), the first historical head movement position corresponding to the delay is extracted from a cache, where a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache.

[0025] As described above, the audio signal play end may periodically detect, through the sensor, a measured head movement position of a user wearing the audio signal play end. For ease of subsequent use, a correspondence between a measured head movement position and detection time may be stored on the audio signal play end. In this way, when the foregoing delay is obtained, a current measured head movement position (referred to as the first historical head movement position in this specification) corresponding to the delay may be obtained by querying the correspondence.

[0026] 2. When the delay information includes the second historical head movement position sent by the audio content provider end (the second case), the second historical head movement position is used as the first historical head movement position.

[0027] It can be learned from the foregoing descriptions of the second case that the audio content provider end directly packages the second historical head movement position into the first information and sends the first information to the audio signal play end. In this way, the audio signal play end can directly obtain a current measured head movement position (the second historical head movement position) corresponding to the delay. Therefore, the audio signal play end can directly determine the second historical head movement position as the first historical head movement position.

[0028] 3. When the delay information includes the second historical head movement position and the first uncertainty coefficient that are sent by the audio content provider end (the third case), the second historical head movement position is used as the first historical head movement position.

[0029] Refer to the foregoing descriptions in 2, in the third case of the delay information, the audio signal play end may alternatively directly determine the second historical head movement position as the first historical head movement position.

[0030] The audio signal play end may obtain a measured head movement position through the sensor, where the measured head movement position is a newly detected head movement position of the user.

[0031] In this application, obtaining the predicted head movement position based on the first historical head movement position and the measured head movement position may include the following two algorithms.

[0032] In a first algorithm, based on the foregoing 1 and 2, the audio signal play end may obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.

[0033] The difference between the first historical head movement position and the measured head movement position may be a distance between the first historical head movement position and the measured head movement position. For example, in a three-dimensional world, a head movement position may be represented by using coordinate axes x, y, and z. Therefore, the foregoing difference may be a distance between a coordinate value of the first historical head movement position and a coordinate value of the measured head movement position. In addition, a head movement position may alternatively be represented in another manner. This is not specifically limited in this application.

[0034] As described above, the audio signal play end may periodically detect, through the sensor, a measured head movement position of a user wearing the audio signal play end. Therefore, the audio signal play end may obtain a head movement change rate based on N measured head movement positions that are previously obtained, where the N measured head movement positions may include N measured head movement positions obtained through counting forward from a current measured head movement position.

[0035] The audio signal play end may predict a head movement change value based on the difference and the head movement change rate through cubic spline interpolation or by another means, and then add up the head movement change value and a current measured head movement position to obtain a predicted head movement position.

[0036] In a second algorithm, based on the foregoing 3, the audio signal play end obtains a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtains a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and inputs the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.

[0037] For a manner of obtaining the head movement change rate, refer to the foregoing descriptions.

[0038] The second uncertainty coefficient may be obtained based on the head movement change rate and the first uncertainty coefficient. The second uncertainty coefficient, the first historical head movement position, and the measured current head movement position are input to the Kalman filter. In the Kalman filter, a Kalman gain coefficient may be determined based on the second uncertainty coefficient and estimated uncertainty of a previous iteration, and the Kalman gain coefficient is used as a latest gain coefficient. In this way, the Kalman filter can output the predicted head movement position.

[0039] It should be noted that, in addition to the foregoing two methods, in this application, the predicted head movement position may alternatively be obtained based on the first historical head movement position and the measured head movement position by using another algorithm. This is not specifically limited herein.

[0040] The audio signal play end sends the first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position, so that the audio content provider end can perform rendering with head movement effect on an audio frame based on the predicted head movement position. In this way, an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position. In this case, when the rendered audio frame is sent to the audio signal play end, predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.

[0041] In a possible implementation, when a first rendering mode is used, the audio content provider end obtains a current frame. When an audio format of an audio source is a non-stereo format, the audio content provider end obtains a predicted head movement position. The audio content provider end renders the current frame based on the predicted head movement position to obtain a first binaural signal. When an audio format of an audio source is a stereo format, the audio content provider end does not perform spatial rendering on the audio source. The audio content provider end sends first information to the audio signal play end, where the first information includes a first audio stream, the audio format of the audio source, and the delay information. When the audio format indicates that the audio source is in the stereo format, the audio signal play end obtains a measured head movement position through the sensor. The audio signal play end renders the current frame based on the measured head movement position to obtain a second binaural signal. The audio signal play end plays audio based on a target binaural signal, where the target binaural signal includes the first binaural signal or the second binaural signal.

[0042] The first rendering mode may be a non-low-delay link mode. In this mode, the audio content provider end may perform rendering with head movement effect on an audio frame based on the predicted head movement position. A user may perform an operation on an application (application, APP) installed on the audio content provider end, to choose whether to select the first rendering mode. Correspondingly, this application further includes a second rendering mode. In this mode, the audio content provider end performs rendering without head movement effect on an audio frame. For this process, refer to the following descriptions. Similarly, the user may perform an operation on the APP installed on the audio content provider end, to choose whether to select the second rendering mode. For the operation of the user, refer to the following embodiments.

[0043] The current frame may be a frame of the audio source, and usually, may be an audio frame that is currently being processed by the audio content provider end.

[0044] The audio format of the audio source includes the stereo format or the non-stereo format, where the non-stereo format may include but is not limited to audio in a multi-channel or multi-object format such as 5.1, 7.1, or 3DA.

[0045] In this application, the audio content provider end may perform high-computing-power rendering on the audio in the multi-channel or multi-object format (non-stereo format) such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized. Based on this, when the audio format of the audio source is the non-stereo format, the audio content provider end may perform rendering with head movement effect on an audio frame. Therefore, the predicted head movement position needs to be obtained.

[0046] The audio content provider end may obtain the predicted head movement position from rendering information that is previously received from the audio signal play end. Optionally, the rendering information may be latest received rendering information. In this way, accuracy of the predicted head movement position can be improved. For obtaining of the predicted head movement position, refer to an embodiment shown in FIG. 3. Details are not described herein.

[0047] In this application, the audio content provider end may separately perform rendering with head movement effect on a direct sound part, an early reflected sound (Early Reflection, ER) part, and a late reverberant sound (Late Reverb, LR) part in the current frame based on the predicted head movement position, to obtain the first binaural signal.

[0048] Optionally, direct sound with head movement effect and reverberation (obtained by mixing ER and LR) with head movement effect are separately rendered by an audio effector and then mixed, and an output result obtained through mixing is rendered by the audio effector again and then transmitted in a form of the first binaural signal. The audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.

[0049] In this application, the audio content provider end does not need to render an audio source in a stereo format, and the audio signal play end performs rendering with head movement effect on the audio source. Therefore, the audio content provider end may perform high-computing-power rendering on the audio in the multi-channel or multi-object format (non-stereo format) such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized. Based on this, when the audio format of the audio source is the non-stereo format, the audio content provider end may perform rendering with head movement effect on an audio frame. Therefore, the audio content provider end only needs to package the audio source into the first audio stream.

[0050] As described above, when the audio format indicates that the audio source is in the non-stereo format, the first audio stream includes the first binaural signal rendered by the audio content provider end; or when the audio format indicates that the audio source is in the stereo format, the first audio stream includes the audio source.

[0051] For an audio source in a stereo format, the audio signal play end performs rendering with a head movement position on the audio source. Because an IMU data transmission link of the audio signal play end has a low delay, the audio signal play end can obtain a current head movement position (referred to as the measured head movement position in this specification) in real time through the sensor.

[0052] The audio signal play end may perform rendering with head movement effect on the direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on the ER and the LR in the current frame, to obtain the second binaural signal.

[0053] Optionally, direct sound with head movement effect and reverberation (obtained by mixing ER and LR) without head movement effect are separately rendered by an audio effector and then mixed, and an output result obtained through mixing is rendered by the audio effector again and then mixed again, to obtain the second binaural signal. The audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.

[0054] The audio signal play end obtains the target binaural signal through processing in the foregoing steps, where the target binaural signal may be the first binaural signal (the audio source is in the non-stereo format) or the second binaural signal (the audio source is in the stereo format).

[0055] In a possible implementation, when the second rendering mode is used, the audio content provider end obtains a current frame, performs rendering without head movement effect on the current frame to obtain a second binaural signal (the second binaural signal is different from the foregoing second binaural signal, and is referred to as a third binaural signal below for differentiation), and sends second information to the audio signal play end, where the second information includes a second audio stream, and the second audio stream includes the third binaural signal. The audio signal play end obtains a measured head movement position through the sensor, and renders a current frame (the current frame is an audio frame obtained by the audio content provider end by performing rendering without head movement effect) based on the measured head movement position to obtain a fourth binaural signal. The audio signal play end plays audio based on a target binaural signal, where the target binaural signal includes the fourth binaural signal.

[0056] In the second rendering mode, the audio content provider end performs rendering without head movement effect on an audio frame, so that computing power advantages of the audio content provider end can be fully utilized. A rendering delay can be shortened without head movement effect. Instead, the audio signal play end performs low-computing-power rendering with head movement effect on the audio frame. Compared with the first rendering mode, this achieves a lower delay, but rendering effect of head movement effect is poor.

[0057] According to a second aspect, this application provides a spatial audio rendering apparatus, including: an obtaining module, configured to: obtain delay information of an audio content provider end; a prediction module, configured to: obtain a predicted head movement position based on the delay information; and a sending module, configured to: send first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position.

[0058] In a possible implementation, the prediction module is specifically configured to: obtain a first historical head movement position corresponding to the delay information; obtain a measured head movement position through a sensor; and obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.

[0059] In a possible implementation, the delay information includes a delay sent by the audio content provider end; and the prediction module is specifically configured to: extract, from a cache, the first historical head movement position corresponding to the delay, where a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache.

[0060] In a possible implementation, the delay information includes a second historical head movement position sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, and the second rendering information is sent earlier than the first rendering information; and the prediction module is specifically configured to: use the second historical head movement position as the first historical head movement position.

[0061] In a possible implementation, the prediction module is specifically configured to: obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.

[0062] In a possible implementation, the delay information includes a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, the first uncertainty coefficient comes from the second rendering information, and the second rendering information is sent earlier than the first rendering information; and the prediction module is specifically configured to: use the second historical head movement position as the first historical head movement position.

[0063] In a possible implementation, the prediction module is specifically configured to: obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtain a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and input the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.

[0064] In a possible implementation, the apparatus further includes: a receiving module, configured to: receive first information sent by the audio content provider end, where the first information includes an audio format, a first audio stream, and the delay information, the audio format indicates that an audio source is in a stereo format or a non-stereo format, and when the audio format indicates that the audio source is in the non-stereo format, the first audio stream includes a first binaural signal rendered by the audio content provider end, or when the audio format indicates that the audio source is in the stereo format, the first audio stream includes the audio source; a rendering module, configured to: when the audio format indicates that the audio source is in the stereo format, obtain a measured head movement position through the sensor; and render a current frame based on the measured head movement position to obtain a second binaural signal, where the current frame is a frame of the audio source; and a playing module, configured to: play audio based on a target binaural signal, where the target binaural signal includes the first binaural signal or the second binaural signal.

[0065] In a possible implementation, the rendering module is specifically configured to: perform rendering with head movement effect on a direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on early reflected sound ER and late reverberant sound LR in the current frame, to obtain the second binaural signal.

[0066] According to a third aspect, this application provides a spatial audio rendering apparatus, including: an obtaining module, configured to: obtain a current frame when a first rendering mode is used, where the current frame is a frame of an audio source; and obtain a predicted head movement position when an audio format of the audio source is a non-stereo format; a rendering module, configured to: render the current frame based on the predicted head movement position to obtain a first binaural signal; and a sending module, configured to: send first information to an audio signal play end, where the first information includes a first audio stream, the audio format of the audio source, and delay information, and the first audio stream includes the first binaural signal.

[0067] In a possible implementation, the rendering module is specifically configured to: separately perform rendering with head movement effect on a direct sound part, an early reflected sound ER part, and a late reverberant sound LR part in the current frame based on the predicted head movement position, to obtain the first binaural signal.

[0068] In a possible implementation, the obtaining module is specifically configured to: obtain the predicted head movement position from rendering information sent by the audio signal play end.

[0069] In a possible implementation, the delay information includes a delay.

[0070] In a possible implementation, the rendering information further includes a measured head movement position in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information includes the measured head movement position.

[0071] In a possible implementation, the rendering information further includes a measured head movement position and a first uncertainty coefficient in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information includes the measured head movement position and the first uncertainty coefficient.

[0072] In a possible implementation, when an audio format of the audio source is a stereo format, the first audio stream includes the audio source.

[0073] In a possible implementation, the obtaining module is further configured to: obtain the current frame when a second rendering mode is used; the rendering module is further configured to: perform rendering without head movement effect on the current frame to obtain a second binaural signal; and the sending module is further configured to: send second information to the audio signal play end, where the second information includes a second audio stream, and the second audio stream includes the second binaural signal.

[0074] According to a fourth aspect, this application provides an audio signal play device, including: one or more processors; and a memory, configured to store one or more programs, where when the one or more programs is / are executed by the one or more processors, the one or more processors is / are enabled to implement the method implemented by the audio signal play end according to any one of the implementations of the first aspect.

[0075] According to a fifth aspect, this application provides an audio content providing device, including: one or more processors; and a memory, configured to store one or more programs, where when the one or more programs is / are executed by the one or more processors, the one or more processors is / are enabled to implement the method implemented by the audio content provider end according to any one of the implementations of the first aspect.

[0076] According to a sixth aspect, this application provides a computer-readable storage medium, including a computer program. When the computer program is executed on a computer, the computer is enabled to perform the method according to any one of the implementations of the first aspect.

[0077] According to a seventh aspect, this application provides a computer program product. The computer program product includes computer program code. When the computer program code is run on a computer, the computer is enabled to perform the method according to any one of the implementations of the first aspect.BRIEF DESCRIPTION OF DRAWINGS

[0078] FIG. 1A and FIG. 1B are a diagram of an application scenario according to this application; FIG. 2 is a diagram of a structure of an audio content provider end 200 according to this application; FIG. 3 is a flowchart of a process 300 of a spatial audio rendering method according to this application; FIG. 4 is a flowchart of a process 400 of a spatial audio rendering method according to this application; FIG. 5 is a diagram of an overall process of a spatial audio rendering method according to this application; FIG. 6 is a diagram of a specific process of a spatial audio rendering method according to this application; FIG. 7 is a schematic flowchart of an algorithm for predicting a head movement position; FIG. 8a and FIG. 8b are a schematic flowchart of an algorithm for predicting a head movement position; FIG. 9a and FIG. 9b are a schematic flowchart of an algorithm for predicting a head movement position; FIG. 10 is a schematic framework flowchart according to this application; FIG. 11 is a diagram of a process in a low-delay mode according to this application; FIG. 12 is a diagram of a process in a non-low-delay mode according to this application; FIG. 13 is a diagram of an example structure of a spatial audio rendering apparatus 1300 according to this application; and FIG. 14 is a diagram of an example structure of a spatial audio rendering apparatus 1400 according to this application. DESCRIPTION OF EMBODIMENTS

[0079] To make objectives, technical solutions, and advantages of this application clearer, the following clearly and completely describes the technical solutions in this application with reference to accompanying drawings in this application. Clearly, the described embodiments are merely some but not all of embodiments of this application. All other embodiments obtained by a person of ordinary skill in the art based on embodiments of this application without creative efforts shall fall within the protection scope of this application.

[0080] In embodiments of this specification, claims, and accompanying drawings of this application, the terms "first", "second", and the like are merely intended for differentiation in descriptions, but shall not be construed as indicating or implying relative importance or indicating or implying a sequence. In addition, the terms "include", "have", and any variant thereof are intended to cover non-exclusive inclusion, for example, include a series of steps or units. A method, system, product, or device is not necessarily limited to those expressly listed steps or units, but may include other steps or units that are not expressly listed or that are inherent to such a process, method, product, or device.

[0081] It should be understood that, in this application, "at least one" means one or more, and "a plurality of" means two or more. "And / or" describes an association relationship between associated objects, and indicates that three relationships may exist. For example, "A and / or B" may indicate the following three cases: Only A exists, only B exists, and both A and B exist, where A and B may be in a singular form or a plural form. The character " / " usually indicates an "or" relationship between the associated objects. "At least one of the following items" or a similar expression thereof indicates any combination of the items, including one of the items or any combination of a plurality of the items. For example, at least one of a, b, or c may indicate a, b, c, "a and b", "a and c", "b and c", or "a, b, and c", where a, b, and c may be in a singular form or a plural form.

[0082] In real three-dimensional space, sound may have two characteristics: a sense of orientation and a sense of space. A spatial audio technology is usually a technology that simulates immersive audio experience with the foregoing two characteristics on a head-mounted play device such as a headset. People's judgment on a sense of orientation for sound is mainly affected by a time difference, a sound level difference, human body filter effect learned during growth, head shaking, and other factors. The time difference, the sound level difference, and the human body filter effect may be comprehensively expressed as a head-related transfer function (Head-Related Transfer Function, HRTF). The head shaking in all directions is of great help for determining a position of a sound source. An indoor sound field may include direct sound, early reflected sound (Early Reflection, ER), and late reverberant sound (Late Reverb, LR). People's sense of space for sound is built based on the ER and the LR.

[0083] Sound sources in different formats, such as stereo, multi-channel surround sound, and a specially produced multi-object sound source, may be rendered into a binaural signal by using the spatial audio technology, so that real and immersive auditory experience can be enjoyed through a headset.

[0084] When a head of a user rotates or moves, an absolute position of a sound source does not change, but a relative direction between the sound source and the head changes. For example, a guitar is being played in front of a user, and if the user turns to the right, sound of the guitar is correspondingly shifted to the left of the user. For another example, there is a guitar on a left side of a stage, and there is a saxophone on a right side. When a user moves to a side of the stage, sound of the guitar and sound of the saxophone overlap, and come from a same direction. Currently, head tracking (Head Tracking) may be performed through a gyroscope or another sensor in a headset, and then a corresponding rotation or displacement change is incorporated in spatial rendering of a sound source.

[0085] However, spatial audio rendering usually needs to be supported by high computing power, and a low delay is needed for head movement tracking. If spatial audio rendering is performed at an audio content provider end (for example, a mobile phone, a tablet computer, or a PC), although a computing power requirement can be met, a long delay occurs. Especially in a head movement scenario, for an audio source in a multi-channel format, a multi-object format, or the like, important spatial information such as an orientation is disordered, leading to degradation of overall quality of spatial audio rendering and affecting overall auditory experience.

[0086] To resolve the foregoing technical problems, this application provides a spatial audio rendering method and apparatus. The following embodiments describe the technical solutions of this application.

[0087] FIG. 1A and FIG. 1B are a diagram of an application scenario according to this application. As shown in FIG. 1A and FIG. 1B, the scenario includes an audio content provider end and an audio signal play end, and an interconnection mode between the audio content provider end and the audio signal play end includes but is not limited to a Bluetooth technology. The audio content provider end may include but is not limited to a mobile phone, a tablet computer, a notebook computer, a desktop computer, and the like, and may provide high computing power, but may have a long delay. The audio signal play end may include but is not limited to a true wireless stereo (true wireless stereo, TWS) headset, a wireless head-mounted headset, a wireless neckband headset, and the like, and has a low delay, but provides low computing power. In addition, a gyroscope or another sensor is disposed in a device serving as the audio signal play end, to capture head movement information of a user.

[0088] In this application, a rendering algorithm includes two parts. One part of the algorithm is deployed at the audio content provider end, and is used to perform high-computing-power rendering on audio in a multi-channel or multi-object format such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized. The other part of the algorithm is deployed at the audio signal play end, and is used to perform low-computing-power rendering on audio in a stereo format. An inertial measurement unit (Inertial Measurement Unit, IMU) data transmission link of the audio signal play end has a low delay, and may also run independently, without relying on a specific audio content provider end, to meet a requirement of a user for connecting to a device of any audio content provider end.

[0089] Based on this, this application further provides an algorithm for predicting a head movement position of a user, to predict a future head movement position of the user. The predicted head movement position is used to provide assistance for the rendering algorithm at the audio content provider end, to reduce impact of a delay of the audio content provider end.

[0090] It should be noted that the application scenario shown in FIG. 1A and FIG. 1B is an example, but this should not constitute any limitation on this application. An application scenario of the spatial audio rendering method is not specifically limited in this application either.

[0091] FIG. 2 is a diagram of a structure of an audio content provider end 200 according to this application. It should be understood that the audio content provider end 200 shown in FIG. 2 is merely an example, and the audio content provider end 200 may have more or fewer components than those shown in the figure, two or more components may be combined, or there may be different component configurations. The components shown in FIG. 2 may be implemented in hardware including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.

[0092] The audio content provider end 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (universal serial bus, USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headset jack 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display 294, a subscriber identity module (subscriber identity module, SIM) card interface 295, and the like. The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, an optical proximity sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, and the like.

[0093] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (application processor, AP), a modem processor, a graphics processing unit (graphics processing unit, GPU), an image signal processor (image signal processor, ISP), a controller, a memory, a video codec, a digital signal processor (digital signal processor, DSP), a baseband processor, and / or a neural-network processing unit (neural-network processing unit, NPU). Different processing units may be independent components, or may be integrated into one or more processors.

[0094] The controller may be a nerve center and a command center of the audio content provider end 200. The controller may generate an operation control signal based on an instruction operation code and a time sequence signal, to control instruction reading and instruction execution.

[0095] A memory may be further disposed in the processor 210 to store instructions and data. In some embodiments, the memory in the processor 210 is a cache. The memory may store instructions or data that has been used or is cyclically used by the processor 210. If the processor 210 needs to use the instructions or the data again, the processor may directly invoke the instructions or the data from the memory. This avoids repeated access, reduces waiting time of the processor 210, and therefore improves system efficiency.

[0096] In some embodiments, the processor 210 may include one or more interfaces. The interface may include an inter-integrated circuit (inter-integrated circuit, I2C) interface, an inter-integrated circuit sound (inter-integrated circuit sound, I2S) interface, a pulse code modulation (pulse code modulation, PCM) interface, a universal asynchronous receiver / transmitter (universal asynchronous receiver / transmitter, UART) interface, a mobile industry processor interface (mobile industry processor interface, MIPI), a general-purpose input / output (general-purpose input / output, GPIO) interface, a subscriber identity module (subscriber identity module, SIM) interface, a universal serial bus (universal serial bus, USB) interface, and / or the like.

[0097] The I2C interface is a two-way synchronous serial bus, and includes a serial data line (serial data line, SDA) and a serial clock line (serial clock line, SCL). In some embodiments, the processor 210 may include a plurality of groups of I2C buses. The processor 210 may be separately coupled to the touch sensor 280K, a charger, a flash, the camera 293, and the like through different I2C bus interfaces. For example, the processor 210 may be coupled to the touch sensor 280K through the I2C interface, so that the processor 210 communicates with the touch sensor 280K through the I2C bus interface, to implement a touch function of the audio content provider end 200.

[0098] The I2S interface may be used for audio communication. In some embodiments, the processor 210 may include a plurality of groups of I2S buses. The processor 210 may be coupled to the audio module 270 through the I2S bus, to implement communication between the processor 210 and the audio module 270. In some embodiments, the audio module 270 may transmit an audio signal to the wireless communication module 260 through the I2S interface, to implement a function of answering a call through a Bluetooth headset.

[0099] The PCM interface may also be used for audio communication, and sampling, quantization, and encoding of an analog signal. In some embodiments, the audio module 270 may be coupled to the wireless communication module 260 through the PCM bus interface. In some embodiments, the audio module 270 may alternatively transmit an audio signal to the wireless communication module 260 through the PCM interface, to implement a function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface may be used for audio communication.

[0100] The UART interface is a universal serial data bus, and is used for asynchronous communication. The bus may be a two-way communication bus. The bus converts to-be-transmitted data between serial communication and parallel communication. In some embodiments, the UART interface is usually configured to connect the processor 210 to the wireless communication module 260. For example, the processor 210 communicates with a Bluetooth module in the wireless communication module 260 through the UART interface, to implement a Bluetooth function. In some embodiments, the audio module 270 may transmit an audio signal to the wireless communication module 260 through the UART interface, to implement a function of playing music through a Bluetooth headset.

[0101] The MIPI interface may be configured to connect the processor 210 to a peripheral component such as the display 294 or the camera 293. The MIPI interface includes a camera serial interface (camera serial interface, CSI), a display serial interface (display serial interface, DSI), and the like. In some embodiments, the processor 210 communicates with the camera 293 through the CSI interface, to implement an image shooting function of the audio content provider end 200. The processor 210 communicates with the display 294 through the DSI interface, to implement a display function of the audio content provider end 200.

[0102] The GPIO interface may be configured by software. The GPIO interface may be configured as a control signal or a data signal. In some embodiments, the GPIO interface may be configured to connect the processor 210 to the camera 293, the display 294, the wireless communication module 260, the audio module 270, the sensor module 280, and the like. The GPIO interface may alternatively be configured as an I2C interface, an I2S interface, a UART interface, an MIPI interface, or the like.

[0103] The USB interface 230 is an interface that conforms to a USB standard specification, and may be specifically a mini USB interface, a micro USB interface, a USB Type-C interface, or the like. The USB interface 230 may be used for connecting a charger to charge the audio content provider end 200, or may be configured to transmit data between the audio content provider end 200 and a peripheral device, or may be used for connecting a headset for playing audio through the headset. The interface may alternatively be used for connecting other user equipment, for example, an AR device.

[0104] It can be understood that an interface connection relationship between the modules shown in this embodiment of this application is merely an example for description, and does not constitute a limitation on a structure of the audio content provider end 200. In some other embodiments of this application, the audio content provider end 200 may alternatively use an interface connection mode different from that in the foregoing embodiment, or use a combination of a plurality of interface connection modes.

[0105] The charging management module 240 is configured to receive charging input from the charger. The charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 240 may receive charging input from a wired charger through the USB interface 230. In some embodiments of wireless charging, the charging management module 240 may receive wireless charging input through a wireless charging coil of the audio content provider end 200. When charging the battery 242, the charging management module 240 may further supply power to user equipment through the power management module 241.

[0106] The power management module 241 is configured to connect to the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240, and supplies power to the processor 210, the internal memory 221, an external memory, the display 294, the camera 293, the wireless communication module 260, and the like. The power management module 241 may be further configured to monitor parameters such as a battery capacity, a quantity of battery cycles, and a battery health status (electric leakage and impedance). In some other embodiments, the power management module 241 may alternatively be disposed in the processor 210. In some other embodiments, the power management module 241 and the charging management module 240 may alternatively be disposed in a same component.

[0107] A wireless communication function of the audio content provider end 200 may be implemented by the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modem processor, the baseband processor, and the like.

[0108] The antenna 1 and the antenna 2 are configured to transmit and receive an electromagnetic wave signal. Each antenna in the audio content provider end 200 may be configured to cover one or more communication frequency bands. Different antennas may be further multiplexed to improve antenna utilization. For example, the antenna 1 may be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.

[0109] The mobile communication module 250 may provide a solution applied to the audio content provider end 200 for wireless communication such as 2G / 3G / 4G / 5G. The mobile communication module 250 may include at least one filter, a switch, a power amplifier, a low noise amplifier (low noise amplifier, LNA), and the like. The mobile communication module 250 may receive an electromagnetic wave through the antenna 1, perform processing such as filtering or amplification on the received electromagnetic wave, and transmit a processed electromagnetic wave to the modem processor for demodulation. The mobile communication module 250 may further amplify a signal modulated by the modem processor, and convert an amplified signal into an electromagnetic wave for radiation through the antenna 1. In some embodiments, at least some functional modules of the mobile communication module 250 may be disposed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 may be disposed in a same component as at least some modules of the processor 210.

[0110] The modem processor may include a modulator and a demodulator. The modulator is configured to modulate a to-be-sent low-frequency baseband signal into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. Then the demodulator transmits the low-frequency baseband signal obtained through demodulation to the baseband processor for processing. The low-frequency baseband signal is processed by the baseband processor and then transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 270A, the receiver 270B, and the like), or displays an image or a video through the display 294. In some embodiments, the modem processor may be an independent component. In some other embodiments, the modem processor may be independent of the processor 210, and is disposed in a same component as the mobile communication module 250 or another functional module.

[0111] The wireless communication module 260 may provide a solution applied to the audio content provider end 200 for wireless communication such as a wireless local area network (wireless local area network, WLAN) (for example, a wireless fidelity (wireless fidelity, Wi-Fi) network), Bluetooth (Bluetooth, BT), a global navigation satellite system (global navigation satellite system, GNSS), frequency modulation (frequency modulation, FM), a near field communication (near field communication, NFC) technology, or an infrared (infrared, IR) technology. The wireless communication module 260 may be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives an electromagnetic wave through the antenna 2, performs frequency modulation and filtering on an electromagnetic wave signal, and sends a processed signal to the processor 210. The wireless communication module 260 may further receive a to-be-sent signal from the processor 210, perform frequency modulation and amplification on the signal, and convert a processed signal into an electromagnetic wave for radiation through the antenna 2.

[0112] In some embodiments, the antenna 1 of the audio content provider end 200 is coupled to the mobile communication module 250, and the antenna 2 is coupled to the wireless communication module 260, so that the audio content provider end 200 can communicate with a network and another device by using a wireless communication technology. The wireless communication technology may include a global system for mobile communications (global system for mobile communications, GSM), a general packet radio service (general packet radio service, GPRS), code division multiple access (code division multiple access, CDMA), wideband code division multiple access (wideband code division multiple access, WCDMA), time-division code division multiple access (time-division code division multiple access, TD-SCDMA), long term evolution (long term evolution, LTE), BT, a GNSS, a WLAN, NFC, FM, an IR technology, and / or the like. The GNSS may include a global positioning system (global positioning system, GPS), a global navigation satellite system (global navigation satellite system, GLONASS), a BeiDou navigation satellite system (BeiDou navigation satellite system, BDS), a quasi-zenith satellite system (quasi-zenith satellite system, QZSS), and / or a satellite-based augmentation system (satellite-based augmentation system, SBAS).

[0113] The audio content provider end 200 implements a display function through the GPU, the display 294, the application processor, and the like. The GPU is a microprocessor for image processing, and is connected to the display 294 and the application processor. The GPU is configured to perform mathematical and geometric computation, and render an image. The processor 210 may include one or more GPUs that execute program instructions to generate or change displayed information.

[0114] The display 294 is configured to display an image, a video, or the like. The display 294 includes a display panel. The display panel may be a liquid crystal display (liquid crystal display, LCD), an organic light-emitting diode (organic light-emitting diode, OLED), an active-matrix organic light-emitting diode (active-matrix organic light-emitting diode, AMOLED), a flexible light-emitting diode (flex light-emitting diode, FLED), a mini-LED, a micro-LED, a micro-OLED, a quantum dot light-emitting diode (quantum dot light-emitting diode, QLED), or the like. In some embodiments, the audio content provider end 200 may include one or N displays 294, where N is a positive integer greater than 1.

[0115] The audio content provider end 200 may implement an image shooting function through the ISP, the camera 293, the video codec, the GPU, the display 294, the application processor, and the like.

[0116] The ISP is configured to process data fed back by the camera 293. For example, during photographing, a shutter is pressed, and light is transmitted to a photosensitive element of the camera through a lens. An optical signal is converted into an electrical signal, and the photosensitive element of the camera transmits the electrical signal to the ISP for processing, to convert the electrical signal into a visible image. The ISP may further perform algorithm optimization on noise, brightness, and complexion of the image. The ISP may further optimize parameters such as exposure and color temperature of an image shooting scene. In some embodiments, the ISP may be disposed in the camera 293.

[0117] The camera 293 is configured to capture a static image or a video. An optical image of an object is generated through the lens, and is projected onto the photosensitive element. The photosensitive element may be a charge coupled device (charge coupled device, CCD) or a complementary metal-oxide-semiconductor (complementary metal-oxide-semiconductor, CMOS) phototransistor. The photosensitive element converts an optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert the electrical signal into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format, for example, RGB or YUV. In some embodiments, the audio content provider end 200 may include one or N cameras 293, where N is a positive integer greater than 1.

[0118] The digital signal processor is configured to process a digital signal, and may further process other digital signals in addition to the digital image signal. For example, when the audio content provider end 200 selects a frequency, the digital signal processor is configured to perform Fourier transform on frequency energy.

[0119] The video codec is configured to compress or decompress a digital video. The audio content provider end 200 may support one or more types of video codecs. In this way, the audio content provider end 200 can play or record videos in a plurality of coding formats, for example, moving picture experts group (moving picture experts group, MPEG)-1, MPEG-2, MPEG-3, and MPEG-4.

[0120] The NPU is a neural-network (neural-network, NN) computing processor. The NPU quickly processes input information with reference to a structure of a biological neural network, for example, a mode of transfer between human brain neurons, and may further continuously perform self-learning. Intelligent cognition applications, such as image recognition, facial recognition, speech recognition, and text understanding, of the audio content provider end 200 may be implemented through the NPU.

[0121] The external memory interface 220 may be used for connecting an external memory card, for example, a microSD card, to extend a storage capability of the audio content provider end 200. The external memory card communicates with the processor 210 through the external memory interface 220, to implement a data storage function. For example, files such as music and videos are stored in the external storage card.

[0122] The internal memory 221 may be configured to store computer-executable program code, and the executable program code includes instructions. The processor 210 runs the instructions stored in the internal memory 221 to implement various function applications and data processing of the audio content provider end 200. The internal memory 221 may include a program storage area and a data storage area. The program storage area may store an operating system, an application for at least one function (for example, a sound play function or an image play function), and the like. The data storage area may store data (for example, audio data and an address book) created during use of the audio content provider end 200, and the like. In addition, the internal memory 221 may include a high-speed random access memory, or may include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory, or a universal flash storage (universal flash storage, UFS).

[0123] The audio content provider end 200 may implement an audio function, for example, music playing or recording, through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the headset jack 270D, the application processor, and the like.

[0124] The audio module 270 is configured to convert digital audio information into an analog audio signal for output, and is also configured to convert analog audio input into a digital audio signal. The audio module 270 may be further configured to encode and decode an audio signal. In some embodiments, the audio module 270 may be disposed in the processor 210, or some functional modules of the audio module 270 are disposed in the processor 210.

[0125] The speaker 270A, also referred to as a "loudspeaker", is configured to convert an electrical audio signal into a sound signal. The audio content provider end 200 may be used to listen to music or answer a call in a hands-free mode through the speaker 270A.

[0126] The receiver 270B, also referred to as an "earpiece", is configured to convert an electrical audio signal into a sound signal. When the audio content provider end 200 is used to answer a call or listen to a voice message, the receiver 270B may be put close to a human ear to listen to a voice.

[0127] The microphone 270C, also referred to as a "mike" or a "mic", is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, a user may make a sound near the microphone 270C through the mouth of the user, to input a sound signal to the microphone 270C. At least one microphone 270C may be disposed at the audio content provider end 200. In some other embodiments, two microphones 270C may be disposed at the audio content provider end 200, to capture a sound signal and implement a noise reduction function. In some other embodiments, three, four, or more microphones 270C may alternatively be disposed at the audio content provider end 200, to capture a sound signal, reduce noise, and recognize a sound source, to implement a directional recording function and the like.

[0128] The headset jack 270D is used for connecting a wired headset. The headset jack 270D may be the USB interface 230, or may be a 3.5 mm open mobile terminal platform (open mobile terminal platform, OMTP) standard interface, or a cellular telecommunications industry association of the USA (cellular telecommunications industry association of the USA, CTIA) standard interface.

[0129] The pressure sensor 280A is configured to sense a pressure signal, and may convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 280A may be disposed on the display 294. There are many types of pressure sensors 280A, for example, a resistive pressure sensor, an inductive pressure sensor, and a capacitive pressure sensor. The capacitive pressure sensor may include at least two parallel plates made of conductive materials. When a force is applied to the pressure sensor 280A, capacitance between electrodes changes. The audio content provider end 200 determines pressure strength based on the capacitance change. When a touch operation is performed on the display 294, the audio content provider end 200 detects strength of the touch operation through the pressure sensor 280A. The audio content provider end 200 may also calculate a touch position based on a detection signal of the pressure sensor 280A. In some embodiments, touch operations that are performed at a same touch position but have different touch operation strength may correspond to different operation instructions. For example, when a touch operation whose touch operation strength is less than a first pressure threshold is performed on an SMS message application icon, an instruction for viewing an SMS message is executed. When a touch operation whose touch operation strength is greater than or equal to the first pressure threshold is performed on the SMS message application icon, an instruction for creating a new SMS message is executed.

[0130] The gyroscope sensor 280B may be configured to determine a motion attitude of the audio content provider end 200. In some embodiments, an angular velocity of the audio content provider end 200 around three axes (that is, axes x, y, and z) may be determined by the gyroscope sensor 280B. The gyroscope sensor 280B may be configured to implement image stabilization during image shooting. For example, when the shutter is pressed, the gyroscope sensor 280B detects an angle at which the audio content provider end 200 shakes, calculates, based on the angle, a distance for which a lens module needs to compensate, and allows the lens to cancel the shake of the audio content provider end 200 through reverse motion, to implement image stabilization. The gyroscope sensor 280B may be further used in a navigation scenario and a somatic game scenario.

[0131] The barometric pressure sensor 280C is configured to measure barometric pressure. In some embodiments, the audio content provider end 200 calculates an altitude based on a value of the barometric pressure measured by the barometric pressure sensor 280C, to provide assistance for positioning and navigation.

[0132] The magnetic sensor 280D includes a Hall effect sensor. The audio content provider end 200 may detect opening and closing of a flip leather case through the magnetic sensor 280D. In some embodiments, when the audio content provider end 200 is a clamshell phone, the audio content provider end 200 may detect opening and closing of a flip cover based on the magnetic sensor 280D. Further, a feature such as automatic unlocking of the flip cover is set based on a detected opening / closing state of the leather case or a detected opening / closing state of the flip cover.

[0133] The acceleration sensor 280E may detect accelerations of the audio content provider end 200 in various directions (usually on three axes), may detect a magnitude and a direction of gravity when the audio content provider end 200 is still, and may be further configured to recognize an attitude of user equipment and used in applications such as landscape / portrait mode switching and a pedometer.

[0134] The distance sensor 280F is configured to measure a distance. The audio content provider end 200 may measure a distance by using infrared or laser. In some embodiments, in an image shooting scenario, the audio content provider end 200 may measure a distance through the distance sensor 280F, to implement quick focusing.

[0135] The optical proximity sensor 280G may include, for example, a light-emitting diode (LED) and an optical detector, for example, a photodiode. The light-emitting diode may be an infrared light-emitting diode. The audio content provider end 200 transmits infrared light to the outside through the light-emitting diode. The audio content provider end 200 detects infrared reflected light from a nearby object through the photodiode. When sufficient reflected light is detected, it can be determined that an object exists near the audio content provider end 200. When insufficient reflected light is detected, the audio content provider end 200 may determine that no object exists near the audio content provider end 200. The audio content provider end 200 may detect, through the optical proximity sensor 280G, that a user holds the audio content provider end 200 close to an ear for a call, to automatically turn off a screen for power saving. The optical proximity sensor 280G may also be used in a leather case mode or a pocket mode for automatic screen unlocking or locking.

[0136] The ambient light sensor 280L is configured to sense ambient light brightness. The audio content provider end 200 may adaptively adjust brightness of the display 294 based on the sensed ambient light brightness. The ambient light sensor 280L may also be configured to automatically adjust white balance during photographing. The ambient light sensor 280L may further cooperate with the optical proximity sensor 280G to detect whether the audio content provider end 200 is in a pocket, to avoid an accidental touch.

[0137] The fingerprint sensor 280H is configured to capture a fingerprint. The audio content provider end 200 may implement fingerprint-based unlocking, application lock access, fingerprint-based photographing, fingerprint-based call answering, and the like by using a feature of the captured fingerprint.

[0138] The temperature sensor 280J is configured to detect temperature. In some embodiments, the audio content provider end 200 executes a temperature processing policy based on the temperature detected by the temperature sensor 280J. For example, when the temperature reported by the temperature sensor 280J exceeds a threshold, the audio content provider end 200 degrades performance of a processor near the temperature sensor 280J, to reduce power consumption for thermal protection. In some other embodiments, when the temperature is lower than another threshold, the audio content provider end 200 heats the battery 242 to avoid abnormal shutdown of the audio content provider end 200 due to low temperature. In some other embodiments, when the temperature is lower than still another threshold, the audio content provider end 200 boosts an output voltage of the battery 242 to avoid abnormal shutdown due to low temperature.

[0139] The touch sensor 280K is also referred to as a "touch panel". The touch sensor 280K may be disposed on the display 294. The touch sensor 280K and the display 294 constitute a touchscreen, which is also referred to as a "touch control screen". The touch sensor 280K is configured to detect a touch operation performed on or near the touch sensor. The touch sensor may transmit the detected touch operation to the application processor to determine a type of a touch event. The display 294 may provide visual output related to the touch operation. In some other embodiments, the touch sensor 280K may alternatively be disposed on a surface of the audio content provider end 200, and a position of the touch sensor 280K is different from a position of the display 294.

[0140] The bone conduction sensor 280M may obtain a vibration signal. In some embodiments, the bone conduction sensor 280M may obtain a vibration signal of a vibration bone of a human vocal-cord part. The bone conduction sensor 280M may also be in contact with a body pulse to receive a blood pressure beating signal. In some embodiments, the bone conduction sensor 280M may alternatively be disposed in a headset, to constitute a bone conduction headset. The audio module 270 may obtain a speech signal through parsing based on the vibration signal, obtained by the bone conduction sensor 280M, of the vibration bone of the vocal-cord part, to implement a speech function. The application processor may parse heart rate information based on the blood pressure beating signal obtained by the bone conduction sensor 280M, to implement a heart rate detection function.

[0141] The button 290 includes a power button, a volume button, and the like. The button 290 may be a mechanical button or a touch button. The audio content provider end 200 may receive key input, and generate key signal input related to user settings and function control of the audio content provider end 200.

[0142] The motor 291 may generate a vibration prompt. The motor 291 may be configured to provide an incoming call vibration prompt or a touch vibration feedback. For example, touch operations performed on different applications (for example, photographing and audio play) may correspond to different vibration feedback effect. The motor 291 may also correspond to different vibration feedback effect for touch operations performed on different areas of the display 294. Different application scenarios (for example, a time reminder, information receiving, an alarm clock, and a game) may also correspond to different vibration feedback effect. Touch vibration feedback effect may be further customized.

[0143] The indicator 292 may be an indicator light, and may be configured to indicate a charging status and a battery level change, or may be configured to indicate a message, a missed call, a notification, and the like.

[0144] The SIM card interface 295 is used for connecting a SIM card. The SIM card may be inserted into the SIM card interface 295 or removed from the SIM card interface 295, to implement contact with or separation from the audio content provider end 200. The audio content provider end 200 may support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 may support a nano-SIM card, a micro-SIM card, a SIM card, and the like. A plurality of cards may be inserted in a same SIM card interface 295 at the same time. The plurality of cards may be of a same type or different types. The SIM card interface 295 is also compatible with different types of SIM cards. The SIM card interface 295 is also compatible with an external memory card. The audio content provider end 200 interacts with a network through the SIM card, to implement functions such as calling and data communication. In some embodiments, the audio content provider end 200 uses an eSIM, namely, an embedded SIM card. The eSIM card may be embedded in the audio content provider end 200, and cannot be separated from the audio content provider end 200.

[0145] It can be understood that the structure shown in this embodiment of the present invention does not constitute a specific limitation on a controlling device. In some other embodiments of this application, the controlling device may include more or fewer components than those shown in the figure, or some components may be combined, or some components may be split, or different component layouts may be used. The components shown in the figure may be implemented by hardware, software, or a combination of software and hardware.

[0146] Based on the foregoing embodiments, FIG. 3 is a flowchart of a process 300 of a spatial audio rendering method according to this application. As shown in FIG. 3, the process 300 may be applied to the application scenario shown in FIG. 1A and FIG. 1B, and is performed by an audio signal play end to obtain a predicted head movement position, so that an audio content provider end uses the predicted head movement position when rendering audio. The process 300 is described as a series of steps or operations. It should be understood that the process 300 may be performed in various sequences and / or simultaneously, and is not limited to an execution sequence shown in FIG. 3. The process 300 includes the following steps.

[0147] Step 301: Obtain delay information of the audio content provider end.

[0148] The delay information indicates a delay status of the audio content provider end, and may include the following cases.

[0149] In a first case, the delay information includes a delay sent by the audio content provider end.

[0150] After determining the delay caused by a transmission link, rendering, and the like of the audio content provider end, the audio content provider end may package the delay into first information to be sent to the audio signal play end, and send the first information to the audio signal play end. In this way, the audio signal play end can directly extract the delay of the audio content provider end from the first information that comes from the audio content provider end.

[0151] In a second case, the delay information includes a second historical head movement position sent by the audio content provider end.

[0152] The audio signal play end (for example, a headset) may periodically detect, through a sensor (for example, a gyroscope or a gravity sensor), a current position (referred to as a measured head movement position in this specification) of a head of a user wearing the headset. In this way, the audio signal play end can send the measured head movement position obtained through measurement to the audio content provider end.

[0153] After receiving the measured head movement position, the audio content provider end does not process the measured head movement position, but still performs related processing on an audio frame according to a specified process. When first information needs to be sent to the audio signal play end after the processing is completed, the measured head movement position is carried in the first information. In this case, because a period of time has elapsed, the measured head movement position becomes a historical head movement position (referred to as the second historical head movement position in this specification). It should be understood that, in a mathematical sense, the measured head movement position is the same as the second historical head movement position.

[0154] Based on this, the audio signal play end may obtain, based on time at which the second historical head movement position is received and time at which the measured head movement position is sent, a time difference between the time at which the second historical head movement position is received and the time at which the measured head movement position is sent (in other words, calculate a time difference between receiving time and sending time of a same head movement position), to indirectly obtain the delay of the audio content provider end.

[0155] In a third case, the delay information includes a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end.

[0156] For the second historical head movement position, refer to the foregoing descriptions of the second case. Details are not described herein again.

[0157] The first uncertainty coefficient is an important parameter in a Kalman filter, and helps improve accuracy of a prediction result.

[0158] This case is related to an algorithm for obtaining a predicted head movement position in step 302. For specific content, refer to the following descriptions.

[0159] Step 302: Obtain a predicted head movement position based on the delay information.

[0160] As described above, to enable a user to enjoy real and immersive auditory experience through the audio signal play end (for example, a headset), sound sources in different formats, such as stereo, multi-channel surround sound, and a specially produced multi-object sound source, may be rendered into a binaural signal by using a spatial audio technology. In the real world, when a head of a user rotates or moves, an absolute position of a sound source does not change, but a relative direction between the sound source and the head changes. For example, a guitar is being played in front of a user, and if the user turns to the right, sound of the guitar is correspondingly shifted to the left of the user. For another example, there is a guitar on a left side of a stage, and there is a saxophone on a right side. When a user moves to a side of the stage, sound of the guitar and sound of the saxophone overlap, and come from a same direction. It can be learned that the so-called real and immersive auditory experience is the auditory experience with the guitar or the stage performance described above. With a change in a position of a head of a user (including but not limited to head displacements in up, down, left, and right directions, head rotation, and the like), sound from a same audio source is rendered into different auditory effect in ears of the user. Such rendering is a rendering method with head movement effect. To be specific, head tracking (Head Tracking) is performed through a sensor in a headset, and then a corresponding rotation or displacement change is incorporated in spatial rendering of a sound source.

[0161] In addition, although the audio content provider end may provide high computing power to implement the foregoing rendering algorithm with head movement effect, due to characteristics of the audio content provider end, a long delay occurs, causing disorder of important spatial information such as an orientation. For example, an acoustic image of an object is designed to be directly in front of space. If a head of a user rotates, the acoustic image first moves to be directly in front of a face of the user, and then is restored to be directly in front of space. A clear delay and "sense of damping" occur, affecting auditory experience.

[0162] To resolve the foregoing problem caused by the delay, the audio signal play end may predict a head movement position of a user, and then send the predicted head movement position to the audio content provider end, so that the audio content provider end can render an audio frame based on the predicted head movement position. In this way, an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position. In this case, when the rendered audio frame is sent to the audio signal play end, predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.

[0163] In this application, the audio signal play end may first obtain a first historical head movement position corresponding to the delay information, and then obtain a measured head movement position through the sensor, and then obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.

[0164] Based on the three cases of the delay information in step 301, the audio signal play end may obtain the first historical head movement position by using three methods. 1. When the delay information includes the delay sent by the audio content provider end (the first case), the first historical head movement position corresponding to the delay is extracted from a cache, where a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache. As described above, the audio signal play end may periodically detect, through the sensor, a measured head movement position of a user wearing the audio signal play end. For ease of subsequent use, a correspondence between a measured head movement position and detection time may be stored on the audio signal play end. In this way, when the foregoing delay is obtained, a current measured head movement position (referred to as the first historical head movement position in this specification) corresponding to the delay may be obtained by querying the correspondence. 2. When the delay information includes the second historical head movement position sent by the audio content provider end (the second case), the second historical head movement position is used as the first historical head movement position. It can be learned from the foregoing descriptions of the second case that the audio content provider end directly packages the second historical head movement position into the first information and sends the first information to the audio signal play end. In this way, the audio signal play end can directly obtain a current measured head movement position (the second historical head movement position) corresponding to the delay. Therefore, the audio signal play end can directly determine the second historical head movement position as the first historical head movement position. 3. When the delay information includes the second historical head movement position and the first uncertainty coefficient that are sent by the audio content provider end (the third case), the second historical head movement position is used as the first historical head movement position.

[0165] Refer to the foregoing descriptions in 2, in the third case of the delay information, the audio signal play end may alternatively directly determine the second historical head movement position as the first historical head movement position.

[0166] The audio signal play end may obtain a measured head movement position through the sensor, where the measured head movement position is a newly detected head movement position of the user.

[0167] In this application, obtaining the predicted head movement position based on the first historical head movement position and the measured head movement position may include the following two algorithms.

[0168] In a first algorithm, based on the foregoing 1 and 2, the audio signal play end may obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.

[0169] The difference between the first historical head movement position and the measured head movement position may be a distance between the first historical head movement position and the measured head movement position. For example, in a three-dimensional world, a head movement position may be represented by using coordinate axes x, y, and z. Therefore, the foregoing difference may be a distance between a coordinate value of the first historical head movement position and a coordinate value of the measured head movement position. In addition, a head movement position may alternatively be represented in another manner. This is not specifically limited in this application.

[0170] As described above, the audio signal play end may periodically detect, through the sensor, a measured head movement position of a user wearing the audio signal play end. Therefore, the audio signal play end may obtain a head movement change rate based on N measured head movement positions that are previously obtained, where the N measured head movement positions may include N measured head movement positions obtained through counting forward from a current measured head movement position.

[0171] The audio signal play end may predict a head movement change value based on the difference and the head movement change rate through cubic spline interpolation or by another means, and then add up the head movement change value and a current measured head movement position to obtain a predicted head movement position.

[0172] In a second algorithm, based on the foregoing 3, the audio signal play end obtains a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtains a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and inputs the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.

[0173] For a manner of obtaining the head movement change rate, refer to the foregoing descriptions.

[0174] The second uncertainty coefficient may be obtained based on the head movement change rate and the first uncertainty coefficient. The second uncertainty coefficient, the first historical head movement position, and the measured current head movement position are input to the Kalman filter. In the Kalman filter, a Kalman gain coefficient may be determined based on the second uncertainty coefficient and estimated uncertainty of a previous iteration, and the Kalman gain coefficient is used as a latest gain coefficient. In this way, the Kalman filter can output the predicted head movement position.

[0175] It should be noted that, in addition to the foregoing two methods, in this application, the predicted head movement position may alternatively be obtained based on the first historical head movement position and the measured head movement position by using another algorithm. This is not specifically limited herein.

[0176] Step 303: Send first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position.

[0177] The audio signal play end sends the first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position obtained in step 302, so that the audio content provider end can perform rendering with head movement effect on an audio frame based on the predicted head movement position. In this way, an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position. In this case, when the rendered audio frame is sent to the audio signal play end, predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.

[0178] In this application, a head movement position of a user is predicted to obtain a predicted head movement position, and after the predicted head movement position is sent to the audio content provider end, the audio content provider end may render an audio frame based on the predicted head movement position. In this way, an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position. In this case, when the rendered audio frame is sent to the audio signal play end, predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.

[0179] FIG. 4 is a flowchart of a process 400 of a spatial audio rendering method according to this application. As shown in FIG. 4, the process 400 may be applied to the communication system shown in FIG. 1A and FIG. 1B, and is jointly performed by an audio content provider end and an audio signal play end, to complete rendering of an audio source, and play audio to a user through the audio signal play end. The process 400 is described as a series of steps or operations. It should be understood that the process 400 may be performed in various sequences and / or simultaneously, and is not limited to an execution sequence shown in FIG. 4. The process 400 includes the following steps.

[0180] Step 401: When a first rendering mode is used, the audio content provider end obtains a current frame.

[0181] The first rendering mode may be a non-low-delay link mode. In this mode, the audio content provider end may perform rendering with head movement effect on an audio frame based on a predicted head movement position. A user may perform an operation on an application (application, APP) installed on the audio content provider end, to choose whether to select the first rendering mode. Correspondingly, this application further includes a second rendering mode. In this mode, the audio content provider end performs rendering without head movement effect on an audio frame. For this process, refer to the following descriptions. Similarly, the user may perform an operation on the APP installed on the audio content provider end, to choose whether to select the second rendering mode. For the operation of the user, refer to the following embodiments.

[0182] The current frame may be a frame of the audio source, and usually, may be an audio frame that is currently being processed by the audio content provider end.

[0183] Step 402: When an audio format of an audio source is a non-stereo format, the audio content provider end obtains a predicted head movement position.

[0184] The audio format of the audio source includes a stereo format or the non-stereo format, where the non-stereo format may include but is not limited to audio in a multi-channel or multi-object format such as 5.1, 7.1, or 3DA.

[0185] In this application, the audio content provider end may perform high-computing-power rendering on the audio in the multi-channel or multi-object format (non-stereo format) such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized. Based on this, when the audio format of the audio source is the non-stereo format, the audio content provider end may perform rendering with head movement effect on an audio frame. Therefore, the predicted head movement position needs to be obtained.

[0186] The audio content provider end may obtain the predicted head movement position from rendering information that is previously received from the audio signal play end. Optionally, the rendering information may be latest received rendering information. In this way, accuracy of the predicted head movement position can be improved. For obtaining of the predicted head movement position, refer to the embodiment shown in FIG. 3. Details are not described herein.

[0187] Step 403: The audio content provider end renders the current frame based on the predicted head movement position to obtain a first binaural signal.

[0188] In this application, the audio content provider end may separately perform rendering with head movement effect on a direct sound part, an early reflected sound (Early Reflection, ER) part, and a late reverberant sound (Late Reverb, LR) part in the current frame based on the predicted head movement position, to obtain the first binaural signal.

[0189] Optionally, direct sound with head movement effect and reverberation (obtained by mixing ER and LR) with head movement effect are separately rendered by an audio effector and then mixed, and an output result obtained through mixing is rendered by the audio effector again and then transmitted in a form of the first binaural signal. The audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.

[0190] In a branch of step 403, a first audio stream may include the first binaural signal.

[0191] Step 404: When an audio format of an audio source is a stereo format, the audio content provider end does not perform spatial rendering on the audio source.

[0192] In this application, the audio content provider end does not need to render an audio source in a stereo format, and the audio signal play end performs rendering with head movement effect on the audio source. Therefore, the audio content provider end may perform high-computing-power rendering on the audio in the multi-channel or multi-object format (non-stereo format) such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized. Based on this, when the audio format of the audio source is the non-stereo format, the audio content provider end may perform rendering with head movement effect on an audio frame. Therefore, the audio content provider end only needs to package the audio source into the first audio stream.

[0193] Step 405: The audio content provider end sends first information to the audio signal play end, where the first information includes the first audio stream, the audio format of the audio source, and delay information.

[0194] As described above, when the audio format indicates that the audio source is in the non-stereo format, the first audio stream includes the first binaural signal rendered by the audio content provider end; or when the audio format indicates that the audio source is in the stereo format, the first audio stream includes the audio source.

[0195] For the delay information, refer to the descriptions of the embodiment shown in FIG. 3. Details are not described herein again.

[0196] Step 406: When the audio format indicates that the audio source is in the stereo format, the audio signal play end obtains a measured head movement position through a sensor.

[0197] For an audio source in a stereo format, the audio signal play end performs rendering with a head movement position on the audio source. Because an IMU data transmission link of the audio signal play end has a low delay, the audio signal play end can obtain a current head movement position (referred to as the measured head movement position in this specification) in real time through the sensor.

[0198] Step 407: The audio signal play end renders the current frame based on the measured head movement position to obtain a second binaural signal.

[0199] The audio signal play end may perform rendering with head movement effect on the direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on the ER and the LR in the current frame, to obtain the second binaural signal.

[0200] Optionally, direct sound with head movement effect and reverberation (obtained by mixing ER and LR) without head movement effect are separately rendered by an audio effector and then mixed, and an output result obtained through mixing is rendered by the audio effector again and then mixed again, to obtain the second binaural signal. The audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.

[0201] Step 408: The audio signal play end plays audio based on a target binaural signal, where the target binaural signal includes the first binaural signal or the second binaural signal.

[0202] The audio signal play end obtains the target binaural signal through processing in the foregoing steps, where the target binaural signal may be the first binaural signal (the audio source is in the non-stereo format) or the second binaural signal (the audio source is in the stereo format).

[0203] In a possible implementation, when the second rendering mode is used, the audio content provider end obtains a current frame, performs rendering without head movement effect on the current frame to obtain a second binaural signal (the second binaural signal is different from the foregoing second binaural signal, and is referred to as a third binaural signal below for differentiation), and sends second information to the audio signal play end, where the second information includes a second audio stream, and the second audio stream includes the third binaural signal. The audio signal play end obtains a measured head movement position through the sensor, and renders a current frame (the current frame is an audio frame obtained by the audio content provider end by performing rendering without head movement effect) based on the measured head movement position to obtain a fourth binaural signal. The audio signal play end plays audio based on a target binaural signal, where the target binaural signal includes the fourth binaural signal.

[0204] In the second rendering mode, the audio content provider end performs rendering without head movement effect on an audio frame, so that computing power advantages of the audio content provider end can be fully utilized. A rendering delay can be shortened without head movement effect. Instead, the audio signal play end performs low-computing-power rendering with head movement effect on the audio frame. Compared with the first rendering mode, this achieves a lower delay, but rendering effect of head movement effect is poor.

[0205] The following describes in detail the technical solutions of this application by using several specific embodiments. In the following embodiments, an audio content provider end may be a mobile phone, and an audio signal play end may be a headset. An audio source may alternatively be a sound source or an input signal. Rendering may alternatively be (binaural) spatial audio rendering. A head movement position may alternatively be IMU data. A rendering algorithm deployed on the mobile phone may alternatively be a first rendering part of a binaural spatial audio rendering algorithm. A rendering algorithm deployed on the headset may alternatively be a second rendering part of the binaural spatial audio rendering algorithm. An audio format may alternatively be a format flag. A delay may alternatively be a dynamic delay of a spatial audio link.Embodiment 1

[0206] FIG. 5 is a diagram of an overall process of a spatial audio rendering method according to this application. As shown in FIG. 5, this application is applied to a spatial audio processing scenario in which a mobile phone and a headset perform joint rendering in a first rendering mode. A joint rendering policy in this embodiment may be designed as follows: If the mobile phone determines that a sound source is stereo, audio is sent to the headset for binaural spatial audio rendering; or if the mobile phone determines that a sound source is non-stereo, audio is rendered on the mobile phone, and after the rendering is completed, rendered audio is sent to the headset, and the headset directly outputs the rendered audio to a speaker.

[0207] FIG. 6 is a diagram of a specific process of a spatial audio rendering method according to this application. As shown in FIG. 6, the process in this embodiment includes the following steps.

[0208] S1: A mobile phone receives an input signal, and detects a format of the input signal, and a headset delivers predicted IMU data to a first rendering part of a binaural spatial audio rendering algorithm of the mobile phone.

[0209] The first rendering part of the binaural spatial audio rendering algorithm includes direct sound (Direct) rendering, early reflection (ER) rendering, and late reverberation (LR) rendering. Head movement effect processing is performed, by using received IMU data, on a reverberation part obtained by mixing a direct sound part with ER and LR. Direct sound with head movement effect and reverberation with head movement effect are separately rendered by an audio effector, and then rendered audio enters a first mixing module (Mixer 1). An output result of the first mixing module is rendered by an audio effector again, and then rendered audio is transmitted to Bluetooth in a dual-channel form. The audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.

[0210] S2: If the format of the input signal is stereo, the mobile phone directly transmits the input signal and a format flag to the first mixing module.

[0211] S3: If the format of the input signal is not stereo, render the input signal based on the first part of the binaural spatial audio rendering algorithm by using the predicted IMU data, and transmit a rendered audio stream and a format flag to the first mixing module.

[0212] S4: Package the audio stream and the format flag of the first mixing module and a dynamic delay of a spatial audio link of the mobile phone, and transmit a package to the headset through Bluetooth.

[0213] S5: The headset stores measured IMU data, selects corresponding historical measured IMU data and measured IMU data of the headset based on the delay reported by the mobile phone, transmits the historical measured IMU data and the measured IMU data of the headset to a prediction module, and transmits the measured IMU data of the headset to a second rendering part of the binaural spatial audio rendering algorithm. The prediction module generates predicted IMU data based on the historical measured IMU data and the measured IMU data of the headset, and delivers the predicted IMU data to the mobile phone again with reference to S1.

[0214] FIG. 7 is a schematic flowchart of an algorithm for predicting a head movement position. As shown in FIG. 7, the process includes the following steps.

[0215] S5.1: A headset stores historical IMU data in a cache, extracts corresponding historical IMU data from the historical cache based on a received mobile phone delay, calculates a difference between measured IMU data and the historical IMU data, and sends the difference to a prediction unit.

[0216] S5.2: Calculate a change rate of previous N pieces of measured IMU data (for example, including a yaw (yaw), a pitch (pitch), a roll (roll), or four elements) of the headset, and transmit a change coefficient to the prediction unit.

[0217] S5.3: Predict a head movement change value through cubic spline interpolation or by another means based on the change rate and the difference calculated in S5.1.

[0218] S5.4: Add up the head movement change value and the measured IMU data of the headset to generate a predicted head movement position.

[0219] S5.5: Deliver the predicted head movement position to the mobile phone.

[0220] S6: The headset detects a format of a sound source, and if the format of the sound source is stereo, renders an input signal based on a second part of a binaural spatial audio rendering algorithm by using the measured IMU data of the headset, and sends a rendered audio stream to a second mixing module.

[0221] The second rendering part of the binaural spatial audio rendering algorithm includes direct sound (Direct) rendering and basic low-computing-power reverberation rendering. Head movement effect processing is performed on a direct sound part by using received IMU data. Direct sound with head movement effect and basic reverberation without head movement effect are separately rendered by an audio effector and then mixed. An overall output result obtained through mixing is rendered by an audio effector again, and then rendered audio enters the second mixing module (Mixer 2). The audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.

[0222] S7: If the format of the sound source is not stereo, the headset directly transmits an input signal to a second mixing module.

[0223] S8: Send the second mixing module to a headset speaker for playing.

[0224] In Embodiment 1, high computing power of the mobile phone and a low delay of the headset are fully utilized, and the headset may run independently. In view of a problem that the mobile phone has a long link and a long delay, an end-to-end delay prediction algorithm is designed to greatly reduce a head movement delay (by more than 50% as predicted).Embodiment 2

[0225] Compared with Embodiment 1, a difference in Embodiment 2 mainly focuses on content transmitted through Bluetooth in S1 and S4 and an algorithm for predicting a head movement position in S5.

[0226] In this embodiment, S1 is changed as follows: A mobile phone receives an input signal, and detects a format of the input signal, and a headset delivers IMU data of the headset and predicted IMU data to a first rendering part of a binaural spatial audio rendering algorithm of the mobile phone.

[0227] S4 is changed as follows: Package the IMU data, the audio stream, and the format flag of the first mixing module, and transmit a package to the headset through Bluetooth.

[0228] S5 is changed as follows: The headset transmits, to a prediction module, historical IMU data reported by the mobile phone, transmits measured IMU data to the prediction module, and transmits the measured IMU data to a second rendering part of the binaural spatial audio rendering algorithm. The prediction module generates predicted IMU data based on the historical IMU data and the measured IMU data, and delivers the measured IMU data and the predicted IMU data to the mobile phone with reference to S1.

[0229] FIG. 8a and FIG. 8b are a schematic flowchart of an algorithm for predicting a head movement position. As shown in FIG. 8a and FIG. 8b, the process includes the following steps.

[0230] S5.1: Calculate a difference between measured IMU data and historical IMU data reported by a mobile phone, and send the difference to a prediction algorithm.

[0231] S5.2: Calculate a change rate of previous N pieces of measured IMU data (for example, including a yaw, a pitch, a roll, or four elements) of a headset, and transmit a change coefficient to a prediction unit.

[0232] S5.3: Predict a head movement change value through cubic spline interpolation or by another means based on the change rate and the difference calculated in S5.1.

[0233] S5.4: Add up the head movement change value and the measured IMU data of the headset to generate a predicted head movement position.

[0234] S5.5: Deliver a measured head movement position and the predicted head movement position to the mobile phone.Embodiment 3

[0235] Compared with Embodiment 1, a difference in Embodiment 3 mainly focuses on content transmitted through Bluetooth in S1 and S4 and a prediction algorithm in S5.

[0236] In Embodiment 3, S1 is changed as follows: A mobile phone receives an input signal, and detects a format of the input signal, and a headset delivers an uncertainty coefficient Rn and predicted IMU data to a first rendering part of a binaural spatial audio rendering algorithm of the mobile phone.

[0237] S4 is changed as follows: Package the IMU data, the uncertainty coefficient Rn, the audio stream, and the format flag of the first mixing module, and transmit a package to the headset through Bluetooth.

[0238] S5 is changed as follows: The headset transmits, to a prediction module, IMU data reported by the mobile phone and the uncertainty coefficient Rn, transmits measured IMU data to the prediction module, and transmits the measured IMU data to a second rendering part of the binaural spatial audio rendering algorithm. The prediction module generates a predicted value of IMU data based on the IMU data reported by the mobile phone, the uncertainty coefficient Rn, and IMU data of the headset at this time, and delivers the measured IMU data and the predicted IMU data to the mobile phone with reference to S1.

[0239] FIG. 9a and FIG. 9b are a schematic flowchart of an algorithm for predicting a head movement position. As shown in FIG. 9a and FIG. 9b, the process includes the following steps.

[0240] S5.1: Calculate a change rate of previous N pieces of measured IMU data (for example, including a yaw, a pitch, a roll, or four elements) of a headset, determine a measurement uncertainty coefficient Rn+1 based on a historical measurement uncertainty coefficient Rn, send Rn+1 to a Kalman filter, and send, to the Kalman filter, a historical head movement position reported by a mobile phone and a measured head movement position measured by a sensor of the headset.

[0241] S5.2: Determine a Kalman gain coefficient Kn+1 based on the uncertainty coefficient Rn+1 and estimated uncertainty Pn+1 of a previous iteration, and send the gain coefficient to a current status update module.

[0242] S5.3: Update a predicted head movement position based on the following formula: Xn+1 = Xn + Kn+1 × (Yn - Xn); output a predicted head movement position; and update the estimated uncertainty based on the following formula: Pn+1 = (1 - Kn) × Pn, where Yn is a current actual head movement position, and Xn is a historical head movement position.Embodiment 4

[0243] FIG. 10 is a schematic framework flowchart according to this application. As shown in FIG. 10, this embodiment provides a compatibility solution for a spatial audio processing scenario in which a mobile phone and a headset perform joint rendering in a non-low-delay case (a first rendering mode) and in a low-delay case (a second rendering mode). In the second rendering mode, the headset no longer performs reverberation rendering, and cannot independently perform spatial audio rendering without the mobile phone, but this mode is characterized by a low requirement for computing power and a simple structure.

[0244] S1: Add a low-delay mode to a user interface (User Interface, UI) of an APP supported by spatial audio, and deliver, based on selection of a user, a flag indicating whether the low-delay mode is used.

[0245] S2: If the low-delay mode is used, the mobile phone performs preprocessing, including downmixing and reverberation processing, on all types of audio signals.

[0246] S3: Transmit a preprocessed signal to the headset for direct sound rendering in a binaural spatial audio rendering algorithm, and output a rendered signal to a headset speaker.

[0247] S4: If a non-low-delay mode is used, perform rendering with head movement effect on the mobile phone, where this part of rendering includes processing of non-stereo direct sound and reverberation processing of signals in all formats; and transmit a rendered signal to the headset.

[0248] S5: The headset determines an input format; and performs direct sound rendering on a stereo sound source, and outputs a rendered sound source to the headset speaker; or directly outputs a non-stereo sound source to the headset speaker.

[0249] In this embodiment, a process in the low-delay mode is shown in FIG. 11 (FIG. 11 is a diagram of a process in a low-delay mode according to this application). After the user selects the low-delay mode, the mobile phone preprocesses a received audio source in a format of stereo, 5.1, 7.1, 3DA, or the like. The preprocessing includes downmixing an audio source in a multi-channel format such as 5.1 or 7.1 or in a multi-object format such as 3DA, to unify the audio source into a stereo format. Then the audio source is transmitted to a first rendering part of the binaural spatial audio rendering algorithm for basic reverberation rendering without head movement effect and audio effector processing. A reverberation part is mixed with downmixed input obtained through an audio effector, and then mixed audio is transmitted to the headset through Bluetooth.

[0250] The headset receives stereo with reverberation effect that is transmitted by the mobile phone, and performs direct sound rendering by using a second part of the binaural spatial audio rendering algorithm based on IMU head movement data input by a headset sensor. The second part of the binaural spatial audio rendering algorithm includes a direct sound orientation rendering module and an audio effector module.

[0251] In this embodiment, a process in the non-low-delay mode is shown in FIG. 12 (FIG. 12 is a diagram of a process in a non-low-delay mode according to this application). After receiving input sources in various formats, the mobile phone performs direct sound rendering on a non-stereo sound source and reverberation rendering on sound sources in all formats by using a first part of the binaural spatial audio rendering algorithm. Both the direct sound rendering and the reverberation rendering in this part include head movement effect, and an implementation solution in which ER and LR are separately rendered and then added up is used for reverberation. For a needed predicted head movement position, a head movement delay may be optimized with reference to the prediction algorithm in Embodiment 1 to Embodiment 3. In FIG. 12, the prediction algorithm in Embodiment 2 is used as an example for illustration.

[0252] The mobile phone processes a direct sound rendering result or stereo raw input and a reverberation rendering result through a corresponding audio effector, mixes audio into a stereo format, and transmits the audio to the headset through Bluetooth.

[0253] The headset determines a format of raw input; and if the raw input is in a non-stereo format, directly outputs the raw input to the speaker; or if the raw input is in a stereo format, performs direct sound rendering and audio effector processing on the raw input by using a second part of the binaural spatial audio rendering algorithm based on IMU head movement data input by a headset sensor, and then outputs processed audio to the speaker for playing.

[0254] In this embodiment, a UI design is added, and a joint rendering solution is selected based on the user's concern about a head movement delay and spatial audio rendering quality.

[0255] Joint rendering by the mobile phone and the headset: A rendering algorithm is divided into two parts: the mobile phone and the headset. The mobile phone renders a multi-channel or multi-object part such as 5.1, 7.1, or 3DA, and performs reverberation effect rendering, to fully utilize computing power advantages of the mobile phone. The headset performs direct sound rendering on stereo, to fully utilize advantages of a short IMU data transmission link and a low delay.

[0256] Head movement position prediction: To resolve a problem that the mobile phone has a long head movement delay, a head movement delay prediction algorithm is designed. The prediction algorithm receives, on the headset, a link delay of the mobile phone, a delay of the headset, and a current IMU value, and generates predicted IMU data. Measured IMU data of the headset is transmitted to the second part of the binaural spatial audio rendering algorithm, and the predicted IMU data is transmitted to the first part of the binaural spatial audio rendering algorithm. In this solution, head movement position prediction is implemented, and impact of a link delay can be reduced by more than 1 / 2.

[0257] FIG. 13 is a diagram of an example structure of a spatial audio rendering apparatus 1300 according to this application. As shown in FIG. 13, the spatial audio rendering apparatus 1300 in this embodiment may be used at an audio signal play end. The spatial audio rendering apparatus 1300 may include an obtaining module 1301, a prediction module 1302, a sending module 1303, a receiving module 1304, a rendering module 1305, and a playing module 1306.

[0258] The obtaining module 1301 is configured to: obtain delay information of an audio content provider end. The prediction module 1302 is configured to: obtain a predicted head movement position based on the delay information. The sending module 1303 is configured to: send first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position.

[0259] In a possible implementation, the prediction module 1302 is specifically configured to: obtain a first historical head movement position corresponding to the delay information; obtain a measured head movement position through a sensor; and obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.

[0260] In a possible implementation, the delay information includes a delay sent by the audio content provider end; and the prediction module 1302 is specifically configured to: extract, from a cache, the first historical head movement position corresponding to the delay, where a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache.

[0261] In a possible implementation, the delay information includes a second historical head movement position sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, and the second rendering information is sent earlier than the first rendering information; and the prediction module 1302 is specifically configured to: use the second historical head movement position as the first historical head movement position.

[0262] In a possible implementation, the prediction module 1302 is specifically configured to: obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.

[0263] In a possible implementation, the delay information includes a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, the first uncertainty coefficient comes from the second rendering information, and the second rendering information is sent earlier than the first rendering information; and the prediction module 1302 is specifically configured to: use the second historical head movement position as the first historical head movement position.

[0264] In a possible implementation, the prediction module 1302 is specifically configured to: obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtain a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and input the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.

[0265] In a possible implementation, the receiving module 1304 is configured to: receive first information sent by the audio content provider end, where the first information includes an audio format, a first audio stream, and the delay information, the audio format indicates that an audio source is in a stereo format or a non-stereo format, and when the audio format indicates that the audio source is in the non-stereo format, the first audio stream includes a first binaural signal rendered by the audio content provider end, or when the audio format indicates that the audio source is in the stereo format, the first audio stream includes the audio source; the rendering module 1305 is configured to: when the audio format indicates that the audio source is in the stereo format, obtain a measured head movement position through the sensor; and render a current frame based on the measured head movement position to obtain a second binaural signal, where the current frame is a frame of the audio source; and the playing module 1306 is configured to: play audio based on a target binaural signal, where the target binaural signal includes the first binaural signal or the second binaural signal.

[0266] In a possible implementation, the rendering module 1305 is specifically configured to: perform rendering with head movement effect on a direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on early reflected sound ER and late reverberant sound LR in the current frame, to obtain the second binaural signal.

[0267] The apparatus in this embodiment may be configured to perform the technical solution performed by the client in the method embodiment shown in FIG. 3 or FIG. 4. An implementation principle and technical effect of the apparatus are similar. Details are not described herein again.

[0268] FIG. 14 is a diagram of an example structure of a spatial audio rendering apparatus 1400 according to this application. As shown in FIG. 14, the spatial audio rendering apparatus 1400 in this embodiment may be used at an audio content provider end. The spatial audio rendering apparatus 1400 may include an obtaining module 1401, a rendering module 1402, and a sending module 1403.

[0269] The obtaining module 1401 is configured to: obtain a current frame when a first rendering mode is used, where the current frame is a frame of an audio source; and obtain a predicted head movement position when an audio format of the audio source is a non-stereo format. The rendering module 1402 is configured to: render the current frame based on the predicted head movement position to obtain a first binaural signal. The sending module 1403 is configured to: send first information to an audio signal play end, where the first information includes a first audio stream, the audio format of the audio source, and delay information, and the first audio stream includes the first binaural signal.

[0270] In a possible implementation, the rendering module 1402 is specifically configured to: separately perform rendering with head movement effect on a direct sound part, an early reflected sound ER part, and a late reverberant sound LR part in the current frame based on the predicted head movement position, to obtain the first binaural signal.

[0271] In a possible implementation, the obtaining module 1401 is specifically configured to: obtain the predicted head movement position from rendering information sent by the audio signal play end.

[0272] In a possible implementation, the delay information includes a delay.

[0273] In a possible implementation, the rendering information further includes a measured head movement position in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information includes the measured head movement position.

[0274] In a possible implementation, the rendering information further includes a measured head movement position and a first uncertainty coefficient in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information includes the measured head movement position and the first uncertainty coefficient.

[0275] In a possible implementation, when an audio format of the audio source is a stereo format, the first audio stream includes the audio source.

[0276] In a possible implementation, the obtaining module 1401 is further configured to: obtain the current frame when a second rendering mode is used; the rendering module 1402 is further configured to: perform rendering without head movement effect on the current frame to obtain a second binaural signal; and the sending module 1403 is further configured to: send second information to the audio signal play end, where the second information includes a second audio stream, and the second audio stream includes the second binaural signal.

[0277] The apparatus in this embodiment may be configured to perform the technical solution performed by the client in the method embodiment shown in FIG. 3 or FIG. 4. An implementation principle and technical effect of the apparatus are similar. Details are not described herein again.

[0278] During implementation, the steps in the foregoing method embodiments may be performed by a hardware integrated logic circuit in a processor or by using instructions in a form of software. The processor may be a general-purpose processor, a digital signal processor (digital signal processor, DSP), an application-specific integrated circuit (application-specific integrated circuit, ASIC), a field programmable gate array (field programmable gate array, FPGA) or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor or the like. The steps of the methods disclosed in embodiments of this application may be directly performed by a hardware encoding processor, or performed by a combination of hardware and a software module in an encoding processor. The software module may be located in a mature storage medium in the art, for example, a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register. The storage medium is located in a memory, and the processor reads information in the memory and performs the steps of the foregoing methods based on hardware of the processor.

[0279] The memory in the foregoing embodiments may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory. The non-volatile memory may be a read-only memory (read-only memory, ROM), a programmable read-only memory (programmable ROM, PROM), an erasable programmable read-only memory (erasable PROM, EPROM), an electrically erasable programmable read-only memory (electrically EPROM, EEPROM), or a flash memory. The volatile memory may be a random access memory (random access memory, RAM), and serves as an external cache. By way of example but not limitative description, RAMs in many forms may be used, for example, a static random access memory (static RAM, SRAM), a dynamic random access memory (dynamic RAM, DRAM), a synchronous dynamic random access memory (synchronous DRAM, SDRAM), a double data rate synchronous dynamic random access memory (double data rate SDRAM, DDR SDRAM), an enhanced synchronous dynamic random access memory (enhanced SDRAM, ESDRAM), a synchronous link dynamic random access memory (synchlink DRAM, SLDRAM), and a direct rambus random access memory (direct rambus RAM, DR RAM). It should be noted that the memory of the systems and methods described in this specification is intended to include but is not limited to these memories and any other appropriate type of memory.

[0280] A person of ordinary skill in the art may be aware that units and algorithm steps in examples described with reference to embodiments disclosed in this specification can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed by hardware or software depends on particular applications and design constraints of technical solutions. A person skilled in the art may use different methods to implement the described functions for each particular application, but it should not be considered that the implementation goes beyond the scope of this application.

[0281] It can be clearly understood by a person skilled in the art that, for ease and brevity of description, for detailed working processes of the foregoing system, apparatus, and unit, reference may be made to corresponding processes in the foregoing method embodiments. Details are not described herein again.

[0282] In the several embodiments provided in this application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described apparatus embodiments are merely examples. For example, division into the units is merely logical function division and may be other division during actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the shown or discussed mutual couplings or direct couplings or communication connections may be implemented through some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electrical, mechanical, or other forms.

[0283] The units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, to be specific, may be located in one place, or may be distributed on a plurality of network units. Some or all of the units may be selected according to actual requirements to achieve the objectives of the solutions of embodiments.

[0284] In addition, functional units in embodiments of this application may be integrated into one processing unit, or each unit may exist alone physically, or two or more units may be integrated into one unit.

[0285] When the functions are implemented in a form of a software functional unit and sold or used as an independent product, the functions may be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions in this application essentially, or the part contributing to the conventional technology, or some of the technical solutions may be implemented in a form of a software product. The computer software product is stored in a storage medium, and includes several instructions for instructing a computer device (a personal computer, a server, a network device, or the like) to perform all or some of the steps of the methods in embodiments of this application. The storage medium includes any medium that can store program code, for example, a USB flash drive, a removable hard disk drive, a read-only memory (read-only memory, ROM), a random access memory (random access memory, RAM), a magnetic disk, or a compact disc.

[0286] The foregoing descriptions are merely specific implementations of this application, but are not intended to limit the protection scope of this application. Any variation or replacement readily figured out by a person skilled in the art within the technical scope disclosed in this application shall fall within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A spatial audio rendering method, comprising: obtaining delay information of an audio content provider end; obtaining a predicted head movement position based on the delay information; and sending first rendering information to the audio content provider end, wherein the first rendering information comprises the predicted head movement position.

2. The method according to claim 1, wherein obtaining the predicted head movement position based on the delay information comprises: obtaining a first historical head movement position corresponding to the delay information; obtaining a measured head movement position through a sensor; and obtaining the predicted head movement position based on the first historical head movement position and the measured head movement position.

3. The method according to claim 2, wherein the delay information comprises a delay sent by the audio content provider end; and obtaining the first historical head movement position corresponding to the delay information comprises: extracting, from a cache, the first historical head movement position corresponding to the delay, wherein a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache.

4. The method according to claim 2, wherein the delay information comprises a second historical head movement position sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, and the second rendering information is sent earlier than the first rendering information; and obtaining the first historical head movement position corresponding to the delay information comprises: using the second historical head movement position as the first historical head movement position.

5. The method according to any one of claims 2 to 4, wherein obtaining the predicted head movement position based on the first historical head movement position and the measured head movement position comprises: obtaining a difference between the first historical head movement position and the measured head movement position; obtaining a head movement change rate, wherein the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtaining the predicted head movement position based on the difference and the head movement change rate.

6. The method according to claim 2, wherein the delay information comprises a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, the first uncertainty coefficient comes from the second rendering information, and the second rendering information is sent earlier than the first rendering information; and obtaining the first historical head movement position corresponding to the delay information comprises: using the second historical head movement position as the first historical head movement position.

7. The method according to claim 6, wherein obtaining the predicted head movement position based on the first historical head movement position and the measured head movement position comprises: obtaining a head movement change rate, wherein the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtaining a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and inputting the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.

8. The method according to any one of claims 1 to 7, further comprising: receiving first information sent by the audio content provider end, wherein the first information comprises an audio format, a first audio stream, and the delay information, the audio format indicates that an audio source is in a stereo format or a non-stereo format, and when the audio format indicates that the audio source is in the non-stereo format, the first audio stream comprises a first binaural signal rendered by the audio content provider end, or when the audio format indicates that the audio source is in the stereo format, the first audio stream comprises the audio source; when the audio format indicates that the audio source is in the stereo format, obtaining a measured head movement position through the sensor; rendering a current frame based on the measured head movement position to obtain a second binaural signal, wherein the current frame is a frame of the audio source; and playing audio based on a target binaural signal, wherein the target binaural signal comprises the first binaural signal or the second binaural signal.

9. The method according to claim 8, wherein rendering the current frame based on the measured head movement position to obtain the second binaural signal comprises: performing rendering with head movement effect on a direct sound part in the current frame based on the measured head movement position, and performing rendering without head movement effect on early reflected sound ER and late reverberant sound LR in the current frame, to obtain the second binaural signal.

10. A spatial audio rendering method, comprising: obtaining a current frame when a first rendering mode is used, wherein the current frame is a frame of an audio source; obtaining a predicted head movement position when an audio format of the audio source is a non-stereo format; rendering the current frame based on the predicted head movement position to obtain a first binaural signal; and sending first information to an audio signal play end, wherein the first information comprises a first audio stream, the audio format of the audio source, and delay information, and the first audio stream comprises the first binaural signal.

11. The method according to claim 10, wherein rendering the current frame based on the predicted head movement position to obtain the first binaural signal comprises: separately performing rendering with head movement effect on a direct sound part, an early reflected sound ER part, and a late reverberant sound LR part in the current frame based on the predicted head movement position, to obtain the first binaural signal.

12. The method according to claim 10 or 11, wherein obtaining the predicted head movement position comprises: obtaining the predicted head movement position from rendering information sent by the audio signal play end.

13. The method according to claim 12, wherein the delay information comprises a delay.

14. The method according to claim 12, wherein the rendering information further comprises a measured head movement position in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information comprises the measured head movement position.

15. The method according to claim 12, wherein the rendering information further comprises a measured head movement position and a first uncertainty coefficient in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information comprises the measured head movement position and the first uncertainty coefficient.

16. The method according to any one of claims 10 to 15, wherein when an audio format of the audio source is a stereo format, the first audio stream comprises the audio source.

17. The method according to any one of claims 10 to 16, further comprising: obtaining the current frame when a second rendering mode is used; performing rendering without head movement effect on the current frame to obtain a second binaural signal; and sending second information to the audio signal play end, wherein the second information comprises a second audio stream, and the second audio stream comprises the second binaural signal.

18. A spatial audio rendering apparatus, comprising: an obtaining module, configured to: obtain delay information of an audio content provider end; a prediction module, configured to: obtain a predicted head movement position based on the delay information; and a sending module, configured to: send first rendering information to the audio content provider end, wherein the first rendering information comprises the predicted head movement position.

19. The apparatus according to claim 18, wherein the prediction module is specifically configured to: obtain a first historical head movement position corresponding to the delay information; obtain a measured head movement position through a sensor; and obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.

20. The apparatus according to claim 19, wherein the delay information comprises a delay sent by the audio content provider end; and the prediction module is specifically configured to: extract, from a cache, the first historical head movement position corresponding to the delay, wherein a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache.

21. The apparatus according to claim 19, wherein the delay information comprises a second historical head movement position sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, and the second rendering information is sent earlier than the first rendering information; and the prediction module is specifically configured to: use the second historical head movement position as the first historical head movement position.

22. The apparatus according to any one of claims 19 to 21, wherein the prediction module is specifically configured to: obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, wherein the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.

23. The apparatus according to claim 19, wherein the delay information comprises a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, the first uncertainty coefficient comes from the second rendering information, and the second rendering information is sent earlier than the first rendering information; and the prediction module is specifically configured to: use the second historical head movement position as the first historical head movement position.

24. The apparatus according to claim 23, wherein the prediction module is specifically configured to: obtain a head movement change rate, wherein the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtain a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and input the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.

25. The apparatus according to any one of claims 18 to 24, further comprising: a receiving module, configured to: receive first information sent by the audio content provider end, wherein the first information comprises an audio format, a first audio stream, and the delay information, the audio format indicates that an audio source is in a stereo format or a non-stereo format, and when the audio format indicates that the audio source is in the non-stereo format, the first audio stream comprises a first binaural signal rendered by the audio content provider end, or when the audio format indicates that the audio source is in the stereo format, the first audio stream comprises the audio source; a rendering module, configured to: when the audio format indicates that the audio source is in the stereo format, obtain a measured head movement position through the sensor; and render a current frame based on the measured head movement position to obtain a second binaural signal, wherein the current frame is a frame of the audio source; and a playing module, configured to: play audio based on a target binaural signal, wherein the target binaural signal comprises the first binaural signal or the second binaural signal.

26. The apparatus according to claim 25, wherein the rendering module is specifically configured to: perform rendering with head movement effect on a direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on early reflected sound ER and late reverberant sound LR in the current frame, to obtain the second binaural signal.

27. A spatial audio rendering apparatus, comprising: an obtaining module, configured to: obtain a current frame when a first rendering mode is used, wherein the current frame is a frame of an audio source; and obtain a predicted head movement position when an audio format of the audio source is a non-stereo format; a rendering module, configured to: render the current frame based on the predicted head movement position to obtain a first binaural signal; and a sending module, configured to: send first information to an audio signal play end, wherein the first information comprises a first audio stream, the audio format of the audio source, and delay information, and the first audio stream comprises the first binaural signal.

28. The apparatus according to claim 27, wherein the rendering module is specifically configured to: separately perform rendering with head movement effect on a direct sound part, an early reflected sound ER part, and a late reverberant sound LR part in the current frame based on the predicted head movement position, to obtain the first binaural signal.

29. The apparatus according to claim 27 or 28, wherein the obtaining module is specifically configured to: obtain the predicted head movement position from rendering information sent by the audio signal play end.

30. The apparatus according to claim 29, wherein the delay information comprises a delay.

31. The apparatus according to claim 29, wherein the rendering information further comprises a measured head movement position in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information comprises the measured head movement position.

32. The apparatus according to claim 29, wherein the rendering information further comprises a measured head movement position and a first uncertainty coefficient in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information comprises the measured head movement position and the first uncertainty coefficient.

33. The apparatus according to any one of claims 27 to 32, wherein when an audio format of the audio source is a stereo format, the first audio stream comprises the audio source.

34. The apparatus according to any one of claims 27 to 33, wherein the obtaining module is further configured to: obtain the current frame when a second rendering mode is used; the rendering module is further configured to: perform rendering without head movement effect on the current frame to obtain a second binaural signal; and the sending module is further configured to: send second information to the audio signal play end, wherein the second information comprises a second audio stream, and the second audio stream comprises the second binaural signal.

35. An audio signal play device, comprising: one or more processors; and a memory, configured to store one or more programs, wherein when the one or more programs is / are executed by the one or more processors, the one or more processors is / are enabled to implement the method according to any one of claims 1 to 9.

36. An audio content providing device, comprising: one or more processors; and a memory, configured to store one or more programs, wherein when the one or more programs is / are executed by the one or more processors, the one or more processors is / are enabled to implement the method according to any one of claims 10 to 17.

37. A computer-readable storage medium, comprising a computer program, wherein when the computer program is executed on a computer, the computer is enabled to perform the method according to any one of claims 1 to 17.

38. A computer program product, wherein the computer program product comprises computer program code, and when the computer program code is run on a computer, the computer is enabled to perform the method according to any one of claims 1 to 17.