Spatial audio
By generating a mono differential audio signal and reconstructing the audio signal in the receiver device, the problem of spatial audio signal delay is solved, resulting in faster response and lower data transmission volume.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-02
- Publication Date
- 2026-03-24
AI Technical Summary
In existing technologies, the delay in spatial audio signals causes the direction of the sound source to appear lagging, making it impossible to respond promptly to changes in the user's viewpoint.
By generating a mono differential audio signal in the transmitter device, estimating the audio signal for the second viewpoint based on the difference between the first audio signal and the second audio signal, reducing the amount of data transmitted, and reconstructing it in the receiver device.
It reduces audio signal transmission delay, improves the real-time response of audio signals to changes in user viewpoint, and reduces data transmission volume.
Smart Images

Figure CN116074729B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to spatial audio. In particular, some embodiments relate to transmitting audio signals between a transmitter device and a receiver device. Background Technology
[0002] Spatial audio adapts to the user's changing viewpoint. For example, when listening with headphones, spatial audio can rotate as the user turns his or her head.
[0003] Humans are highly adept at detecting the direction of sound sources. They use varying viewpoints (such as head rotation) to improve this ability. For example, a user can rotate their head to position the desired sound at the center of their head, where their ability to detect the direction of the sound source is optimal. Furthermore, head rotation can be used to differentiate between sound sources in front of and behind the user. Using a left-to-right head rotation, sound sources in front move from right to left, while sound sources behind move from left to right.
[0004] In existing solutions, viewpoint data that tracks the user's viewpoint is transmitted or acquired by a transmitter that modifies the audio signal to rotate the audio scene based on the viewpoint data. The transmitter then performs low-bit-rate encoding on the audio signal and sends the encoded audio to a receiver for rendering. In some examples, the receiver may be a headset. The receiver decodes the audio and plays it back to the user. These steps can cause a delay in rendering the modified audio to the user in response to changes in their viewpoint. Typically, this delay can be several hundred milliseconds. Therefore, the direction of the sound source will appear laggy. It is desirable to reduce this delay. Summary of the Invention
[0005] According to various, but not necessarily all, embodiments, an apparatus for implementing adaptive playback is provided, comprising components configured to perform the following operations:
[0006] A first audio signal is obtained for at least the first and second channels from a first viewpoint;
[0007] A second audio signal is obtained for at least the first and second channels from a second viewpoint;
[0008] Based at least on the difference between the first audio signal and the second audio signal, a mono-channel differential audio signal is determined for the second viewpoint;
[0009] Depending on the mono differential audio signal and the first audio signal, it is possible to estimate both the first and second channels of the second audio signal for the second viewpoint.
[0010] According to some, but not necessarily all, examples, the component configured to determine a mono-channel differential audio signal based on the difference between a first audio signal and a second audio signal is configured to determine the difference between a reference channel of the first audio signal and the second audio signal, wherein the reference channel is the first channel, the second channel, or a synthesized channel based on the first channel and the second channel, and wherein the component configured to enable estimation of both the first channel and the second channel of the second audio signal depends on the mono-channel differential audio signal and the reference channel of the first audio signal to enable estimation.
[0011] According to some, but not necessarily all, examples, components configured to determine mono-channel difference audio signals for a second viewpoint are configured to determine the difference between a first audio signal and a second audio signal in the time domain.
[0012] According to some, but not necessarily all, examples, the device further includes a smoothing component configured to smooth the mono-channel differential audio signal in the frequency domain to obtain a smoothed mono-channel differential audio signal, and, depending on the smoothed mono-channel differential audio signal and the first audio signal, to enable estimation of at least a second audio signal.
[0013] Based on some, but not necessarily all, examples, the smoothing component is configured to replicate frequency bins within one or more different frequency bands.
[0014] According to some, but not necessarily all, examples, the smoothing component is configured to perform dynamic smoothing, wherein the dynamic smoothing of the mono-channel differential audio signal depends on the probability of viewpoint change from the first viewpoint to the second viewpoint, based at least on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint.
[0015] According to some, but not necessarily all, examples, the device includes components configured to perform the following operations:
[0016] When the second viewpoint is offset from the first viewpoint by a positive first angle and the third viewpoint is offset from the first viewpoint by a negative first angle, a mono-channel differential audio signal is obtained for the second viewpoint but not for the third viewpoint.
[0017] According to various, but not necessarily all, embodiments, a method for implementing adaptive playback is provided, comprising:
[0018] A first audio signal is obtained for at least the first and second channels from a first viewpoint;
[0019] A second audio signal is obtained for at least the first and second channels from a second viewpoint;
[0020] Based at least on the difference between the first audio signal and the second audio signal, a mono-channel differential audio signal is determined for the second viewpoint;
[0021] Depending on the mono differential audio signal for the second viewpoint and the first audio signal, it is possible to estimate the first and second channels of the second audio signal for the second viewpoint.
[0022] According to various, but not necessarily all, embodiments, an apparatus for adaptive playback is provided, comprising components configured to perform the following operations:
[0023] At least depending on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint, a monophonic differential audio signal is obtained for the second viewpoint; and
[0024] Based on the mono differential audio signal and the first audio signal, estimate the first and second channels of the second audio signal for the second viewpoint.
[0025] According to some, but not necessarily all, examples, the apparatus includes components configured to perform the following operations: obtaining a monophonic difference audio signal in the time domain for a second viewpoint, depending at least on the time-domain difference between a first audio signal for a first viewpoint and a second audio signal for a second viewpoint; and
[0026] Depending on the mono differential audio signal in the time domain and the first audio signal in the time domain, the first and second channels of the second audio signal for the second viewpoint are estimated in the time domain.
[0027] According to some, but not necessarily all, examples, the device is configured such that if the second viewpoint corresponds to a head rotation relative to the first viewpoint, the second audio signal is estimated based at least on the addition of one of the first and second channels of the mono audio signal involving the mono audio signal differential with the first audio signal, and the subtraction of the other of the first and second channels of the mono audio signal differential with the first audio signal.
[0028] According to some, but not necessarily all, examples, the device is configured such that if the second viewpoint corresponds to a head translation relative to the first viewpoint, the second audio signal is estimated at least based on the addition of one of the first and second channels of the mono audio signal involving the mono audio signal differential with the first audio signal and the other of the first and second channels of the mono audio signal differential with the first audio signal, or at least based on the subtraction of one of the first and second channels of the mono audio signal differential with the first audio signal and the other of the first and second channels of the mono audio signal differential with the first audio signal.
[0029] According to some, but not necessarily all, examples, the device includes components configured to perform the following operations:
[0030] When the second viewpoint shifts positively by a first angle from the first viewpoint and the third viewpoint shifts negatively by a first angle from the first viewpoint, the inverse of the mono-differential audio signal for the second viewpoint is reused as the mono-differential audio signal for the third viewpoint.
[0031] According to various, but not necessarily all, embodiments, a method is provided, comprising:
[0032] At least depending on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint, a monophonic differential audio signal is obtained for the second viewpoint; and
[0033] Based on the mono differential audio signal and the first audio signal, estimate the first and second channels of the second audio signal for the second viewpoint.
[0034] According to various, but not necessarily all, embodiments, an apparatus for implementing adaptive playback is provided, comprising components for performing the following operations:
[0035] Obtain the first audio signal from the first viewpoint;
[0036] Obtain the second audio signal from the second viewpoint;
[0037] Based on the difference between the first audio signal and the second audio signal, at least the difference audio signal is determined for the second viewpoint;
[0038] Smooth the differential audio signals in the frequency domain to obtain a smoothed first differential audio signal;
[0039] Depending on the smoothed differential audio signal and the first audio signal, it is possible to estimate at least the second audio signal.
[0040] According to various, but not necessarily all, embodiments, a method is provided that includes:
[0041] Obtain the first audio signal from the first viewpoint;
[0042] Obtain the second audio signal from the second viewpoint;
[0043] Based on the difference between the first audio signal and the second audio signal, at least the difference audio signal is determined for the second viewpoint;
[0044] Smooth the differential audio signals in the frequency domain to obtain a smoothed first differential audio signal;
[0045] Depending on the smoothed differential audio signal and the first audio signal, it is possible to estimate at least the second audio signal.
[0046] Examples as claimed in the appended claims are provided according to various, but not necessarily all, embodiments. Attached Figure Description
[0047] Some examples will now be described with reference to the accompanying drawings, in which:
[0048] Figure 1 Examples of the topics described in this article are shown;
[0049] Figure 2 This provides another example of the topic described in this article;
[0050] Figure 3 This provides another example of the topic described in this article;
[0051] Figure 4 This provides another example of the topic described in this article;
[0052] Figure 5 This provides another example of the topic described in this article;
[0053] Figure 6 This provides another example of the topic described in this article;
[0054] Figure 7 This provides another example of the topic described in this article;
[0055] Figure 8 This provides another example of the topic described in this article;
[0056] Figure 9A and Figure 9B This provides another example of the topic described in this article;
[0057] Figure 10A and Figure 10B This provides another example of the topic described in this article;
[0058] Figure 11 This provides another example of the topic described in this article;
[0059] Figure 12A This provides another example of the topic described in this article;
[0060] Figure 12B This provides another example of the topic described in this article;
[0061] Figure 13A This provides another example of the topic described in this article;
[0062] Figure 13B This provides another example of the topic described in this article;
[0063] Figure 14 This provides another example of the topic described in this article;
[0064] Figure 15A This provides another example of the topic described in this article;
[0065] Figure 15B This provides another example of the topic described in this article;
[0066] Figure 16 This provides another example of the topic described in this article;
[0067] Figure 17 This provides another example of the topic described in this article;
[0068] Figure 18A This provides another example of the topic described in this article;
[0069] Figure 18B This provides another example of the topic described in this article;
[0070] Figure 19 Here is another example of the topic described in this article. Detailed Implementation
[0071] Figure 1 An example of a system 10 for playing back audio signal 60 is shown. System 10 includes a transmitter device 20 that communicates with receiver device 30 via interface 12. In some examples, interface 12 may be a wireless interface, such as a radio interface.
[0072] The transmitter device 20 is configured to: obtain a first audio signal 601 for a first viewpoint 401; obtain a second audio signal 602 for a second viewpoint 402; determine at least a difference audio signal 70 for the second viewpoint 402 based on the difference between the first audio signal 601 and the second audio signal 602; and, depending on the difference audio signal 70 and the first audio signal 601, make it possible to estimate at least the second audio signal 602.
[0073] For example, the estimation of at least the second audio signal 602 can be achieved by transmitting the difference audio signal 70 to the receiver device 30 via the interface 12.
[0074] Receiver device 30 includes components configured to perform the following operations:
[0075] At least depending on the difference between the first audio signal 601 for the first viewpoint 401 and the second audio signal 602 for the second viewpoint 402, a difference audio signal 70 is obtained for the second viewpoint 402; and
[0076] Based on the difference audio signal 70 and the first audio signal 601, the second audio signal 602 for the second viewpoint 402 is estimated.
[0077] The differential audio signal 70 can be defined in a variety of different ways.
[0078] Figure 2 It shows Figure 1 An example of system 10 is shown. In this example, the first audio signal 601 is an audio signal for at least the first channel 51 and the second channel 52, and the second audio signal 602 is an audio signal for at least the first channel 51 and the second channel 52. In this example, but not necessarily in all examples, the first channel 51 is the left (L) channel, and the second channel 52 is the right (R) channel.
[0079] In this example, the transmitter device 20 is configured to estimate both the first channel 51 and the second channel 52 of the second audio signal 602, depending on the difference audio signal 70 and the first audio signal 601. Furthermore, the receiver device 30 is configured to estimate the first channel 51 and the second channel 52 of the second audio signal 602 for the second viewpoint 402, depending on the difference audio signal 70 and the first audio signal 601.
[0080] In this example, the differential audio signal 70 can have various different forms. For example, the following abbreviations can be used:
[0081]
[0082] First, reduce the number of channels in the signal surrounding the user's line of sight to minimize delay. As an example, one of the following three formulas can be used:
[0083] (Equation 1 for the differential audio signal 70)
[0084] X 90 =L 90 -L0 (Equation 2 for the differential audio signal 70)
[0085] X 90 =R0-R 90 (Equation 3 for the differential audio signal 70)
[0086] The difference audio signal 70 based on the difference between the first audio signal 601 and the second audio signal 602 can be considered as the difference between the reference channels of the first audio signal 601 and the second audio signal 602, wherein the reference channel is the first channel 51, the second channel 52, or a synthesized channel based on the first channel 51 and the second channel 52.
[0087] In Equation 1, the reference channel is the right minus the left. R0-L0 is the reference channel for the first audio signal, and R 90 -L 90 This is the reference channel for the second audio signal. The difference between the reference channel of the first audio signal and the reference channel of the second audio signal is X. 90 .
[0088] In Equation 2, the reference channel is the left channel L. L0 is the reference channel for the first audio signal, and L... 90 X is the reference channel of the second audio signal, and the difference between the reference channel of the first audio signal and the reference channel of the second audio signal is X. 90 .
[0089] In Equation 3, the reference channel is the right channel R. R0 is the reference channel for the first audio signal, and R 90 This is the reference channel for the second audio signal. The difference between the reference channel of the first audio signal and the reference channel of the second audio signal is X. 90 .
[0090] from Figure 2 The general understands that transmitter device 20 creates a channel difference audio signal 70 from the first audio signal 601 and the second audio signal 602, but then only transmits the first audio signal 601 and the channel difference audio signal 70. It does not transmit the second audio signal 602. Then, transmitter device 30 reverses this process and uses the difference audio signal 70 and the first audio signal 601 to recreate or estimate the second audio signal 602. Therefore, it will be understood that by transmitting the difference audio signal 70 instead of the second audio signal 602, the amount of information transmitted on interface 12 is reduced.
[0091] exist Figure 2 In the example shown, the difference audio signal 70 is obtained by subtracting only the second audio signal 602 from the first audio signal 601. This is performed at the difference component 22. Then, the estimator 32 in the transmitter device 30 adds only the first audio signal 601 to the difference audio signal 70 to recover the estimate of the second audio signal 602. Figure 2 In the example, the difference 22 and the estimate 32 occur independently for each of the first channel 51 and the second channel 52.
[0092] Below Figure 3In the example, system 10 is further refined so that only the mono differential audio signal 70 is transmitted from transmitter device 20 to receiver device 30. By comparison... Figure 2 and Figure 3 It can be seen that, Figure 2 In this process, two differential audio signals 70 are transmitted from the transmitter device 20 to the receiver device 30 for the reconstruction of the second audio signal 602; that is, there is a differential audio signal 70 for each channel 51, 52. However, in Figure 3 In this process, a single differential audio signal 70 is transmitted from the transmitter device 20 to the receiver device 30 for the reconstruction of the second audio signal 602. That is, a single differential audio signal 70 is sent from the transmitter device 20 to the receiver device 30. This single differential audio signal 70 is used for one channel and is referred to as a mono differential audio signal 70.
[0093] The term "mono-differential audio signal 70" can be replaced with:
[0094] "Mono, Differential Audio Signal 70" or
[0095] "Mono representation of differential audio signals 70" or
[0096] "Channel 70, including differential audio signals" or
[0097] "The differential audio signal 70 is configured to be transmitted in mono."
[0098] Differences can be based on one or more channels, but the representation of a difference is mono.
[0099] refer to Figure 8 For example, transmitter device 20 can perform smoothing of the differential audio signal 70 in the frequency domain to obtain a smoothed first differential audio signal 70'. This smoothing operation can, for example, be applied to... Figure 1 , Figure 2 or Figure 3 Example execution.
[0100] exist Figure 3 In the example, transmitter device 20 includes components configured to perform the following operations:
[0101] A first audio signal 601 for at least the first channel 51 and the second channel 52 is obtained for the first viewpoint 401;
[0102] A second audio signal 602 for at least the first channel 51 and the second channel 52 is obtained for the second viewpoint 402;
[0103] Based at least on the difference between the first audio signal 601 and the second audio signal 602, a monophonic differential audio signal 70 is determined for the second viewpoint 402; and
[0104] Depending on the mono differential audio signal 70 for the second viewpoint 402 and the first audio signal 601, it is possible to estimate both the first channel 51 and the second channel 52 of the second audio signal 60 for the second viewpoint 402.
[0105] Receiver device 30 includes components configured to perform the following operations:
[0106] At least depending on the difference between the first audio signal 601 for the first viewpoint 401 and the second audio signal 602 for the second viewpoint 402, a mono differential audio signal 70 is obtained for the second viewpoint 402; and
[0107] Based on the mono differential audio signal 70 for the second viewpoint 402 and the first audio signal 601, the first channel 51 and the second channel 52 of the second audio signal 602 for the second viewpoint 402 are estimated.
[0108] exist Figure 3 In the example shown, receiver device 30 is configured to estimate the first channel 51 of second audio signal 602 using at least the first channel 51 of first audio signal 601 and the mono differential audio signal 70 for second viewpoint 402, and to estimate the second channel 52 of second audio signal 602 using at least the second channel 52 of first audio signal 601 and the same mono differential audio signal 70 for second viewpoint 402.
[0109] As previously described, the difference component 22 for determining the difference audio signal 70 (e.g., mono difference audio signal 70) for the second viewpoint 402 based on the difference between the first audio signal 601 and the second audio signal 602 is configured to determine the difference between reference channels of the first audio signal 601 and the second audio signal 602, wherein the reference channel is the first channel 51, the second channel 52, or a synthesized channel based on the first channel 51 and the second channel 52, and wherein the estimator 32 is configured to depend on the mono difference audio signal 70 and the reference channel of the first audio signal 601, such that both the first channel 51 and the second channel 52 of the second audio channel 602 can be estimated.
[0110] exist Figure 3 In the example shown, the reference channel is the second channel 52.
[0111] In the transmitter device 20, a component configured to determine the mono difference audio channel 70 for a second viewpoint 402 is configured to determine the difference between the first audio signal 601 and the second audio signal 602 in the time domain or in the frequency domain.
[0112] The advantage of determining the difference in the time domain is that it provides a lower latency. In an example using time-domain differences, receiver device 30 includes components configured to: obtain a mono-channel difference audio signal 70 for the second viewpoint 402 in the time domain, depending at least on the time-domain difference between a first audio signal 601 for a first viewpoint 401 and a second audio signal 602 for a second viewpoint 402; and estimate, in the time domain, a first channel 51 and a second channel 52 of the second audio signal 602 for the second viewpoint 402, depending on the mono-channel difference audio signal 70 in the time domain and the first audio signal 601 in the time domain.
[0113] The advantage of calculating differences in the frequency domain is that only a portion of the available frequencies can be used. For example, differences can be calculated for high frequencies, while signals for the corresponding viewpoints can be transmitted for low frequencies.
[0114] Figure 4 An example of a method 200 for implementing adaptive playback of audio is shown. Method 200 includes: at block 202, obtaining a first audio signal 601 for at least a first channel 51 and a second channel 52 for a first viewpoint 401. At block 204, method 200 includes obtaining a second audio signal 602 for at least the first channel 51 and the second channel 52 for a second viewpoint 402. At block 206, method 200 includes determining a mono differential audio signal 70 for the second viewpoint 402, at least based on the difference between the first audio signal 601 and the second audio signal 602. At block 208, method 200 includes, depending on the mono differential audio signal 70 for the second viewpoint 402 and the first audio signal 601, enabling estimation of both the first channel 51 and the second channel 52 of the second audio signal 602 for the second viewpoint 402.
[0115] In block 206, the estimation can be provided by transmitting the mono differential audio signal 70 from transmitter device 20 to receiver device 30.
[0116] Method 200 can be performed by transmitter device 20.
[0117] Figure 5 An example of a method 210 for adaptive playback audio is shown. Method 210 includes: in block 212, obtaining a mono differential audio signal 70 for a second viewpoint 402, depending at least on the difference between a first audio signal 601 for a first viewpoint 401 and a second audio signal 602 for a second viewpoint 402.
[0118] In block 214, method 210 includes estimating a first channel 51 and a second channel 52 of a second audio signal 602 for a second viewpoint 402, depending on the mono differential audio signal 70 and the first audio signal 601. The mono differential audio signal 70 may be a mono differential audio signal 70 for the second viewpoint 402.
[0119] In the preceding examples, a single alternative viewpoint 402 and a single second audio signal 602 have been described. However, the preceding description can be applied to any number of different viewpoints 402. i and the corresponding audio signal 60 i Used together. Therefore, although the preceding examples illustrate a main stream associated with the first viewpoint 401 and the first audio signal 601, and a single side stream associated with the second viewpoint 402 and the second audio signal 602, in other examples, there may be multiple such side streams, each associated with a different viewpoint 401. i and the corresponding audio signal 60 i Related.
[0120] Information indicating the direction of the main flow and the side flow can be transmitted between the transmitter device 20 and the receiver device 30.
[0121] Typically, side stream directions can be + / - 20°, 40°, 60°, 90°, or 120° to the left (positive) or right (negative) from the main stream direction. This provides enough direction that switching between different streams doesn't cause auditory problems, and these directions are far enough to the left and right that even if the user quickly moves their viewpoint, there will usually be a side stream near the changing user viewpoint. For many use cases, such as watching movies or other non-360° content, a smaller number of side streams is often sufficient. For example, you can use only the main stream and any single 30° side stream.
[0122] Figure 6A method 220 for selecting a stream for use is shown. In box 222, method 220 checks whether the main stream direction is closest to the current user's line of sight. If yes, the method moves to box 224; otherwise, it moves to box 230. In box 224, the method plays the main stream (channels of the first audio signal 601) to the user. In box 230, method 220 determines which side stream is "closest" to the current line of sight in that direction. In box 232, method 220 combines the side stream with the main stream, as described in the preceding example. This may include, for example, estimating the first channel 51 and the second channel 52 of the second audio output signal 602 for the second viewpoint 402, depending on the mono differential audio signal 70 and the first audio signal 601 for the second viewpoint 402. Or, more generally, depending on the i-th viewpoint 40... i Mono differential audio signal 70 i And the first audio signal 601, estimated for the i-th viewpoint 40 i The i-th signal 60 i The first channel 51 and the second channel 52. Then, in box 234, the estimated audio signal is rendered to the user.
[0123] In some examples, the choice of how to render the main stream and side streams to the user is made after all streams have been decoded. In this example, all audio samples are available in the temporal domain, and the selection can be done sample by sample. In an alternative example, the selection can be made before decoding, thus saving processing power since not all streams need to be decoded. However, with this option, the latency will be longer (due to audio decoding latency). This can be reduced by using low-latency audio and encoder / decoder for the side streams.
[0124] Figure 7 The system 10 described above is shown, which has been further developed to reduce the amount of information transmitted from the transmitter device 20 to the receiver device 30. In this example, a mono differential audio signal 70 is used for multiple viewpoints 402, 403.
[0125] In this example, the second viewpoint 402 is offset from the first viewpoint 401 by a positive first angle +α, and the third viewpoint 403 is offset from the first viewpoint 401 by a negative first angle (-α). The transmitter device 20 is configured to obtain a mono differential audio signal 70 for the second viewpoint 402 but not for the third viewpoint 403. The mono differential audio signal 701 for the second viewpoint 402 is transmitted from the transmitter device 20 to the receiver device 30 and can be used to estimate audio signals for both the second viewpoint 402 and the third viewpoint 403. The mono differential audio signal 701 for the third viewpoint 403 is not transmitted from the transmitter device 20 to the receiver device 30.
[0126] Receiver device 30 uses the mono differential audio signal 70 for the second viewpoint 402 to estimate the second audio signal 602 for the second viewpoint 402, as previously described. Furthermore, receiver device 30 reuses the inverse of the mono differential audio signal 70 for the second viewpoint 402 as the mono differential audio signal 70 for the third viewpoint 403. The mono differential audio signal 70 for the third viewpoint 403 is then used, as previously described, to estimate the third audio signal 603 for the third viewpoint 403 by combining it with the first audio signal 601.
[0127] The symmetry between the second viewpoint 402 and the third viewpoint 403 allows a single difference signal 70 to be used to estimate the audio signal for these different viewpoints. Figure 7 In this context, a mono differential audio signal of 70 is equivalent to X. 90 =R0-R 90 That is to say, X α =R0-R α .
[0128] exist Figure 7 At the top, regarding the second viewpoint 402, adaptation is achieved by adding the mono differential audio signal 70 for the second viewpoint 402 to the second channel 52 of the main stream (first audio signal 601) and subtracting the mono differential audio signal 70 for the second viewpoint 402 from the first channel 51 of the main stream (first audio signal 601). This estimates the second channel 52 and the first channel 51 of the second audio signal 602, respectively.
[0129] For two different viewpoints 402 and 403, the mono differential audio signal 70 is sent only once, because the symmetry of the problem allows the differential signal 70 and its inverse to be used in the direction α° to the left and α° to the right from the current viewing direction.
[0130] exist Figure 7 At the bottom, regarding the third viewpoint 403, adaptation is achieved by subtracting the mono differential audio signal 70 for the second viewpoint 402 from the second channel 52 of the main stream (first audio signal 601) and adding the mono differential audio signal 70 for the second viewpoint 402 to the first channel 51 of the main stream (first audio signal 601). This estimates the second channel 52 and the first channel 51 of the third audio signal 603, respectively.
[0131] In this way, by using the same mono differential audio signal 70 for both symmetrical directions, the number of mono differential audio signals 70 transmitted is reduced by half.
[0132] Although the mono differential audio signal 70 has been described above, it should be understood that this method can also be used when using a multi-channel differential audio signal 70.
[0133] Any of the aforementioned examples of transmitter device 20 can be modified to introduce, for example Figure 8 The smoothing component 100 is shown. In this example, the smoothing component 100 is configured to smooth the difference audio signal 70 in the frequency domain to obtain a smoothed difference audio signal 70', and depending on the smoothed difference audio signal 70' and the first audio signal 601, to enable the estimation of at least a second audio signal 602.
[0134] The difference audio signal 70 can be a mono difference audio signal 70. Then, in this example, the smoothing component 100 is configured to smooth the mono difference audio signal 70 in the frequency domain to obtain a smoothed mono difference audio signal 70', and depending on the smoothed mono difference audio signal 70' and the first audio signal 601, it is possible to estimate at least the second audio signal 602.
[0135] Figure 9A An example of a differential audio signal 70 is shown, which could be, for example, a mono differential audio signal 70. The signal is displayed as a spectrum, where the x-direction indicates increasing frequencies. The signal is divided into multiple distinct frequency compartments. The frequency domain is divided into different frequency ranges, as shown by dashed lines, and each frequency range includes one or more frequency compartments. Figure 9A The difference audio signal 70 before smoothing is shown. Figure 9B The smoothed difference audio signal is shown. After smoothing, the difference audio signal 70 is referred to as the smoothed difference audio signal 70'. After smoothing, the mono difference audio signal 70 is referred to as the smoothed mono difference audio signal 70'.
[0136] exist Figure 9A and Figure 9B and equivalent Figure 10A and Figure 10B In the example shown, the smoothing component 100 is configured to replicate frequency bins within one or more different frequency bands. This results in each frequency band having a single value on all frequency bins within that band, such as... Figure 9B As shown.
[0137] Figure 10A Is with Figure 9A Equivalent graph Figure 10B yes Figure 10A A smooth equivalent diagram. It shows the mono-channel differential audio signal 70. Figure 10B It shows Figure 10A The smoothed version of the mono differential audio signal 70 shown.
[0138] In some examples, the smoothing component 100 is configured to perform dynamic smoothing. The dynamic smoothing of the (mono)difference audio signal 70 depends on the probability of viewpoint change from the first viewpoint 401 to the second viewpoint 402, at least based on the difference between the first audio signal 601 for the first viewpoint 401 and the second audio signal 602 for the second viewpoint 402. Therefore, different smoothing parameters (e.g., bandwidth size and quantity) can vary with the probability of viewpoint change. A smoothed (mono)difference audio signal 70' for a more probable viewpoint 40 can have a larger but smaller bandwidth than a smoothed (mono)difference audio signal 70' for a less probable viewpoint 40.
[0139] Therefore, in some examples, transmitter device 30 includes components configured to perform the following operations:
[0140] A first audio signal 601 is obtained from the first viewpoint 401;
[0141] A second audio signal 602 is obtained for the second viewpoint 402;
[0142] Based on the difference between the first audio signal 601 and the second audio signal 602, at least a difference audio signal 70 is determined for the second viewpoint 402;
[0143] Smooth the differential audio signal 70 in the frequency domain to obtain a smoothed differential audio signal 70'; and
[0144] Depending on the smoothed differential audio signal 70 and the first audio signal 601, it is possible to estimate at least the second audio signal 602.
[0145] The transmitter device 30 also performs an equivalent method, namely, obtaining a first audio signal 601 for the first viewpoint 401;
[0146] A second audio signal 602 is obtained for the second viewpoint 402;
[0147] Based on the difference between the first audio signal 601 and the second audio signal 602, at least a difference audio signal 70 is determined for the second viewpoint 402;
[0148] In the frequency domain, the differential audio signal 70 at the second viewpoint 402 is smoothed to obtain a smoothed differential audio signal 70'; and
[0149] Depending on the smoothed differential audio signal 70' and the first audio signal 601, it is possible to estimate at least the second audio signal 602.
[0150] exist Figure 9BIn the example, the value representing all bins in the frequency band is used to replace all bins within that band. This significantly reduces the bit rate required to encode the smooth, differential audio signal 70'.
[0151] In some embodiments, binning values near the frequency band boundary can be smoothed toward binning values in adjacent frequency bands.
[0152] Depending on the time-frequency transformation, the value can be real or complex.
[0153] In some examples, the smoothing component 100 performs averaging. The average value of the cells within a frequency band can be used to represent all cells within that band. The average value can be a direct average of complex-valued cells, where the average of the absolute value and angle value of a cell is the average of the values of the cells, or a cell that is closer to the average value, etc.
[0154] However, other smoothing methods are possible. For example, any low-pass filter would also be suitable. The intention is to reduce the variance of the differential audio signal 70 through smoothing.
[0155] In some examples, codebooks or other parameterized implementations can be used to represent the values of smoothed frequency bands.
[0156] Frequency bands can be selected based on any suitable method. For example, they can be triple octave bands, block bands, or ERB equivalent rectangular bands. In some examples, the bands are narrower at low frequencies and wider at high frequencies.
[0157] Figure 8 An encoder 110 is shown to optionally be present to encode the smooth, differential audio signal 70.
[0158] In some examples, the encoders used are MPEG AAC, MP3, MPEG AAC+, and MPEG AAC-LD encoders. Voice encoders, such as AMR-WB, can also be used. Even a mono encoder can be used to encode each channel of a multi-channel audio signal individually.
[0159] Multi-channel audio codecs, such as MPEG, AAC, and Dolby Digital, can also be used. A single multi-channel audio codec can be used to encode several streams.
[0160] Alternatively, a specially designed audio encoder can be used to perform low-bit-rate encoding of the differential audio signal 70 with a copy bay. This encoder can be designed to fully utilize the structure of the smooth differential audio signal 70.
[0161] In some examples, the number of sidestreams and the angles chosen for them depends on the extent to which the user can rotate his head and / or how good the quality is expected. Sidestreams deemed less likely (e.g., those angularly furthest from the user's line of sight, where the user is unlikely to turn his head) can be encoded at a smaller bit rate than more likely sidestreams. Additionally, for less important / less frequently used sidestreams, the frequency band used for the replication bin can be wider.
[0162] In some examples, the differential audio signal 70 is generated in the time domain and processed in the time domain at the receiver device 30. In this example, after the smoothed differential audio signal 70' is determined in the frequency domain, it is converted from the frequency domain to the time domain. The conversion to the time domain can occur, for example, at the transmitter device 20 or the receiver device 30. In some examples, frequency bin copying can occur within the audio encoder using a time-frequency transformation employed by the encoder.
[0163] Figures 11 to 16 An example of a receiver device 30 is shown, which operates to generate different audio signals 60 for different viewpoints. In these examples, the viewpoint is a combination of angle and / or movement. The angle can be a two-dimensional angle or a three-dimensional angle. In this particular example, the angle is a two-dimensional rotation in the horizontal plane. And in this example, the movement is a small movement in the horizontal plane, such as tilting forward, tilting backward, tilting left, or tilting right.
[0164] Receiver device 30 receives a main stream comprising a first audio signal 601 for both the first channel R and the second channel L. It also receives multiple side streams. Each side stream is a mono differential audio signal 70 for a specific viewpoint. It can be used to estimate the audio signal 60 for that viewpoint, and its inverse can be used to estimate the audio signal 60 for a symmetrically opposite viewpoint.
[0165] For example, a mono differential audio signal 70 TL (For left turns, rotation α = 90°) can be used to estimate the audio signal 60 for user rotation α. TL ( Figure 12A -Turn left [TL], viewpoint 40 TL ), and its inverse can be used to estimate the audio signal rotated 60° for -α user. TR ( Figure 12B -Turn right [TR], viewpoint 40 TR ).
[0166] Mono differential audio signal 70 LF (For forward tilt) can be used to estimate the audio signal 60 for a user leaning forward. LF ( Figure 13A - Lean forward [LF], viewpoint 40LF ), and its inverse can be used to estimate the audio signal 60 for a user tilting backward. LB ( Figure 13B - Lean back [LB], viewpoint 40 LB ).
[0167] Mono differential audio signal 70 LL (For left-leaning) can be used to estimate the audio signal 60 for a user leaning to the left. LL ( Figure 15A -Leftward [LL], viewpoint 40 LL ), and its inverse can be used to estimate the audio signal 60 for a user tilting to the right. LR ( Figure 15B -Right tilt [LR], viewpoint 40 LR ).
[0168] As will be noted from the figure, for rotation, the mono differential audio signal 70 for rotational viewpoint 40 and the first audio signal 601 of the main stream are combined in the estimator 32 for different channels L and R in opposite senses. Figure 12A , Figure 12B ).
[0169] It will also be noted that, for tilt, the mono differential audio signal 70 for tilt viewpoint 40 and the first audio signal 601 of the main stream are combined in the estimator 32 for different channels L, R in the same sense. Figure 13A , Figure 13B ; Figure 15A , Figure 15B ).
[0170] Figure 11 This illustrates the situation when the user is looking directly ahead. The main stream, including the first audio signal 601, is rendered to the user in the L and R channels.
[0171] Figure 12A This shows the user turning left (i.e., viewpoint 40). TL The estimator 32 uses the data specific to that viewpoint 40. TL Mono differential audio signal 70 TL To amplify the right channel R and attenuate the left channel L. (Regarding this viewpoint 40) TL Mono differential audio signal 70 TL The right channel R of the first audio signal 601 of the main stream is added, and the left channel L of the first audio signal 601 of the main stream is subtracted to estimate the signal for viewpoint 40. TL audio signal 60 TL .
[0172] Figure 12BThis shows the user turning to the right (i.e., viewpoint 40). TR Its relationship with viewpoint 40 TL (The case of symmetrical relative viewpoints). Estimator 32 uses a viewpoint 40 for symmetrical relative viewpoints. TL Mono differential audio signal 70 TL This is used to attenuate the right channel R and amplify the left channel L. (Targeting viewpoint 40) TL Mono differential audio signal 70 TL The left channel L of the first audio signal 601 of the main stream is added, and the right channel R of the first audio signal 601 of the main stream is subtracted to estimate the signal for viewpoint 40. TR (and viewpoint 40) TL 60 (symmetrically relative) audio signals TR .
[0173] Figure 13A This shows the user tilting forward (i.e., viewpoint 40). LF Example 32. Estimator 32 uses data from viewpoint 40. LF Mono differential audio signal 70 LF This amplifies both the right channel (R) and the left channel (L). (Regarding this viewpoint 40) LF Mono differential audio signal 70 LF The right channel R of the first audio signal 601 of the main stream is added, and the left channel L of the first audio signal 601 of the main stream is added, to estimate for that viewpoint 40. LF audio signal 60 LF .
[0174] Figure 13B This shows the user tilting backward (i.e., viewpoint 40). LB Its relationship with viewpoint 40 LF An example of (symmetrically relative). Estimator 32 uses a viewpoint 40 for symmetrically relative positions. LF Mono differential audio signal 70 LF To attenuate the right channel (R) and left channel (L). For viewpoint 40. LF Mono differential audio signal 70 LF Subtract from the left channel L of the first audio signal 601 of the main stream and subtract from the right channel R of the first audio signal 601 of the main stream to estimate for viewpoint 40. LB (and viewpoint 40) LF 60 (symmetrically relative) audio signals LB .
[0175] Figure 15A This shows the user tilting to the left (i.e., viewpoint 40). LL Example 32. Estimator 32 uses data from viewpoint 40. LLMono differential audio signal 70 LL This amplifies both the right channel (R) and the left channel (L). (Regarding this viewpoint 40) LL Mono differential audio signal 70 LL The right channel R of the first audio signal 601 of the main stream is added, and the left channel L of the first audio signal 601 of the main stream is added, to estimate for that viewpoint 40. LL audio signal 60 LL .
[0176] Figure 15B This shows the user tilting to the right (i.e., viewpoint 40). LR Its relationship with viewpoint 40 LL An example of (symmetrically relative). Estimator 32 uses a viewpoint 40 for symmetrically relative positions. LR Mono differential audio signal 70 LR To attenuate the right channel (R) and left channel (L). For viewpoint 40. LR Mono differential audio signal 70 LR Subtract from the left channel L of the first audio signal 601 of the main stream and subtract from the right channel R of the first audio signal 601 of the main stream to estimate for viewpoint 40. LR (and viewpoint 40) LL 60 (symmetrically relative) audio signals LR .
[0177] There can also be different independent combinations of rotation, forward / backward tilt, and left / right tilt. For example, rotation can be defined independently as + / -α, and / or forward / backward tilt can be defined as forward or backward, and / or left / right tilt can be defined as left or right.
[0178] For example, Figure 14 The diagram shows a combination of forward tilting and left turn. The diagram actually combines... Figure 13A and Figure 12A .
[0179] Figure 16 The diagram shows a combination of tilting to the left and turning left. The diagram actually combines... Figure 15A and Figure 12A .
[0180] Other combinations are possible, such as leaning forward, leaning left, and turning left, which would combine... Figure 13A , Figure 15A and Figure 12ATherefore, it should be understood that the receiver device 30 is configured to manage head rotation. The receiver device 30 is configured to estimate the second audio signal 602, if the second viewpoint 402 corresponds to a head rotation relative to the first viewpoint 401, based at least on the addition involving one of the first channel L and the second channel R of the mono differential audio signal 70 and the first audio signal 601, and the subtraction involving the other of the first channel L and the second channel R of the mono differential audio signal 70 and the first audio signal 601.
[0181] Alternatively or additionally, receiver device 30 is configured to manage head translation. Receiver device 30 is configured to estimate a second audio signal 602 if the second viewpoint 402 corresponds to a head translation relative to the first viewpoint 401, based at least on the addition of one of the first channel L and the second channel R of the mono differential audio signal 70 and the first audio signal 601, and the addition of the other of the mono differential audio signal 70 and the first channel L and the second channel R; or at least on the subtraction of one of the first channel L and the second channel R of the mono differential audio signal 70 and the first audio signal 601, and the subtraction of the other of the first channel L and the second channel R of the mono differential audio signal 70 and the first audio signal 601.
[0182] In these examples, the rotation mono signal (mono difference audio signal 70 for rotating viewpoints) represents the difference between the current and future head-viewing direction binaural signals, and the translation mono signal (mono difference audio signal 70 for different tilts) represents the difference between the current and future head-translation binaural signals. When the user rotates their head, the corresponding rotation mono signal 70 is added to the left channel of the current head-viewing direction binaural signal and subtracted from the right channel of the current head-viewing direction binaural signal, respectively. When the user translates their head, the corresponding translation mono signal 70 is added to both channels L and R of the current head-viewing direction binaural signal 601. The rotation mono signal and translation mono signal are independent and are combined (added / subtracted) with the current head-viewing direction binaural signal 601 independently of each other.
[0183] Typically, there will be more side flows for more viewpoints (especially orientations) than shown. This is indicated by the ellipsis "...".
[0184] Multiple side streams can also be blended into the main stream. For example, if the future translation is towards the right front and the right front side stream is unavailable, device 30 can blend the front and right side streams and add the blended side stream to the main stream. The amount of translation side signal added can depend on the amount of user head movement.
[0185] Side streams can be mixed in varying amounts to obtain an interpolated version for directions where no side stream is available. For example, if a user is looking at a 30° direction and no 30° side stream is available, the device can mix the available side streams to create an interpolated version for the 30° direction. For instance, mixing one-third of a 10° side stream with two-thirds of a 40° side stream can give an approximate 30° side stream.
[0186] Figure 17 An example of a transmitter device 20 is shown, which is configured to generate multiple side streams (multiple mono differential audio signals 70 for different viewpoints).
[0187] There are three mono differential audio signals 70 for three different rotations α, 2α, and 3α. There are also mono differential audio signals 70 for four different translations (forward, backward, left, and right). The respective mono differential audio signals 70 are smoothed by the respective smoothing components 100 before being encoded by the respective encoders 110 for transmission to the receiver device 30, as described above.
[0188] As will be understood from the foregoing, this disclosure introduces viewpoint tracking audio with near-zero latency and a significantly low bit rate. This is achieved by transmitting one or more difference audio signals 70 (which have differences between the current viewpoint and possible future viewpoints) for different future viewpoints, in addition to transmitting the current gaze direction audio (first audio signal 601). Zero latency can be achieved by adding and / or subtracting the difference audio signals 70 from the current gaze direction audio signal 601 in the time domain.
[0189] Low bitrates can be achieved by making the difference signal 70 mono and repeating it in the frequency domain after smoothing by 100%. Using an efficient codec, bandwidths of less than 150 kb / s for near-CD quality music and less than 64 kb / s for speech can be achieved.
[0190] In some embodiments, due to symmetry, the difference signal 70 is set only once for two different (symmetrical) viewpoints.
[0191] Therefore, it will be understood that sidestreams can be modified to be more compatible with low bit rate coding by reducing the number of channels in the sidestream. Sidestreams can also be modified to be more compatible with low bit rate coding by replicating frequency bins in the sidestream. In sidestreams associated with the direction in which the user might turn his or her head, even more frequency bin replication (more bandwidth) can be used.
[0192] Figure 18AAn example of controller 400 is shown. Such a controller can be used in transmitter device 20 and / or receiver device 30. Controller 400 can be implemented as controller circuitry. Controller 400 can be implemented solely in hardware, have certain aspects solely in software (including firmware), or be a combination of hardware and software (including firmware).
[0193] like Figure 18A As shown, the controller 400 can be implemented using instructions that implement hardware functions, for example, by using executable instructions 406 of a computer program in a general-purpose or special-purpose processor 402, which can be stored on a computer-readable storage medium (disk, memory, etc.) for execution by such processor 402.
[0194] Processor 402 is configured to read from and write to memory 404. Processor 402 may also include an output interface (through which data and / or commands are output by processor 402) and an input interface (through which data and / or commands are input to processor 402).
[0195] Memory 404 stores a computer program 406 comprising computer program instructions (computer program code) that controls the operation of devices 20 and 30 when loaded into processor 402. The computer program instructions of computer program 406 provide the logic and routines that enable the devices to perform the required methods. Processor 402 can load and execute computer program 406 by reading from memory 404.
[0196] Therefore, device 20 may include:
[0197] At least one processor 402; and
[0198] At least one memory 404 includes computer program code;
[0199] At least one memory 404 and computer program code are configured, together with at least one processor 402, to cause the devices 20, 30 to perform at least the following:
[0200] A first audio signal is obtained for at least the first and second channels from a first viewpoint;
[0201] A second audio signal is obtained for at least the first and second channels from a second viewpoint;
[0202] Based at least on the difference between the first audio signal and the second audio signal, a monophonic differential audio signal is determined for the second viewpoint, and
[0203] Depending on the mono differential audio signal for the second viewpoint and the first audio signal, estimate both the first and second channels of the second audio signal for the second viewpoint.
[0204] Therefore, device 30 may include:
[0205] At least one processor 402; and
[0206] At least one memory 404 includes computer program code;
[0207] At least one memory 404 and computer program code are configured, together with at least one processor 402, to cause the devices 20, 30 to perform at least the following:
[0208] At least depending on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint, a monophonic differential audio signal is obtained for the second viewpoint; and
[0209] Based on the mono differential audio signal and the first audio signal, estimate the first and second channels of the second audio signal for the second viewpoint.
[0210] Therefore, device 20 may include:
[0211] At least one processor 402; and
[0212] At least one memory 404 includes computer program code;
[0213] At least one memory 404 and computer program code are configured, together with at least one processor 402, to cause the devices 20, 30 to perform at least the following:
[0214] Obtain the first audio signal from the first viewpoint;
[0215] Obtain the second audio signal from the second viewpoint;
[0216] Based on the difference between the first audio signal and the second audio signal, at least the difference audio signal is determined for the second viewpoint;
[0217] Smooth the differential audio signals in the frequency domain to obtain a smoothed first differential audio signal;
[0218] Depending on the smoothed differential audio signal and the first audio signal, it is possible to estimate at least the second audio signal.
[0219] like Figure 18BAs shown, computer program 406 can reach devices 20 and 30 via any suitable delivery mechanism 408. Delivery mechanism 408 can be, for example, a machine-readable medium, a computer-readable medium, a non-transitory computer-readable storage medium, a computer program product, a memory device, a recording medium (such as a compact optical disc read-only memory (CD-ROM) or a digital versatile optical disc (DVD) or solid-state memory), or an article of manufacture that includes or tangibly embodies computer program 406. The delivery mechanism can be a signal configured to reliably transmit computer program 406. Devices 20 and 30 can propagate or transmit computer program 406 as a computer data signal.
[0220] Computer program instructions are used to cause device 20 to perform at least the following operations or to perform at least the following operations:
[0221] A first audio signal is obtained for at least the first and second channels from a first viewpoint;
[0222] A second audio signal is obtained for at least the first and second channels from a second viewpoint;
[0223] Based at least on the difference between the first audio signal and the second audio signal, a monophonic differential audio signal is determined for the second viewpoint, and
[0224] Depending on the mono differential audio signal for the second viewpoint and the first audio signal, it is possible to estimate both the first and second channels of the second audio signal for the second viewpoint.
[0225] Computer program instructions are used to cause device 30 to perform at least the following operations or to perform at least the following operations:
[0226] Based on at least the difference between a first audio signal for a first viewpoint and a second audio signal for a second viewpoint, a monophonic difference audio signal is obtained for the second viewpoint; and
[0227] Based on the mono differential audio signal and the first audio signal, estimate the first and second channels of the second audio signal for the second viewpoint.
[0228] Computer program instructions may be included in a computer program, a non-transitory computer-readable medium, a computer program product, or a machine-readable medium. In some, but not necessarily all, examples, computer program instructions may be distributed across more than one computer program.
[0229] Although memory 404 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which may be integrated / removable and / or provide permanent / semi-permanent / dynamic / cached storage.
[0230] Although processor 402 is shown as a single component / circuit, it can be implemented as one or more separate components / circuits, some or all of which may be integrated / removable. Processor 402 may be a single-core or multi-core processor.
[0231] References to “computer-readable storage medium,” “computer program product,” “tangible computer program,” or “controller,” “computer,” “processor,” etc., should be understood to include not only computers with different architectures (such as single / multiprocessor architectures and sequential (von Neumann) / parallel architectures), but also special-purpose circuits, such as field-programmable gate arrays (FPGAs), special-purpose circuits (ASICs), signal processing devices, and other processing circuitry systems. References to computer programs, instructions, code, etc., should be understood to include software or firmware for programmable processors, such as, for example, the programmable content of hardware devices, whether instructions for processors or configuration settings for fixed-function devices, gate arrays, or programmable logic devices, etc.
[0232] As used in this application, the term "circuit" may refer to one or more, or all of the following:
[0233] (a) Pure hardware circuit system implementations (such as implementations using purely analog and / or digital circuits), and
[0234] (b) A combination of hardware circuitry and software, such as (if applicable):
[0235] (i) A combination of analog and / or digital hardware circuitry and software / firmware, and
[0236] (ii) Any part of a hardware processor (including one or more digital signal processors), software, and memory (one or more) having software, which work together to enable a device (such as a mobile phone or server) to perform various functions, and
[0237] (c) (one or more) hardware circuits and / or (one or more) processors, such as (one or more) microprocessors or a portion thereof, which require software (e.g. firmware) for operation, but may be absent when operation is not required.
[0238] This definition of "circuit" applies to all uses of the term in this application, including in any claim. As a further example, as used in this application, the term "circuit" also covers only the implementation of hardware circuitry or processors with their accompanying software and / or firmware. For example, and if applicable to a particular claim element, the term "circuit" also covers baseband integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or networking devices.
[0239] Figure 19 This is an example of a receiver device 30, which is configured to not only generate an audio signal 60, but also render the audio signal 60. The receiver device 30 includes a controller 400 as described above, and also includes an audio rendering device 420 for rendering audio based on the audio signal 60. As previously mentioned, the audio signal 60 can vary depending on the user's current viewpoint.
[0240] In this example, receiver device 30 is headphones 410. For example, headphones could be a pair of over-ear headphones or a pair of augmented reality or virtual reality glasses.
[0241] In some examples, the headset 410 can communicate via a wireless interface 12 that provides wireless data connectivity (such as Bluetooth connectivity). In some examples, the transmitter device 20 is a mobile phone or similar or other personal electronic device.
[0242] In some examples, the user's viewpoint used in the foregoing examples can be determined by the viewpoint of the headset 410. The viewpoint of the headset 410 can be tracked using sensors in the headset 410. In this case, the headset 410 sends head tracking information to the transmitter device 20.
[0243] In an alternative implementation, the user's viewpoint can be tracked by using sensors at the transmitter device 20 or elsewhere to track the user.
[0244] The sensor used for head tracking can be, for example, an accelerometer built into the headset 410, but it can also be of other types, such as optical, camera, infrared, Bluetooth, LT antenna array, 3D camera, etc. The tracking sensor can reside outside the headset 410. For example, a device similar to Microsoft Connect can be used to track the user's head position from outside the headset 410.
[0245] The headset 410 has applications such as augmented reality, virtual reality, and teleconferencing. The transmitter device 20 can modify the audio based on head tracking information. The transmitter device 20 sends the modified audio to the headset 410, and the headset further modifies / selects the audio played to the user. The head tracking information is delayed upon arrival at the transmitter device 20 (due to transmission delay) compared to the actual current user gaze direction. The transmitter device 20 uses the delayed head tracking information to create different audio streams. A high-quality stereo audio stream (first audio signal 601) is optimized for the user's delayed gaze direction. This is the main stream. Other side streams (other different audio signals 70 for different viewpoints 40) are of lower quality and can be used to modify the main stream so that it becomes optimized for other viewpoints, one of which is typically close to the current user gaze direction. The headset 410 continuously modifies the main stream in this way based on the current head tracking information.
[0246] As mentioned earlier, the modification (for rotation) is completed by adding the side stream associated with the current user's gaze direction to the left channel of the main stream and subtracting the side stream from the right channel of the main stream.
[0247] The audio signal 60 used can be stereo, two-channel, 5.1, or Ambisonics (ambient stereo), where Ambisonics or 5.1 has more than two channels and may not be able to reduce different signals to a mono signal. Instead, some channels in 5.1 or Ambisonics can be divided into stereo pairs, and each pair uses a different signal.
[0248] Typically, the side viewpoint 40 will be fixed; for example, typical choices for orientation could be + / -20°, + / -40°, + / -60°, + / -90°, or + / -120°, as these are close to the possible directions in which the user can turn their head. However, in some cases, there may be other reasons for choosing these directions. The selection can be made in the mobile phone or the headset 410. Either of these devices can determine the more probable direction. If the determination is made in the headset 410, the determined direction needs to be transmitted to the mobile phone 20 so that it can be used in the determined direction, including the side viewpoint 70.
[0249] Either device 20 or 30 can determine the direction of a sound source from the audio signal transmitted from mobile phone 20 to headset 30 or from a real-world sound environment. For real-world sound sources, the device requires at least two (typically three or four) microphones to detect the direction of the sound source. Methods such as beamforming or time difference can be used to detect the direction of the sound source. The direction of the sound source (such as a speaker in a conference call or another real-world person (not the user)) is the possible direction when the user can turn their head. These directions can be used to create more probable lateral flow directions, and these more probable lateral flows can be encoded at a higher bit rate than other lateral flows.
[0250] The boxes shown in the figure may represent steps in a method and / or code segment of computer program 406. The illustration of a specific order of boxes does not necessarily imply a required or preferred order for these boxes, and the order and arrangement of the boxes may vary. Furthermore, some boxes may be omitted.
[0251] Where a structural feature has been described, it may be replaced by a component that performs one or more functions of that structural feature, whether or not the function or those functions are explicitly described or implicitly described.
[0252] The above example is applied as an enabling component for the following systems:
[0253] Automotive systems; telecommunications systems; electronic systems including consumer electronics; distributed computing systems; media systems for generating or rendering media content, including audio, video, and audio-visual content, as well as mixed, mediated, virtual, and / or augmented reality; personal systems including personal health systems or personal fitness systems; navigation systems; user interfaces, also known as human-computer interfaces; networks including cellular, non-cellular, and optical networks; self-organizing networks; the Internet of Things; the Internet of Things; virtualized networks; and related software and services.
[0254] The term “includes” is used in this document in an inclusive rather than exclusive sense. That is, any reference to X including Y indicates that X may include only one Y, or may include more than one Y. If the intention is to use “includes” in an exclusive sense, it will be made clear in the context by referring to “includes only one…” or by using “consisting of…”.
[0255] Various examples have been referenced in this description. Descriptions of features or functions of an example indicate those features or functions present in that example. The terms “example,” “for example,” “may,” or “can” are used throughout to indicate that, whether explicitly stated or not, such a feature or function exists at least in the described example, whether or not it is described as an example, and that it may, but not necessarily, exist in some or all other examples. Therefore, “example,” “for example,” “may,” or “can” refers to a specific instance of a class of examples. An instance’s property may be a property belonging only to that instance, or it may be a property of the class, or it may be a property of the class that includes some, but not all, instances of that class. Therefore, it is implicitly disclosed that features described with reference to one example but not to another example may, where possible, be used as part of a composition of work for that other example, but are not necessarily required to be used for that other example.
[0256] Although various examples have been described in the preceding paragraphs, it should be understood that modifications may be made to the given examples without departing from the scope of the claims.
[0257] The features described above can be used in combinations other than those explicitly described above.
[0258] Although some features have been described with reference to certain characteristics, those functions can be performed by other features, whether or not they are described.
[0259] Although features have been described with reference to some examples, these features may also exist in other examples, whether or not they are described.
[0260] The terms “a” or “that” are used in this document in an inclusive rather than exclusive sense. That is, any reference to X including one / that Y indicates that X may include only one Y or may include more than one Y, unless the context clearly indicates the opposite. If “a” or “that” is intended to have an exclusive meaning, then it will be clearly stated in the context. In some cases, “at least one” or “one or more” may be used to emphasize an inclusive meaning, but no exclusive meaning should be inferred from the absence of these terms.
[0261] The appearance of a feature (or combination of features) in a claim refers to the feature or combination of features itself, and also to features that achieve substantially the same technical effect (equivalent features). Equivalent features include, for example, features that are variations and achieve substantially the same result in substantially the same manner. Equivalent features include, for example, features that perform substantially the same function in substantially the same manner to achieve substantially the same result.
[0262] In this description, adjectives or adjectival phrases have been used to refer to various examples to describe the characteristics of the examples. Such descriptions of the characteristics of the examples indicate that the characteristics exist exactly as described in some examples, and exist substantially the same as described in other examples.
[0263] While efforts have been made in the foregoing specification to draw attention to those features deemed important, it should be understood that an applicant may seek protection by means of any patentable feature or combination of features mentioned above and / or shown in the drawings, whether or not such features or combinations of features are emphasized.
Claims
1. An apparatus for implementing adaptive playback, comprising at least one processor and at least one memory, the at least one memory including computer program code, the at least one memory and the computer program code being configured together with the at least one processor such that the apparatus at least: A first audio signal is obtained for at least the first and second channels from a first viewpoint; A second audio signal is obtained for at least the first channel and the second channel from a second viewpoint; Based at least on the difference between the first audio signal and the second audio signal, a monophonic difference audio signal is determined for the second viewpoint; and Depending on the mono differential audio signal and the first audio signal, it is possible to estimate the first channel and the second channel of the second audio signal for the second viewpoint.
2. The apparatus according to claim 1, wherein, The device determines the difference between the reference channels of the first audio signal and the second audio signal by determining the mono-channel differential audio signal, wherein the reference channel is the first channel, the second channel, or a synthesized channel based on the first channel and the second channel, and wherein the device determines the difference between the reference channels of the first audio signal and the second audio signal by depending on the mono-channel differential audio signal and the reference channel of the first audio signal, thereby enabling the estimation of the first channel and the second channel of the second audio signal.
3. The apparatus according to claim 1, wherein, The device determines the difference between the first audio signal and the second audio signal in the time domain by determining the mono-channel difference audio signal.
4. The apparatus of claim 1 is further configured such that: the mono differential audio signal is smoothed to obtain a smoothed mono differential audio signal, and depending on the smoothed mono differential audio signal and the first audio signal, at least the second audio signal can be estimated.
5. The apparatus according to claim 4, wherein, This allows the device to smooth the monophonic differential audio signal in the frequency domain.
6. The apparatus according to claim 5, wherein, This allows the device to smooth out frequency bins within one or more different frequency bands.
7. The apparatus according to claim 4, wherein, The device is made to smooth dynamically, wherein the dynamic smoothing of the mono-channel differential audio signal depends on the probability of viewpoint change from the first viewpoint to the second viewpoint, at least based on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint.
8. The apparatus according to claim 1, further comprising: When the second viewpoint is offset from the first viewpoint by a positive first angle and the third viewpoint is offset from the first viewpoint by a negative first angle, the monophonic differential audio signal is obtained for the second viewpoint but not for the third viewpoint.
9. A method for implementing adaptive playback, comprising: A first audio signal is obtained for at least the first and second channels from a first viewpoint; A second audio signal is obtained for at least the first channel and the second channel from a second viewpoint; Based at least on the difference between the first audio signal and the second audio signal, a monophonic difference audio signal is determined for the second viewpoint; and Depending on the mono differential audio signal for the second viewpoint and the first audio signal, it is possible to estimate the first channel and the second channel of the second audio signal for the second viewpoint.
10. The method according to claim 9, wherein, Determining the mono-channel differential audio signal includes: determining the difference between a reference channel of the first audio signal and the second audio signal, wherein the reference channel is the first channel, the second channel, or a synthesized channel based on the first channel and the second channel, and wherein the first channel and the second channel of the second audio signal can be estimated depending on the reference channel of the mono-channel differential audio signal and the first audio signal.
11. The method of claim 9, further comprising: The mono-channel differential audio signal is smoothed to obtain a smoothed mono-channel differential audio signal, and depending on the smoothed mono-channel differential audio signal and the first audio signal, at least the second audio signal can be estimated.
12. The method according to claim 11, wherein, Smoothing includes dynamic smoothing, wherein the dynamic smoothing of the mono-channel differential audio signal depends on the probability of viewpoint change from the first viewpoint to the second viewpoint, based at least on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint.
13. The method of claim 9, further comprising: When the second viewpoint is offset from the first viewpoint by a positive first angle and the third viewpoint is offset from the first viewpoint by a negative first angle, the monophonic differential audio signal is obtained for the second viewpoint but not for the third viewpoint.
14. An apparatus for adaptive playback, comprising at least one processor and at least one memory, the at least one memory including computer program code, the at least one memory and the computer program code being configured together with the at least one processor such that the apparatus at least: At least depending on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint, a monophonic differential audio signal is obtained for the second viewpoint; and Based on the mono differential audio signal and the first audio signal, the first and second channels of the second audio signal for the second viewpoint are estimated.
15. The apparatus of claim 14, wherein obtaining the mono-channel differential audio signal for the second viewpoint depends at least in the time domain on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint; and Based on the mono differential audio signal in the time domain and the first audio signal in the time domain, the first and second channels of the second audio signal for the second viewpoint are estimated in the time domain.
16. The apparatus of claim 14, wherein if the second viewpoint corresponds to a head rotation relative to the first viewpoint, the second audio signal is estimated at least based on addition involving one of the first and second channels of the mono audio signal and subtraction involving the other of the first and second channels of the mono audio signal.
17. The apparatus of claim 14, wherein if the second viewpoint corresponds to a head translation relative to the first viewpoint, the second audio signal is estimated at least based on addition involving one of the first and second channels of the mono audio signal and addition involving the other of the first and second channels of the mono audio signal, or at least based on subtraction involving one of the first and second channels of the mono audio signal and subtraction involving the other of the first and second channels of the mono audio signal.
18. The apparatus according to claim 14, wherein: When the second viewpoint deviates from the first viewpoint by a positive first angle and the third viewpoint deviates from the first viewpoint by a negative first angle, the inverse of the mono-channel differential audio signal for the second viewpoint is reused as the mono-channel differential audio signal for the third viewpoint.
19. A method for adaptive playback, comprising: At least depending on the difference between the first audio signal for the first viewpoint and the second audio signal for the second viewpoint, a mono-channel differential audio signal is obtained for the second viewpoint; as well as Based on the mono differential audio signal and the first audio signal, the first and second channels of the second audio signal for the second viewpoint are estimated.
20. The method of claim 19, comprising: When the second viewpoint deviates from the first viewpoint by a positive first angle and the third viewpoint deviates from the first viewpoint by a negative first angle, the inverse of the mono-channel differential audio signal for the second viewpoint is reused as the mono-channel differential audio signal for the third viewpoint.
Citation Information
Patent Citations
Wireless Head Mounted Display with Differential Rendering and Sound Localization
US20170045941A1
Predictive head-tracked binaural audio rendering
WO2019067445A1