Method and system for maintaining track length of pre-rendered spatial audio

By applying HRTF and attenuation operations in the audio system to generate and maintain the track length of the binaural audio version, the problem that pre-rendered spatial audio in the prior art is difficult to maintain track length, and more stable audio playback is achieved.

CN115442734BActive Publication Date: 2025-05-06APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210623737.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-04
Filing Date
2022-06-02
Publication Date
2025-05-06
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

The prior art has difficulty maintaining track lengths of audio tracks when pre-rendering spatial audio, resulting in challenges in audio playback devices in terms of processing and storage.

Method used

Generate binaural audio versions by receiving tracks with specific track lengths in the audio system and applying head-dependent transfer function (HRTF) and other audio signal processing operations. Then, an attenuation operation is performed to gradually lower the signal level so that the track length of the binaural audio version is the same as the original track.

Benefits of technology

Effectively maintain the track length of pre-rendered audio, ensuring that the audio playback device can correctly process and store audio content, and provide a more stable audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115442734B_ABST
    Figure CN115442734B_ABST
Patent Text Reader

Abstract

The present disclosure relates to "Method and system for maintaining track length of pre-rendered spatial audio." A method performed by a programmed processor of an audio system, the method comprising: receiving an audio track having a track length; generating a binaural audio version of the audio track, the binaural audio version having an extended track length; performing a decay operation on the binaural audio version to gradually reduce a signal level of the binaural audio version to below a signal threshold level at a time corresponding to an end time of the track length of the audio track along the extended track length; and storing the binaural audio version having the track length of the audio track in a memory for later transmission to an audio playback device for driving one or more speakers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of and priority to U.S. Provisional Patent Application Serial No. 63 / 197,210, filed on June 4, 2021, which is hereby incorporated by reference in its entirety. Technical Field

[0003] One aspect of the disclosure relates to maintaining track length of a pre-rendered spatial audio track. Other aspects are also described. Background Art

[0004] Binaural audio recording provides the listener with three-dimensional (3D) stereo sound that can be reproduced by a pair of speakers or headphones. To create this 3D sound, microphones are positioned at each ear of a mannequin's head in a room to capture the sound. When played back, the reproduced sound gives the listener the impression of listening to the sound in the room where it was originally recorded. Summary of the invention

[0005] One aspect of the present disclosure is a method performed by a programmed processor of an audio system, such as a remote electronic server. The audio system receives an audio track having a track length and generates a binaural audio version of the audio track, the binaural audio version having an extended track length. For example, the system may generate the binaural audio version by applying at least one spatial filter (e.g., a predefined head-related transfer function (HRTF), which may be generic and not personalized for any particular user). The extended track length may be a result of applying the HRTF and / or other audio signal processing operations, such as applying reverberation and performing equalization operations. For example, the application of the operation may add a reverberation tail to the audio track, resulting in an extended length. The system performs an attenuation operation on the binaural audio version to gradually reduce (e.g., fade out) a portion of the binaural audio version. For example, the attenuation operation may reduce the signal level of the binaural audio version to below (or equal to) a signal threshold level (e.g., 0 dB) at a time corresponding to an end time of the (original) track length of the audio track along the extended track length. The system stores the binaural audio version having the track length of the audio track in a memory for later transmission to an audio playback device for playback (e.g., driving one or more speakers). Thus, the system maintains the track length between the original audio track and the binaural audio version.

[0006] In one aspect, the signal threshold level is a first signal threshold level, the system further determines whether a portion of the binaural audio version has a signal level exceeding a second signal threshold level (e.g., such as 0 full decibel scale (dBFS)), and in response to determining that the portion exceeds the second signal threshold level, applies dynamic range compression based on a difference between the signal level of the portion and the second signal threshold level. In another aspect, the system may receive a request for streaming a binaural audio version of the audio track from an audio playback device via a computer network, retrieve the binaural audio version from a memory, and transmit the binaural audio version to the audio playback device via the computer network. In some aspects, the system, while performing the attenuation operation, trims an end portion of the binaural audio version along the extended track length beginning at a time corresponding to an end time of the track length of the audio track so that both the audio track and the binaural audio version have the same track length.

[0007] Another aspect of the disclosure is a method performed by a programmed processor of an electronic server, the method comprising: receiving an audio track having a track length, applying a head-related transfer function (HRTF) on the audio track to produce a binaural rendered track having an extended track length, determining whether a signal level of an end portion of the binaural rendered track that exceeds the track length of the audio track is below a signal threshold level, and clipping the end portion of the binaural rendered track in response to the signal level of the end portion being below the signal threshold level. In one aspect, determining whether the signal level of the end portion is below the signal threshold level includes determining whether the end portion of the audio track is below the signal threshold level and has a track length that is longer than the track length of the end portion of the binaural rendered track. In some aspects, determining whether the signal level of the end portion is below the signal threshold level includes determining whether the end portion of the binaural rendered track that begins at a time corresponding to an end time of the audio track remains below the signal threshold level along its track length.

[0008] In one aspect, in response to the signal level of the end portion not being below a signal threshold level, the electronic server reduces the end portion of the binaural rendering track by fading out a portion of the applied reverb on the binaural rendering track and cutting off the reduced end portion so that the binaural rendering track has the same track length as the audio track.

[0009] Another aspect of the present disclosure is a method performed by a programmed processor of an electronic server, the method comprising: receiving an ordered plurality of tracks of an audio album, each track having a corresponding track length; combining the ordered plurality of tracks to form a concatenation of tracks; spatially rendering the concatenation of tracks, wherein the spatial rendering causes at least one of the concatenated tracks to have an extended track length; spatially rendering the concatenation of separated tracks to form an ordered plurality of spatially rendered tracks, each spatially rendered track having a corresponding track length of its corresponding track in the ordered plurality of tracks, wherein a beginning portion of a spatially rendered track is an end portion of a previous spatially rendered track that exceeds its corresponding track length. In one aspect, the spatially rendered concatenation of separated tracks includes fading out a last spatially rendered track separated from the concatenation along the corresponding extended track length of the spatially rendered track at a time corresponding to an end time of the corresponding track in the ordered plurality of tracks.

[0010] Another aspect of the present disclosure is a method comprising: receiving an audio track having audio content; generating a binaural rendering audio track having the same track length as the audio track and having an end portion in which reverberation fades out; and transmitting the binaural rendering audio track to an audio playback device over a computer network for playback. In one aspect, generating the binaural rendering audio track comprises: applying an HRTF to the audio track to generate the binaural rendering audio track, applying reverberation to the binaural rendering audio track, and fading out the end portion of the binaural rendering audio track. In some aspects, the faded-out end portion is a first end portion, the binaural rendering audio track includes a second end portion starting at the end of the first end portion, wherein the method further comprises cutting off the second end portion of the binaural rendering audio track. In some aspects, the method also receives a request for streaming the binaural rendering audio track from the audio playback device over a computer network, wherein the track is transmitted in response to the request.

[0011] The above summary does not include an exhaustive list of all aspects of the present disclosure. It is contemplated that the present disclosure includes all systems and methods that can be practiced by all suitable combinations of the various aspects summarized above and disclosed in the detailed description below and specifically pointed out in the claims. Such combinations may have specific advantages not specifically set forth in the above summary of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Aspects are shown by way of example and not limitation in the illustrations of the accompanying drawings, where similar reference numerals indicate similar elements. It should be noted that reference to "one" or "an" aspect in the present disclosure is not necessarily the same aspect, and it means at least one. In addition, for the sake of brevity and to reduce the total number of drawings, a certain drawing may be used to illustrate features of more than one aspect, and for a certain aspect, not all elements in the drawing may be required.

[0013] Figure 1 A block diagram of an audio system including an audio content server for maintaining track lengths of pre-rendered spatial audio is shown according to one aspect.

[0014] Figure 2 An example of an audio content server according to one aspect is shown.

[0015] Figure 3 is a flow chart of one aspect of a process for maintaining a track length of a spatially rendered audio track according to one aspect.

[0016] Figure 4 is a flow chart of one aspect of a process for spatially rendering a concatenation of several audio tracks according to one aspect.

[0017] Figure 5 is a flow chart of one aspect of a process for transmitting a rendered audio track to an audio playback device according to one aspect.

[0018] Figure 6 Several stages are shown for applying a decay operation on a spatially rendered audio track to maintain the track length, according to one aspect.

[0019] Figure 7 Several stages are shown for trimming end portions of a spatially rendered audio track to maintain track length according to another aspect.

[0020] Figure 8 Several stages are shown for maintaining track lengths for concatenation of spatially rendered audio tracks according to some aspects.

[0021] Fig. 9 is a flow chart of a process for maintaining a track length of a spatially rendered audio track according to one aspect. DETAILED DESCRIPTION

[0022] The various aspects of the present disclosure will now be explained with reference to the accompanying drawings. As long as the shapes, relative positions and other aspects of the components described in a certain aspect are not clearly defined, the scope of the present disclosure is not limited to the components shown here, and the components shown are only for illustrative purposes. In addition, although many details are set forth, it should be understood that some embodiments can be implemented without these details. In other cases, well-known circuits, structures and technologies are not shown in detail to avoid blurring the understanding of the description. In addition, unless the meaning is clearly contrary, all ranges shown herein are considered to include the end values ​​of each range.

[0023] Some consumer devices are capable of creating three-dimensional (3D) sound effects (e.g., via headphones worn by the listener), thereby giving the impression that there are virtual sound sources in the 3D space around the listener. To achieve this effect, the device may retrieve (or stream) audio content, such as a soundtrack of a musical work (e.g., from an online music streaming platform), and render the audio content spatially using a spatial audio filter (e.g., head-related transfer function (HRTF)) that can be personalized to binaural audio signals for the listener. When these signals are used to drive the speakers of the headphones, the sound produced by the speakers may be perceived as originating from a specific source (e.g., from behind the listener). Rendering audio content on the listener device has disadvantages. For example, in order to spatialize the audio content, the device may require a large amount of audio data and may require a considerable amount of processing power. Devices that do not have a sufficient amount of processing resources may not be able to effectively produce 3D sound.

[0024] In this case, another device (e.g., a remote server) that is communicatively coupled to the listener device and provides 3D sound to the listener can render (or pre-render) audio content spatially. Pre-rendering audio content has some disadvantages. For example, in order to render the audio track spatially, the remote server can convolve the track with a head-related impulse response (HRIR) (which is a time domain representation of the HRTF). The convolution of the audio track can result in an extension or stretching of the track (e.g., a few milliseconds), which can be caused by the added reverberation tail of the reverberation reflected by the HRTF. The track can also be extended due to other audio signal processing operations performed during pre-rendering to provide a better listening experience. For example, reverberation can be used to render audio content spatially (e.g., to provide more space and depth to the sound). The application of reverberation can extend the track (e.g., hundreds of milliseconds), which can also be the result of adding (or extending) the reverberation tail at the end of the track. Therefore, the application of audio signal processing operations can extend the original track length (e.g., by adding additional sounds, such as the reverberation tail at the end of the track).

[0025] To overcome these deficiencies, the present disclosure describes an audio system that is capable of maintaining the track length of pre-rendered audio, such as the track of a music album. Specifically, the system receives a track with a specific track length and generates a binaural audio version of the track. The system determines that the binaural audio version has an extended track length, which can be attributed to the reverberation tail associated with the HRTF applied to the original track to generate the binaural audio version as described herein. The system performs an attenuation operation on the binaural audio version to gradually reduce the signal level of the binaural audio version to below the signal threshold level along the extended track length at a time corresponding to the end time of the track length of the original track. The system then stores the binaural audio version with the track length of the original track in a memory for later transmission to an audio playback device for driving one or more speakers. Therefore, the pre-rendered track can have the same track length as the original track.

[0026] Figure 1 A block diagram of an audio system 1 according to one aspect is shown, the audio system including an audio content server 5 for maintaining track lengths of pre-rendered spatial audio. Specifically, the system includes an audio playback device 2, an audio output device 3, a (e.g., computer) network (e.g., the Internet) 4, and the audio content server 5. In one aspect, the system may include more or fewer elements, such as having additional audio content servers, or not including an audio playback device. In this case, the audio output device can stream the audio content for output, as described herein.

[0027] In some aspects, audio content server 5 can be an independent electronic server, computer (for example, desktop computer) or a server computer cluster configured to pre-render audio content, as described herein. In one aspect, the audio content of server can be any type of (for example, user desired) audio content, such as the sound track of a musical work, the sound track of a moving picture, etc. In this case, server can be a part of a cloud computer system, which can pre-render audio content and stream the rendered audio content as a cloud-based service (for example, a subscription service) provided to one or more user devices. For example, an audio content server can be configured to stream audio content by an online media content (for example, audio and / or video) streaming platform. Therefore, audio content can be in the form of a sound track, wherein the sound track can be associated with other media content, such as the sound track of a moving picture. As shown in the figure, the server is communicatively coupled (for example, via a network) to an audio playback device, so that pre-rendered audio content is streamed for playback (for example, via an audio output device). More content about the operation performed by the server is described herein.

[0028] In one aspect, the audio playback device can be any electronic device (e.g., having electronic components, such as a processor, memory, etc.) that is capable of streaming (e.g., pre-rendering) audio content, such as a spatially rendered audio track (e.g., in the form of one or more binaural audio signals), for playback (e.g., via one or more speakers integrated into the playback device and / or via one or more audio output devices, as described herein). For example, the playback device can be a desktop computer, a laptop computer, a digital media player, etc. In one aspect, the device can be a portable electronic device (e.g., handheld), such as a tablet computer, a smartphone, etc. In another aspect, the device can be a head-mounted device, such as smart glasses, or a wearable device, such as a smart watch.

[0029] In one aspect, the audio output device 3 can be any electronic device including at least one speaker and configured to output (or playback) sound by driving a speaker. For example, as shown, the device is a wireless headset (e.g., in-ear headset or earplug), which is designed to be positioned on (or in) the user's ear, and is designed to output sound to the user's ear canal. In some aspects, the headset can be a sealing type with a flexible headset end, which is used to acoustically seal the entrance of the user's ear canal relative to the surrounding environment by blocking or occluding in the ear canal. As shown, the output device includes a left earphone for the user's left ear and a right earphone for the user's right ear. In this case, each headset can be configured to output at least one audio channel of video content (e.g., the right earphone outputs the right audio channel in the dual-channel input of the stereo recording of a musical work, and the left earphone outputs the left channel). On the other hand, the output device can be any electronic device including at least one speaker and arranged to be worn by the user and arranged to output sound by driving the speaker with an audio signal. As another example, the output device may be any type of earphone, such as a circum-aural (or supra-aural) earphone that at least partially covers the user's ear and is arranged to direct sound into the user's ear.

[0030] In some aspects, the audio output device may be a head mounted device, as shown herein. In another aspect, the audio output device may be any electronic device arranged to output sound into the surrounding environment. Examples may include a stand-alone speaker, a smart speaker, a home theater system, or an infotainment system integrated into a vehicle.

[0031] In one aspect, the audio output device can be a wireless device that can be communicatively coupled to an audio playback device to exchange audio data. For example, the playback device can be configured to establish a wireless connection with the audio output device via a wireless communication protocol (e.g., a Bluetooth protocol or any other wireless communication protocol). During the established wireless connection, the playback device can exchange (e.g., transmit and receive) data packets (e.g., Internet Protocol (IP) packets) with the audio output device, which data packets can include audio digital data in any audio format.

[0032] In another aspect, the audio playback device 2 may be communicatively coupled to the audio output device 3 via other methods. For example, both devices may be coupled via a wired connection. In this case, one end of the wired connection may be (e.g., fixedly) connected to the audio output device, while the other end may have a connector, such as a media jack or a universal serial bus (USB) connector, that is inserted into a socket of the audio playback device. Once connected, the playback device may be configured to drive one or more speakers of the audio output device using one or more audio signals via a wired connection. For example, the playback device may transmit the audio signal as digital audio (e.g., PCM digital audio). In another aspect, the audio may be transmitted in an analog format.

[0033] In some aspects, audio playback device 2 and audio output device 3 can be different (independent) electronic devices, as shown in this article. In another aspect, playback device can be a part of audio output device (or integrated with audio output device). For example, at least some of the parts of playback device (e.g., one or more processors, memory, etc.) can be a part of audio output device, and / or at least some of the parts of audio output device can be a part of playback device. In this case, at least some of the operations performed by audio playback device (e.g., streaming audio content from audio content server 5) can be performed by audio output device.

[0034] Figure 2An example of an audio content server 5 according to one aspect is shown. In one aspect, the audio content server can be operated by one or more audio content providers (e.g., via an online streaming platform), and can provide (e.g., streaming) audio content to one or more audio playback devices of, for example, equipment 2. For example, the server can (e.g., via a network 4) receive an audio content from equipment (e.g., audio playback device 2), such as (e.g., a musical work) track is streamed. The server can use any audio codec (e.g., MP3, AAC, etc.) to encode audio content, and the encoded audio content can be transferred to an audio playback device to be decoded and output (e.g., by utilizing one or more driver signals including audio content to drive one or more loudspeakers). In addition, the server can be configured to pre-render audio content for transmission to one or more audio playback devices later. More content about pre-rendering audio content is described herein.

[0035] The server includes a network interface 20, one or more processors 21 and a non-transitory machine-readable storage medium 22 (or memory). The network interface 20 provides an interface for the server 5 to communicate with the audio playback device 2 so as to stream the audio content. For example, the network interface is configured to establish a communication link with the audio playback device (for example, in response to receiving a request to stream the audio content, as described herein), and once established, the audio content is transmitted, as described herein. Examples of non-transitory machine-readable storage media may include read-only memory, random access memory, CD-ROMS, DVD, magnetic tape, optical data storage device, flash memory device and phase change memory. Although shown as being included in the server, one or more of the components can be part of a separate electronic device, such as medium 22 is a separate data storage device. For example, the storage medium can be part of an online database (or include an online database), and the server is coupled to the online database in communication. As shown, the non-transitory machine-readable storage medium stores a server software program 23, audio content 27 and one or more spatial audio filters 28 stored therein.

[0036] As described herein, audio content 27 may include tracks (e.g., multiple pieces of audio content), each track being a musical work of one or more (e.g., music) albums. In another aspect, the tracks may include any type of audio content, such as podcasts, tracks of motion pictures, and the like. In one aspect, audio content 27 may include one or more (e.g., different) versions of the same track. For example, audio content 27 may include different audio formats of the track. For example, audio content 27 may include a mono version (e.g., as one channel) or a stereo version (e.g., as two stereo channels) of the track. In one aspect, audio content may include a multi-channel version in any surround sound multi-channel format (e.g., 5.1, 7.1, etc.). As another example, audio content may include a sound space representation of a virtual sound source, such as, for example, a high-order Ambisonics (HOA) representation of a sound space including audio content (e.g., positioned at a virtual position in the space), a vector-based amplitude translation (VBAP) representation of sound, and the like. In another example, audio content may include an object-based representation of sound, which includes one or more audio channels, which include sound (at least a portion) and metadata describing the sound (e.g., the spatial characteristics of the sound). In some aspects, audio content may include other audio formats. On the other hand, audio content 27 may include one or more identical versions of tracks. For example, an audio content server may include several stereo versions with identical tracks, wherein each stereo version may be processed differently. As an example, an audio content server may include a stereo version of a track to which reverberation has been added, and include another stereo version that does not include applied reverberation. As described herein, audio content 27 may also include a spatial rendering version (e.g., a binaural audio version) of one or more tracks.

[0037] The spatial audio filter 28 includes one or more spatial filters for performing spatial rendering, as described herein. In one aspect, the spatial filters may include one or more HRTFs, equivalently, one or more HRIRs. In some aspects, the spatial filters may be predefined (or default) filters (e.g., defined in a controlled setting such as a laboratory), and therefore may not be personalized for any particular user of an audio playback device to which an audio content server streams audio content, as described herein. In another aspect, at least some of the spatial filters may be personalized for a user of an audio playback device (e.g., playback device 2 of audio system 1) to take into account the user's anthropometry.

[0038] When executed by one or more processors 21 of the content server to perform audio signal processing operations to maintain the track length of the pre-rendered spatial audio, the server software program 23 includes one or more operation blocks, such as a spatial renderer 24, an audio signal processing 25, and an attenuator 26. In one aspect, the server software program may include more or fewer operation blocks. In another aspect, at least some of the operations described herein may be performed by one or more other electronic devices communicatively coupled to the audio content server. For example, the audio playback device 2 may be configured to perform one or more audio signal processing operations on the audio content.

[0039] The spatial renderer 24 is configured to receive an audio track (e.g., retrieve the audio track from the audio content 27) and is configured to render the audio track spatially. Specifically, the renderer may apply one or more spatial filters 28 to the audio track (e.g., one or more channels constituting the audio track) to produce a spatially rendered (or processed) audio track. For example, the renderer may apply one or more HRTFs to the audio track to produce a binaural audio version of the audio track (e.g., as at least one binaural signal). In one aspect, the spatial renderer may generate several spatial renderings of the audio track based on the type of audio output device that outputs the track. For example, the renderer may spatially render a stereo recording using one or more HRTFs to produce a binaural audio signal for a head-mounted audio output device (e.g., headphones), as described herein. As another example, when the audio output device includes one or more speakers, the spatial renderer may render an HOA representation of the audio track to produce one or more speaker driver signals (e.g., based on a predefined speaker configuration).

[0040] The audio signal processing 25 may perform one or more audio signal processing operations on one or more audio tracks. Specifically, the server software program may perform operations before and / or after the audio track has been spatially rendered by the renderer 24. In one aspect, the processing 25 may be configured to perform a (dynamic) compression operation on the audio track. In one aspect, the compression may be based on the spatial rendering of the audio track. For example, before the audio track is spatially rendered, the audio track may have a dynamic range that does not exceed the signal threshold level. For example, the audio track may be a digital audio signal that does not exceed the digital domain range from a positive threshold (e.g., "+1") to a negative threshold ("-1"). In one aspect, the (e.g., positive) signal threshold may represent a maximum signal level, such as an audio system (or more specifically, an audio output device) of 0 full decibel scale (decibelsrelative to full scale / full scale relative decibel) (dBFS), and if used to drive one or more speaker drivers, the portion of the digital signal above the signal threshold may be cut. When the audio track is spatially rendered, the resulting digital signal may exceed the digital range (e.g., spanning +1 and / or -1). This may be due to the type of digital audio signal processing performed by the audio signal processor 25. For example, the server software program may be configured to perform floating point digital audio signal processing, where the audio data is in a floating point audio format with a high dynamic range (e.g., a dynamic range of more than 1,500 dB for a 32-bit audio sample). Thus, the dynamic range of the digital audio signal in the floating point audio format may exceed the digital domain range. Thus, the audio signal processor may be configured to determine whether the signal level of a portion of the audio track exceeds a signal threshold level of the digital domain. If so, dynamic range compression may be applied to the portion to produce a compressed audio signal having a digital waveform that remains within the digital range (e.g., does not exceed the signal threshold level). In one aspect, the dynamic range compression applied may be based on the difference between the signal level of the portion and the signal threshold level.

[0041] In another aspect, the audio signal processing 25 may perform other operations, such as adding (or applying) reverberation (or reverberation) to the (e.g., rendered) audio track. For example, the audio signal processing may apply a convolution reverberation to at least a portion of the spatially rendered audio track. In another aspect, the processing may apply a reverberant audio filter, wherein the processing may adjust the gain of the filter so that a certain amount of reverberation is applied (or added) to the track. As another example, the processing 25 may apply an equalization operation and a spectral shaping operation. For example, one or more (e.g., linear) filters (e.g., a low-pass filter, a band-pass filter, a high-pass filter, etc.) may be applied to (at least a portion of) the track. In another aspect, a scalar gain value may be applied to the audio track (e.g., to reduce the signal level of the track). In some aspects, any signal processing operation may be performed on the audio track.

[0042] As described herein, applying a spatial filter on a track may extend the length of the track. Thus, a spatially rendered track may have an extended track length relative to the track length of the (original or pre-rendered) track. For example, this may be caused by the reverberation tail of the reverberation reflected by the applied spatial filter. In another aspect, the length of the track may be extended based on one or more audio signal processing operations performed by the audio signal processing 25 of, for example, the server software program 23. For example, applying reverberation (or more specifically, the later reflections of the reverberation) on the track may extend the length of the track. Thus, the extended track length of the spatially rendered track to which reverberation has been applied may be based at least in part on 1) the spatial filter applied to the track and / or 2) the applied reverberation. In one aspect, the track may be extended by stretching the track to a new extended track length, as described herein. In another aspect, the track may be extended by adding an additional portion (e.g., a reverberation tail) to the end portion of the original track length (e.g., added at the end or stop time of the original track).

[0043] The attenuator 26 is configured to perform an attenuation operation on the spatially rendered (and / or audio processed) track so as to maintain the track length of the "original" track (e.g., the track before spatial rendering and / or audio processing). Specifically, the attenuation operation may be performed so that the spatially rendered track has the same track length (or approximately the same track length, e.g., within a threshold of the original track) as the original track (or more precisely, the version of the track before spatial rendering). In one aspect, the attenuation operation may be a "fade-out" of the track so that the signal level of the spatially rendered track gradually decreases to below a signal threshold level (e.g., -90 dBFS) along the spatially rendered extended track length at a time corresponding to the end time of the track length of the original track. In one aspect, the attenuator may begin fading out the spatially rendered track at a predetermined time before the end time of the original track. In another aspect, the attenuation operation may begin based on the difference between the track length of the original track and the extended track length of the spatially rendered track. In this case, the attenuator may be configured to determine that the spatially rendered track has an extended track length. For example, the fader may compare the track length of the processed track to the track length of the original audio track to determine the difference.

[0044] In one aspect, the fader can be configured to trim the track length of the processed track. Specifically, when performing the fade operation, the spatially rendered track may still include an end portion that extends beyond the length of the original track. In this case, the fader can trim an additional end portion of the processed track that begins along the extended track length of the track at a time corresponding to the end time of the track length of the original track, so that the processed track has the same track length as the original track. More information about trimming the additional portion is described herein.

[0045] In some aspects, the attenuator may trim the extra end portion of the processed track without performing an attenuation operation. Specifically, the attenuator may determine whether the end portion of the processed track (naturally) decays to silence before or at the end of the original track. For example, the attenuator may determine whether the signal level of the end portion of the spatially rendered track that exceeds the track length of the original track is below a signal threshold level (e.g., -90dBFS). If so, the attenuator cuts off the end portion of the spatially rendered track. In one aspect, without performing an attenuation operation, determining whether to trim the end portion may be based on whether the entirety of the end portion (of the rendered track that exceeds the original track) remains below the signal threshold level along the length of the end portion. However, if not, it means that at least some of the end portion is above the signal threshold level, and the attenuator may perform an attenuation operation.

[0046] In another aspect, the attenuator may perform an attenuation operation relative to the audio signal processing operations performed on the track. Specifically, the attenuator may fade out one or more audio signal processing operations (e.g., in addition to or instead of attenuating the signal level of the rendered track) to prevent the performed operations from extending the length of the track. For example, when a convolution reverb is applied to the track, the length of the track may be extended. In response, the attenuator may fade out the reverb (e.g., by gradually reducing the amount of reverb applied to the track) to prevent the reverb tail of the reverb from extending the track beyond the original length of the track or reducing the reverb tail of the reverb. In one aspect, the attenuator may fade out the processing operations at the end portion of the processed track. In another aspect, the attenuator may fade out the audio processing operations and signal levels of the track. More content about fading out one or more audio signal processing operations is described herein.

[0047] In some aspects, when both a spatial filter (e.g., HRTF filter) and reverberation are applied, the attenuator may attenuate the reverberation (e.g., at the end portion of the track), but may not fade out the spatial filtering. In one aspect, this may provide minimal auditory impact. In one aspect, the rendered track may still have an extended length based on the natural reverberation of the HRTF. Therefore, the renderer may only truncate the end portion of the track that extends beyond its original length.

[0048] In another aspect, the attenuator may cross-fade from spatial (e.g., binaural) to non-spatial (e.g., non-binaural) audio near the end portion of the track in order to get rid of (or reduce) the filtering of the rendered track (e.g., and thus extend it). As a result, the rendered track will be spatially rendered to the end portion of the track, at which point the spatial characteristics of the rendered track will be reduced. For example, at a specific time period along the length of the track, the attenuator may begin to cross-fade non-spatial audio into spatial audio, and may increase the amount of non-spatial audio added to the spatial rendering as the track increases (e.g., moving from a specific time period). In one aspect, at the end of the spatially rendered track, non-spatial rendered audio content may fade into the track (e.g., completely) instead of spatially rendered content. As a result of the reduced spatial aspect, the rendered track will have the same (or substantially the same) track length as the original track. In one aspect, performing cross-fading does not reduce the signal level of the rendered track (e.g., the end portion).

[0049] Figure 3-Figure 5 and Fig. 9 Flowcharts including processes 30, 40, 80 and 90, respectively, which may be performed by the processor 21 (e.g., the server software program 23 when executed by the processor) of the audio content server 5. In particular, at least some of the operations may be performed by at least some of the operation blocks 24-26, which are performed by the server software program, as described herein.

[0050] about Figure 3, this figure is a flow chart of a process 30 for maintaining (e.g., separately) a track length of a spatially rendered audio track according to one aspect. The process 30 begins with receiving an audio track by a server software program, wherein the audio track has a track length (at box 31). For example, the audio track may be received from audio content 27 stored in a memory of an audio content server (or from a remote memory device, as described herein). The software program spatially renders the audio track to produce a spatially rendered audio track (e.g., a binaural audio version of the audio track) having an extended track length (at box 32). Specifically, the software program applies (e.g., predefined) spatial filters (e.g., at least one HRTF) on the audio track to produce a rendered track (e.g., the rendered track may include one or more binaural audio signals). The software program performs one or more signal processing operations, such as dynamic range compression, applying reverberation, performing equalization operations, and / or applying at least one scalar gain, etc., as described herein (at box 33). In one aspect, the performance of the signal processing operations may be optional (as shown by box 33 having a dashed boundary). The software program determines whether the signal level of the end portion of the rendered track that exceeds the track length of the (original) audio track is below the signal threshold level (at decision box 34). Specifically, the software program determines whether the end portion of the spatially rendered track (e.g., the end portion of the rendered track that exceeds the track length of the original track) that begins at a time corresponding to the end time of the original audio track is below the threshold. If so, the software program trims the end portion of the rendered track so that the rendered track has the same track length as the original audio track (at box 35). The software program then stores the rendered track in a memory (at box 36). Specifically, the software program stores the spatially rendered track in the audio content 27 of the server 5.

[0051] However, if the end portion does not have a signal level below the signal threshold level, then the software program performs a decay operation on the spatially rendered track (at box 37). Specifically, the software program applies the decay operation to gradually reduce the signal level of the spatially rendered track to below the signal threshold level at a time corresponding to the end time of the track length of the audio track along the extended track length of the spatially rendered track. Thus, the program fades out the rendered track (e.g., to -90 dB) along its extended track length at a time corresponding to the end time of the original track. In one aspect, when performing the decay operation, the software program may cut off the end portion of the spatially rendered track so that the two tracks have the same track length.

[0052] In one aspect, the attenuation operation can reduce the (e.g., overall) signal level of the spatially rendered track. In another aspect, the attenuation operation can be applied to the spatial filter while the audio content is spatially rendered. For example, the software program can attenuate (e.g., reduce) the HRIR being convolved with the audio signal so as to attenuate the spatialization (e.g., with respect to time) of the audio track. Specifically, the reduction in the spatialization of the audio track can increase starting from a period of time (e.g., at the end of the audio track).

[0053] Figure 4 4 is a flow chart of one aspect of a process 40 for spatially rendering a concatenation of several tracks according to one aspect. Specifically, the process describes spatially rendering an entire (or at least a portion) of an audio (or music) album including a concatenation of several tracks, and maintaining a respective track length for each of the tracks. The process 40 begins by a server software program 23 receiving a plurality of ordered tracks of an audio (or music) album, each track having a respective track length (at block 41). For example, the software program may retrieve the entire (or at least a portion) of the audio album from the audio content 27. The software program combines the ordered tracks to form a concatenation of tracks (at block 42). Specifically, the tracks are added together in the order in which they appear in their audio album to create one track (e.g., a digital audio signal) (e.g., where track one is at the front end of the concatenation and the last track is at the back end of the concatenation). The software program spatially renders (e.g., as described herein, by applying one or more spatial filters to) the concatenation of tracks, wherein the rendering causes at least one of the tracks in the concatenation to have an extended track length (at block 43). For example, the concatenation may be stretched (e.g., linearly) such that each track within the concatenation extends at least a portion of the total extension of the concatenation. The software program (optionally) performs one or more audio signal processing operations, such as dynamic range compression, applying reverb and / or equalization (at block 44).

[0054] The software program 23 separates the spatially rendered audio tracks in series to form ordered spatially rendered audio tracks (e.g., where the order of the audio tracks is the same as the order of the original audio tracks), each of the spatially rendered audio tracks having a corresponding track length of its corresponding audio track of the ordered original audio track (at box 45). Specifically, the software program separates the audio tracks based on the order (and track length) of the audio tracks in the series before performing spatial rendering (and / or audio signal processing operations). For example, starting at the first audio track in the series, the program cuts along the length of the spatially rendered series at a time corresponding to the end time of the original first audio track. However, since the audio tracks in the series have been spatially rendered, each (or at least some) of the tracks have an extended track length. Therefore, once cut, the beginning portion of the second spatially rendered audio track (e.g., as the second track in the album and after the first track) is the end portion of the previous spatially rendered audio track (e.g., the first track in the album), which extends beyond its corresponding track length, and both of the spatially rendered audio tracks are part of the series. In one aspect, each successive spatially rendered track will (or may) begin at the end of the adjacent previous spatially rendered track. For example, the next cut to separate the second rendered track may be at a time along the length of the concatenation that corresponds to the combined track length of both the first original track and the second original track. Thus, the third (potential) rendered track will begin at the end portion of the second rendered track. The software program stores the plurality of ordered spatially rendered tracks in memory (at block 46).

[0055] In one aspect, the software program may perform the following with respect to the last spatially rendered track in the concatenation: Figure 3 , so that the tracks maintain their original track lengths. For example, fader 26 may fade out the last rendered track along its respective extended track length at a time corresponding to the end time of the corresponding track of the ordered track. In another aspect, the fader may determine when to fade out the last rendered track based on the track length of the concatenation. For example, once the concatenation is rendered, the fader may fade out the end portion of the concatenation (which is the end portion of the last rendered track) so that the spatially rendered concatenation has the same length as the original track concatenation. In some aspects, the end portion of the concatenation may be trimmed, as described herein.

[0056] In one aspect, the spatially rendered ordered series of tracks may not each be part of a particular audio album. Specifically, the tracks may be part of a collection of tracks organized in a particular order (e.g., by the server 5). For example, the order may be based on a listener preference or setting.

[0057] As described thus far, the server software program 23 can render tracks individually or a concatenation of rendered tracks (e.g., belonging to the same album) and maintain the track length of the rendered tracks in order to transmit the (pre-)rendered tracks to one or more audio playback devices. In one aspect, the server software program 23 can transmit individually rendered tracks or rendered tracks separated from the rendered concatenation based on how the listener plays back (or requests playback of) the tracks. For example, if the listener wants to listen to one rendered track, the server can transmit the individually rendered track (e.g., Figure 3 ), rather than being rendered as part of a track concatenation (as described in Figure 4 ). For example, because the length of the concatenation is stretched when rendering, the track length of each spatially rendered audio track in the concatenation may extend beyond its original length. Therefore, when the spatially rendered audio tracks are separated at their original length to be transmitted to the listener's device, each consecutive track will start at the end portion of the previous track of the album, which has been extended beyond its original length in the concatenation. Although a track starting at the end of the previous track may be imperceptible to listeners who listen to the tracks in the order they appear in their music album, problems may arise when the listener listens to the tracks out of order or individually. In this case, the listener may hear and perceive the last moments of the previous sound at the beginning of the track as an unexpected audio distortion or glitch (for example, when listening to the second rendered audio track of the album instead of the first rendered audio track).

[0058] Figure 5 8 is a flow chart of one aspect of a process 80 for transmitting a rendered audio track to an audio playback device according to one aspect. The process 80 begins by the server software program 23 receiving a request for streaming a rendered audio track from an audio playback device over a computer network (e.g., network 4) at block 81. For example, the audio playback device may include an audio playback application (e.g., a music streaming application) executed by the audio playback device (e.g., one or more processors of the audio playback device). Through the application (e.g., a graphical user interface (GUI) of the application displayed on a display screen of the audio playback device), a device user may select a particular audio track for playback. Once selected, the application may transmit the request to the server.

[0059] The server determines whether to transmit a separately rendered audio track, wherein the rendered audio track maintains the same track length as the corresponding original audio track (at decision box 82). Specifically, the server can make this determination based on a request received from an audio playback device. For example, the request may indicate that a user of the audio playback device is requesting streaming of a particular audio album. On the other hand, the request (e.g., metadata within the request) may indicate that the user is requesting the order in which the audio tracks are to be listened to (e.g., via user settings on the audio playback device). If so, the software program 23 retrieves the separately rendered audio tracks (e.g., audio content 27) from the memory (at box 83). The software program then transmits the rendered audio tracks to the audio playback device via a computer network (at box 84). Otherwise, if the software program determines that the user is requesting to listen to a pre-rendered audio track sequence, the software program retrieves a rendered audio track (at box 85) that is part of a serially rendered audio track sequence (e.g., an audio album), and transmits the rendered audio tracks (e.g., in a continuous order). Thus, the audio content server transmits the pre-rendered audio track having the same track length as the original audio track to one (or at least one) playback device.

[0060] Some aspects may perform variations of processes 30, 40, and 80. For example, certain operations may not be performed in the exact order shown and described. Certain operations may not be performed in a continuous series of operations, and different certain operations may be performed in different aspects.

[0061] Figure 6 Three stages 50-52 are shown according to one aspect, in which the audio content server 5 applies a decay operation on the spatially rendered audio track to maintain the track length (e.g., as Figure 3 30). The first stage 50 shows a digital waveform of a (e.g., original) audio track 53. Specifically, this stage shows the audio track along the track length from T0 to T1 (e.g., the positive portion of the digital waveform of ). The second stage 51 shows the digital waveform of a spatially rendered audio track 54. In one aspect, track 54 can be the result of applying one or more HRTFs to track 53. Thus, the rendered audio track 54 has an extended track length relative to the length of the original audio track. Specifically, the rendered track has an end portion (e.g., T1-T1') that extends beyond the original track end time T1. Thus, the rendered track has an extended track length T0-T1'.

[0062] It is also shown that the end portion of the rendered track drops to a signal threshold level ("Th"). Specifically, the signal level of the end portion of the rendered track is attenuated to Th. In one aspect, the signal threshold level may correspond to -90dBFS. In another aspect, the threshold level may be another sound pressure value. In one aspect, the drop in signal level may correspond to the reverberation tail reflected by the spatial filter applied to the track, which causes the end portion of the track to attenuate to -90dB (or silence). In another aspect, the end portion of the spatially rendered track may not attenuate to Th or below Th. The third stage 52 shows a spatially rendered track 54, which includes an attenuated end portion 55 that attenuates to Th at T1. For example, the attenuator 26 of the server software program 23 performs an attenuation operation in which the rendered track begins to attenuate at T2 and fades the track out so that the signal level of the track reaches (or becomes below) Th at T1. For example, the software program may attenuate the spatial rendering (e.g., reduce the applied spatial filter) on the track (e.g., at or before T2) in order to reduce the track length. In another aspect, the software program may attenuate the signal level of the entire audio signal to reduce the level to be equal to or below Th at T1.

[0063] Figure 7 Three stages 60-62 are shown according to another aspect, in which the audio content server 5 trims the end portion of the spatially rendered audio track to maintain the track length (e.g., Figure 3 30). The first stage 60 shows a digital waveform of an audio track 63 having a track length of T0-T2 and including an end portion that drops below Th. Specifically, this stage shows that the signal level of the audio track 63 drops below Th at T1 and remains below Th until T2. In one aspect, this end portion may represent silence at the end of the audio track. The second stage 61 shows a digital waveform of a spatially rendered audio track 64 (e.g., this is the result of applying one or more spatial filters, as described herein). Therefore, the rendered audio track 64 has an extended track length that extends beyond the end time T2 of the original audio track 63. Specifically, due to the spatial rendering, the signal level of the audio track drops below Th at T1' and remains below Th until the end time -T2' of the rendered track.

[0064] The third stage 62 shows that the end portion 65 of the spatially rendered audio track has been cut from the rendered audio track. Specifically, the server software program may determine that the end portion can be trimmed based on whether the end portion drops below a signal threshold level before the end of the original audio track. For example, the software program may determine to trim the end portion based on whether the signal level of the rendered audio track drops below Th before T2. In other words, the software program determines whether the rendered audio track naturally decays to silence before T2 (or at T2). In this case, since T1' (e.g., the time when the signal level drops below Th) is before T2, the software program has trimmed a portion of the rendered track between T2-T2. As described herein, if the end portion after T2' is still below T2, then the end portion can be trimmed. However, if the signal level increases to above Th between T2' and T2, then the software program may perform an attenuation operation, such as Figure 6 Shown.

[0065] Figure 8 Four stages 70-73 of track length for maintaining a concatenation of spatially rendered audio tracks are shown according to some aspects (e.g., as Figure 4 4). The first stage 70 shows a first track 74 having a first track length L1, and a second track 75 having a second track length L2. In one aspect, these tracks may be part of a sequence of tracks, such as tracks that are part of an audio album. In this case, the first track 74 may be the first track in the album, and the second track 75 may be the second track of the album. The second stage 71 shows the result of combining the two tracks into a concatenated track 76 (or a concatenation of the first track and the second track), thereby adding the second track to the end portion of the first track. Thus, the length of the concatenated tracks may be L1+L2.

[0066] The third stage 72 shows the spatially rendered concatenated audio tracks 77. Specifically, this stage shows the result of concatenating the spatially rendered audio tracks (and / or performing one or more additional audio signal processing operations on the audio tracks in concatenation, as described herein), such as Figure 4 As shown, the lengths of both tracks have been extended, and thus the lengths of the rendered concatenated tracks have also been extended due to the rendering. Specifically, the track length of a portion of the rendered concatenated tracks associated with the first track has been increased to L1', and similarly, the track length of another portion of the rendered concatenated tracks associated with the second track has been increased to L2'.

[0067] The fourth stage 73 shows the result of separating the concatenated track 77 into two spatially rendered tracks 78 and 79. Specifically, this stage shows how to separate the rendered tracks 78 and 79 so that both tracks have the same respective track lengths as their corresponding original tracks 74 and 75. For example, when separating the concatenated tracks, the server software program 23 may start at the first track and separate the first portion of the concatenated track 77 between the beginning of the track to L1 to create track 78. From L1, the server software program may separate the second portion between L1 and L2 to create track 79. As shown, since the rendered track 78 extends beyond L1, the end portion 90 of the spatially rendered first track is the beginning portion of the rendered second track 79. In one aspect, the server software program 23 may trim the end portion of the last rendered track separated from the concatenated track 77, as described herein. Here, the last track is track 79, and therefore the end portion 91 has been trimmed so that the second track 79 has a length of L2.

[0068] Fig. 9 Flowchart of process 90 for maintaining track length of spatially rendered audio track according to one aspect. Process 90 begins with receiving an audio track with audio content by server 5 (e.g., server software program 23 of the server) (at box 91). As described herein, the audio track may have any type of audio content in any type of audio format. In one aspect, the server receives the audio track in response to receiving a request for a spatially rendered version of the audio track, as described herein. The server generates a binaurally rendered audio track having the same track length as the audio track and having an end portion in which reverberation fades out (at box 92). Specifically, the server may apply an HRTF to the audio track to generate a binaurally rendered audio track. The server may apply reverberation to the binaurally rendered audio track and fade out the end portion of the binaurally rendered audio track. For example, the server may fade out the reverberation along the track length of the rendered track at a time before the end time of the original track. In one aspect, the reverberation may be reduced relative to time. For example, the server may reduce the reverberation (e.g., the gain of the reverberation) linearly relative to the time before the end time along the track length. In one aspect, the server may perform a fade-out operation such that any reverb tail added by applying the reverb is removed. In one aspect, the server may fade out the reverb after it is applied to the track. In another aspect, the server may fade out the reverb while it is applied to the track.

[0069] In one aspect, a binaural rendered audio track whose reverb has been faded out may still have an extended track length relative to the original track length. For example, as described herein, the application of an HRTF may extend the track length due to the reverb reflected in the HRTF. Thus, when the audio track is binaurally rendered and reverb is applied, there will be a first extended end portion due to the reverb tail added by the HRTF and a second extended end portion due to the reverb tail added by the reverb. As the reverb fades out, the second extended end portion may be reduced (or eliminated). Due to having the first extended end portion, the server may trim the portion (cut off the portion), thereby generating a binaural rendered audio track having the same track length as the original track length. The server transmits the binaural rendered audio track to an audio playback device via a computer network for playback (at box 93).

[0070] In one aspect, at least some of the operations described herein may be performed in response to determining that a signal level for a track length of a binaural rendered audio track is greater than a signal threshold level. For example, in response to determining that a signal level for an end portion (which extends beyond the original track length) is greater than a threshold level, the server may be configured to reduce the end portion by fading out a portion of the applied reverb on the rendered audio track (e.g., fading out the reverb at a certain time along the extended track length). As described herein, the end portion may still extend beyond the original track length due to the applied HRTF. Therefore, the server may trim off the reduced end portion, as described herein.

[0071] As described so far, the operations may be performed by the audio content server 5 (e.g., the server software program 23 of the audio content server). In another aspect, at least some of the operations may be performed by another electronic device, such as Figure 2 An audio playback device 2 is shown.

[0072] In one aspect, a method performed by a programmed processor of an electronic server includes receiving an ordered plurality of tracks of an audio album, each track having a corresponding track length; combining the ordered plurality of tracks to form a track concatenation; spatially rendering the track concatenation, wherein the spatial rendering causes at least one of the tracks in the concatenation to have an extended track length; and separating the spatially rendered track concatenation to form an ordered plurality of spatially rendered tracks, each spatially rendered track having a corresponding track length of its corresponding track of the ordered plurality of tracks, wherein a beginning portion of a spatially rendered track is an end portion of a previous spatially rendered track that extends beyond its corresponding track length so that both spatially rendered tracks are part of the concatenation.

[0073] In another aspect, the previous spatially rendered audio track is a first spatially rendered audio track, and the spatially rendered audio track is a second spatially rendered audio track, wherein both the first and second spatially rendered audio tracks have an extended track length that constitutes the length of the spatially rendered series, wherein separating the spatially rendered audio track series includes cutting along the length of the spatially rendered series at a time corresponding to the end time of the corresponding track length of the audio track corresponding to the first spatially rendered audio track in the ordered plurality of audio tracks. In some aspects, the spatially rendered series of the separated audio tracks includes fading out the last spatially rendered audio track separated from the series along the corresponding extended track length of the spatially rendered audio track at a time corresponding to the end time of the corresponding audio track in the ordered plurality of audio tracks. In another aspect, the electronic server performs one or more signal processing operations on the series, wherein the extended track length of at least some of the audio tracks in the audio tracks is based at least in part on the operations performed. In one aspect, the one or more signal processing operations include applying reverb and performing equalization operations.

[0074] It is understood that the use of personally identifiable information should be subject to privacy policies and practices that are generally recognized to meet or exceed industry or government requirements for maintaining user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use, and the nature of the authorized use should be clearly stated to users.

[0075] As previously mentioned, one aspect of the present disclosure may include a non-transitory machine-readable medium (e.g., a microelectronic memory) having stored thereon instructions that program one or more data processing components (generally referred to herein as "processors") to perform network operations, spatial rendering operations, and audio signal processing operations, as described herein. In other aspects, some of these operations may be performed by specific hardware components that include hardwired logic. Alternatively, those operations may be performed by any combination of programmed data processing components and fixed hardwired circuit components.

[0076] Although certain aspects have been described and shown in the drawings, it is to be understood that such aspects are merely illustrative of the broad disclosure and not limiting, and the disclosure is not limited to the specific constructions and arrangements shown and described, since various other modifications may occur to one of ordinary skill in the art. Therefore, the description is to be regarded as illustrative and not limiting.

[0077] In some aspects, the present disclosure may include language such as “at least one of [element A] and [element B]”. The language may refer to one or more of these elements. For example, “at least one of A and B” may refer to “A”, “B”, or “A and B”. Specifically, “at least one of A and B” may refer to “at least one of A and at least one of B” or “at least either A or B”. In some aspects, the present disclosure may include language such as “[element A], [element B], and / or [element C]”. The language may refer to any one of these elements or any combination thereof. For example, “A, B, and / or C” may refer to “A”, “B”, “C”, “A and B”, “A and C”, “B and C”, or “A, B, and C”.

Claims

1. A method performed by a programmed processor of an audio system, the method comprising: receiving an audio track having a track length; generating a binaural audio version of the audio track, the binaural audio version having an extended track length; performing a fade operation on the binaural audio version to gradually reduce a signal level of the binaural audio version to below a signal threshold level at a time along the extended track length corresponding to an end time of the track length of the audio track; as well as The binaural audio version having the track length of the audio track is stored in a memory for later transmission to an audio playback device for driving one or more speakers.

2. The method of claim 1, wherein generating the binaural audio version comprises applying a predefined head-related transfer function (HRTF) that is not personalized to a user of the audio playback device.

3. The method of claim 1 , wherein the signal threshold level is a first signal threshold level, wherein the method further comprises: determining whether a portion of the binaural audio version has a signal level exceeding a second signal threshold level; as well as In response to determining that the portion has a signal level exceeding the second signal threshold level, dynamic range compression is applied based on a difference between the signal level of the portion and the second signal threshold level. The method of claim 3 , wherein the second signal threshold level is 0 decibels relative to full scale (dBFS).

5. The method of claim 1, further comprising applying reverberation to the binaural audio version of the audio track, wherein the extended track length is based at least in part on the applied reverberation. 6 . The method of claim 5 , wherein performing the decay operation comprises fading out a portion of the applied reverb on the binaural audio version of the audio track.

7. The method of claim 1 , further comprising, while performing the decay operation, trimming an end portion of the binaural audio version of the audio track starting from the time along the extended track length corresponding to the end time of the track length of the audio track so that both the audio track and the binaural audio version have the same track length.

8. The method according to claim 1, further comprising: receiving a request from the audio playback device over a computer network to stream the binaural audio version of the audio track; retrieving the binaural audio version of the audio track from the memory; as well as The binaural audio version of the audio track is transmitted to the audio playback device via the computer network.

9. A method performed by a programmed processor of an electronics server, the method comprising: receiving an audio track having a track length; applying a head-related transfer function (HRTF) on the audio track to produce a binaural rendered track having an extended track length; determining whether a signal level of an end portion of the binaural rendering track exceeding the track length of the audio track is below a signal threshold level; as well as In response to the signal level of the terminal portion being below the signal threshold level, the terminal portion of the binaural rendering track is clipped.

10. The method of claim 9, wherein determining whether the signal level of the end portion is below the signal threshold level comprises determining whether an end portion of a binaurally rendered audio track exceeding the track length of the audio track is below the signal threshold level.

11. The method of claim 9, wherein determining whether the signal level of the end portion is below the signal threshold level comprises determining whether the end portion of the binaural rendering track starting at a time corresponding to an end time of the audio track remains below the signal threshold level along its track length.

12. The method of claim 9, further comprising applying reverb to the binaural rendering track, wherein the extended track length of the binaural rendering track is based at least in part on the applied reverb and the applied HRTF.

13. The method of claim 12, further comprising, in response to the signal level of the end portion not being below the signal threshold level, reducing the end portion of the binaural rendering track by fading out a portion of the reverb applied on the binaural rendering track; and The reduced end portion is trimmed so that the binaural rendering track has the same track length as the audio track.

14. The method of claim 9, further comprising fading out the binaural rendering track at a time along the extended track length corresponding to an end time of the track length of the audio track in response to the signal level of the end portion not being below the signal threshold level.

15. The method of claim 9, wherein the signal threshold level is a first signal threshold level, wherein the method further comprises: determining whether a portion of the binaural rendered track has a signal level exceeding a second signal threshold level; as well as In response to determining that the portion has a signal level exceeding the second signal threshold level, dynamic range compression is applied based on a difference between the signal level of the portion and the second signal threshold level.

16. The method of claim 15, wherein the second signal threshold level is 0 decibels relative full scale (dBFS).

17. A method performed by a programmed processor of an electronics server, the method comprising: receiving an audio track having audio content; generating a binaural rendered audio track having the same track length as the audio track and having an end portion in which reverb fades out; as well as The binaural rendered audio track is transmitted to an audio playback device over a computer network for playback.

18. The method of claim 17, wherein generating the binaural rendering audio track comprises: applying a head-related transfer function (HRTF) to the audio track to produce the binaural rendered audio track; applying reverberation to the binaural rendered audio track; as well as The end portion of the binaural rendered audio track is faded out.

19. The method of claim 18, wherein the faded-out end portion is a first end portion, wherein the binaural rendering audio track includes a second end portion starting from an end of the first end portion, wherein the method further comprises cutting off the second end portion of the binaural rendering audio track.

20. The method of claim 17, further comprising receiving a request from the audio playback device over a computer network to stream the binaural rendered audio track, wherein the binaural rendered audio track is streamed in response to the request.

Citation Information

Patent Citations

  • Method and system for music information retrieval

    US20070276733A1

  • Audio crossfading

    US20120053710A1