Audio processing method, device and storage medium

CN115701777BActive Publication Date: 2026-08-18TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280004249.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-05-31
Filing Date
2022-06-02
Publication Date
2026-08-18
Estimated Expiration
2042-06-02

Smart Images

  • Figure CN115701777B_ABST
    Figure CN115701777B_ABST
Patent Text Reader

Abstract

Aspects of the disclosure provide methods and apparatuses (e.g., client devices and server devices) for audio processing. In some examples, a client device includes processing circuitry. The processing circuitry transmits, to a server device, a selection signal indicating an audio encoding configuration for encoding audio content in an audio input. The processing circuitry receives, from the server device, an encoded bitstream in response to the transmission of the selection signal. The encoded bitstream includes the audio content that has been encoded according to the audio encoding configuration. The processing circuitry renders an audio signal based on the encoded bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation

[0002] This application claims priority to U.S. Patent Application No. 17 / 828,755, filed May 31, 2022, entitled "Adaptive Audio Transmission and Rendering," which in turn claims priority to U.S. Provisional Application No. 63 / 196,066, filed June 2, 2021, entitled "Adaptive Audio Transmission and Rendering." The entire disclosure of the earlier applications is incorporated herein by reference. Technical Field

[0003] This disclosure describes embodiments that are generally related to audio processing. Background Technology

[0004] The background description provided herein is intended to present the overall context of this disclosure. The work described by the currently identified inventors within the scope of the work described in this background section, and aspects of the description that may not be considered prior art at the time of submission, are neither explicitly nor implicitly acknowledged as prior art to this disclosure.

[0005] In virtual reality or augmented reality applications, to give users the feeling of being immersed in the application's virtual world, audio within the application's virtual scene is perceived as originating from associated virtual characters within the real world. In some examples, the user's physical movements in the real world are perceived as having matching movements within the application's virtual scene. Furthermore, and importantly, users can interact with the virtual scene using audio that is perceived as realistic and matches their experience in the real world. Summary of the Invention

[0006] Various aspects of this disclosure provide methods and apparatus for audio processing (e.g., client devices and server devices). In some examples, the client device includes processing circuitry. The processing circuitry transmits a selection signal to the server device, the selection signal indicating an audio encoding configuration for encoding audio content in an audio input. The processing circuitry receives an encoded bitstream from the server device in response to the transmission of the selection signal. The encoded bitstream includes audio content encoded according to the audio encoding configuration. The processing circuitry renders an audio signal based on the encoded bitstream.

[0007] In some embodiments, the audio encoding configuration includes a bitrate for encoding the audio content. In some examples, the audio encoding configuration includes a classification layer corresponding to portions of the audio content in the audio input.

[0008] In some examples, identifiers associated with audio encoding configurations are transmitted from the client device to the server device.

[0009] In some examples, the audio encoding configuration is determined based on at least one of the client device's media processing capabilities, the client device's network connectivity, and the user's preference input.

[0010] In some examples, the audio encoding configuration includes a bitrate used to encode the audio content. In one example, the encoded bitstream includes one or more audio channels that have been encoded according to the bitrate. In another example, the encoded bitstream includes one or more audio objects that have been encoded according to the bitrate. In yet another example, the encoded bitstream includes a set of high-order ambisonic (HOA) audio signals that have been encoded according to the bitrate.

[0011] For example, the audio encoding configuration includes a classification layer corresponding to a portion of the audio content in the audio input. In one example, the encoded bitstream is encoded based on a subset of audio channels in the audio content of the audio input. The subset of audio channels corresponds to a classification layer of the audio content in the audio input. In another example, the encoded bitstream is encoded based on a subset of audio objects in the audio content of the audio input. The subset of audio objects corresponds to a classification layer of the audio content in the audio input. In yet another example, the encoded bitstream is encoded based on a reduced-order set of HOA signals in the audio content of the audio input. The reduced-order set of HOA signals corresponds to a classification layer of the audio content in the audio input.

[0012] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform an audio processing method. Attached Figure Description

[0013] Further features, nature, and various advantages of the subject matter of this disclosure will become more apparent from the following detailed description and accompanying drawings, wherein:

[0014] Figure 1 A block diagram of a media system according to some embodiments of the present disclosure is shown.

[0015] Figure 2 The layout of a vertical three-layer sound system is shown in some examples.

[0016] Figures 3A-3C The speaker arrangements in some example audio systems are shown.

[0017] Figure 4 Examples of multiple sound sources in the sound field of some example scenarios are shown.

[0018] Figure 5 A flowchart illustrating an example of an overview process according to an embodiment of this disclosure is shown.

[0019] Figure 6A flowchart illustrating another example of a process according to an embodiment of this disclosure is shown.

[0020] Figure 7 This is a schematic diagram of a computer system according to an embodiment. Detailed Implementation

[0021] This disclosure provides techniques for adaptive audio content delivery and rendering. According to this disclosure, audio content delivery and rendering are typically limited by various factors, such as rendering device capabilities, network conditions, and user preferences. To address these limitations, adaptive audio content delivery and rendering schemes can be used.

[0022] Figure 1 A block diagram of a media system (100) according to an embodiment of the present disclosure is shown. The media system (100) can be used for various applications, such as immersive media applications, augmented reality (AR) applications, virtual reality applications, video game applications, sports game animation applications, teleconferencing and telepresence applications, media streaming applications, etc.

[0023] The media system (100) includes a media server device (110) and multiple media client devices that can be connected via a network (not shown), such as... Figure 1 The media client devices (160A) and (160B) are shown. In one example, the media server device (110) may include one or more devices with audio and video encoding capabilities. In one example, the media server device (110) includes a single computing device, such as a desktop computer, laptop computer, server computer, and tablet computer. In another example, the media server device (110) includes one or more data centers and one or more server farms. The media server device (110) can receive video and audio content and compress the video and audio content into one or more encoded bitstreams according to appropriate media encoding standards. The encoded bitstreams can be transmitted to the media client devices (160A) and (160B) via a network.

[0024] Multiple media client devices (e.g., media client devices (160A) and (160B)) each include one or more devices with video encoding and audio encoding capabilities for media applications. In one example, each of the multiple media client devices includes a computing device, such as a desktop computer, laptop computer, server computer, tablet computer, wearable computing device, and head-mounted display (HMD) device. The media client devices can decode the encoded bitstream according to a suitable media encoding standard. The decoded video content and the decoded audio content can be used for media playback.

[0025] The media server device (110) can be implemented using any suitable technology. Figure 1 In the example, the media server device (110) includes processing circuitry (130) and interface circuitry (111) coupled together.

[0026] The processing circuitry (130) may include any suitable processing circuitry system, such as one or more central processing units (CPUs), one or more graphics processing units (GPUs), and application-specific integrated circuits (ASICs). Figure 1 In one example, the processing circuitry (130) can be configured to include various encoders, such as an audio encoder (140), a video encoder (not shown), etc. In one example, one or more CPUs and / or GPUs can execute software to function as the audio encoder (140). In another example, the audio encoder (140) can be implemented using an application-specific integrated circuit (ASIC).

[0027] Interface circuit (111) can connect media server device (110) to a network. Interface circuit (111) may include a receiving part for receiving signals from the network and a transmitting part for sending signals to the network. For example, interface circuit (111) can transmit signals carrying encoded bitstreams to other devices, such as media client devices (160A), media client devices (160B), etc., via the network. Interface circuit (111) can receive signals from media client devices such as media client devices (160A) and (160B).

[0028] The network is suitably coupled to the media server device (110) and media client devices (e.g., media client devices (160A) and (160B)) via wired and / or wireless connections (e.g., Ethernet, fiber optic, WiFi, cellular, etc.). The network may include network server devices, memory devices, and network devices, etc. The components of the network are suitably coupled together via wired and / or wireless connections.

[0029] Media client devices (e.g., media client devices (160A) and (160B)) are configured to decode the encoded bitstream. In one example, each media client device can perform video decoding to reconstruct a displayable sequence of video frames and can perform audio decoding to generate an audio signal for playback.

[0030] Media client devices such as the (160A) and (160B) can be implemented using any suitable technology. Figure 1The example shows a media client device (160A), but is not limited to a head-mounted display (HMD) with headphones that can be used by user A. Figure 1 The example also shows a media client device (160B), but is not limited to the smartphone used by user B.

[0031] exist Figure 1 In the middle, the media client device (160A) includes, for example, Figure 1 The interface circuitry (161A) and processing circuitry (170A) coupled together, as well as the media client device (160B), include, as shown... Figure 1 The interface circuit (161B) and processing circuit (170B) are shown coupled together.

[0032] The interface circuit (161A) connects the media client device (160A) to a network. The interface circuit (161A) may include a receiving section for receiving signals from the network and a transmitting section for sending signals to the network. For example, the interface circuit (161A) can receive signals carrying data from the network, such as signals carrying encoded bitstreams.

[0033] The processing circuitry (170A) may include a suitable processing circuitry system, such as a CPU, GPU, and application-specific integrated circuits (ASICs). The processing circuitry (170A) may be configured to include various components, such as an audio decoder (171A), a renderer (172A), etc.

[0034] In some examples, the audio decoder (171A) can decode the audio content in the encoded bitstream by selecting a decoding tool suitable for the encoding scheme of the audio content. Furthermore, the renderer (172A) can generate a final digital product suitable for a media client device (160A) based on the audio content decoded from the encoded bitstream. It should be noted that the processing circuitry (170A) may include other suitable components (not shown) for further audio processing, such as a mixer, post-processing circuitry, etc.

[0035] Similarly, the interface circuit (161B) can connect the media client device (160B) to the network. The interface circuit (161B) may include a receiving section for receiving signals from the network and a transmitting section for sending signals to the network. For example, the interface circuit (161B) can receive signals carrying data from the network, such as signals carrying encoded bitstreams.

[0036] The processing circuitry (170B) may include a suitable processing circuitry system, such as a CPU, GPU, and application-specific integrated circuits (ASICs). The processing circuitry (170B) may be configured to include various components, such as an audio decoder (171B), a renderer (172B), etc.

[0037] In some examples, the audio decoder (171B) can decode the audio content in the encoded bitstream by selecting a decoding tool suitable for the audio content by the encoding scheme. Furthermore, the renderer (172B) can generate a final digital product suitable for a media client device (160B) based on the audio content decoded from the encoded bitstream. It should be noted that the processing circuitry (170A) may include other suitable components (not shown) for further audio processing, such as a mixer, post-processing circuitry, etc.

[0038] According to one aspect of this disclosure, media client devices can have different media processing capabilities, such as different CPU configurations and different memory configurations. For the same encoded bitstream, some media client devices can render audio from the encoded bitstream without any problems; however, some media client devices may fail to render audio successfully due to insufficient processing power. According to another aspect of this disclosure, network conditions such as bandwidth and latency may also affect rendering. Furthermore, media client device users may prefer personalization and have preferences regarding how audio is rendered.

[0039] According to some aspects of this disclosure, the media system (100) is configured with adaptive audio transmission and rendering technology. Adaptive audio transmission and rendering technology can adjust audio transmission and rendering while taking into account various constraints such as media processing capacity constraints, network condition constraints, and user preference constraints, thereby optimizing the auditory experience.

[0040] According to some aspects of this disclosure, audio input can be encoded into an encoded bitstream with different audio encoding configurations. A media server device (110) and / or a media client device can select an encoded bitstream with a suitable audio encoding configuration for the media client device based on various constraints, and the encoded bitstream can be transmitted to the media client device, and the media client device can render the audio output based on the encoded bitstream.

[0041] In some embodiments, the media server device (110) is configured to select appropriate audio encoding configurations for each of the multiple media client devices. In some examples, the processing circuitry (130) includes an adaptive controller (135) configured to select appropriate audio encoding configurations for each of the multiple media client devices.

[0042] In some examples, a media server device (110) receives audio input from an audio source (101) (e.g., an audio injection server in the example). An audio encoder (140) can encode the audio input into an encoded bitstream with different audio encoding configurations. The audio encoding configuration may include one or more parameters that affect audio encoding, such as bitrate, classification layer, etc.

[0043] In some examples, the audio encoding configuration has different bitrates, and the audio input is encoded into the encoded bitstream according to the different bitrates. In some examples, the audio encoding configuration has different classification layers, and the audio input is encoded into the encoded bitstream according to the different classification layers. In some examples, the audio encoding configuration may include both bitrate and classification layers. The audio encoding configuration has different bitrates and / or different classification layers, and the audio input is encoded into the encoded bitstream according to the different bitrates and / or different classification layers.

[0044] In some video-on-demand (VOD) applications, a media server device (110) can encode the audio content of the entire program according to different audio encoding configurations and can store the encoded bitstreams. Typically, the media server device (110) can be configured to have a relatively large storage capacity (compared to media client devices) to store encoded bitstreams with different audio encoding configurations. For example, encoded bitstreams with different audio encoding configurations can be adaptively provided to the corresponding media client devices based on their respective media processing capabilities, network conditions, and user preferences.

[0045] In some real-time streaming applications, a media server device (110) can receive portions of the audio content of a program in real time and encode those portions according to different audio encoding configurations. The encoded streams can be buffered. Typically, the media server device (110) can be configured to have relatively large media processing capabilities (compared to media client devices) to encode portions of the audio content in real time according to different audio encoding configurations, and the media server device (110) can be configured to have relatively large storage capabilities (compared to media client devices) to buffer encoded streams with different audio encoding configurations. For example, encoded streams with different audio encoding configurations can be adaptively provided to the respective media client devices based on their respective media processing capabilities, network conditions, and user preferences.

[0046] For example, in Figure 1In the example, the first encoded bitstream is encoded based on a first audio encoding configuration such as the lowest bitrate, lowest classification layer, and lowest quality; the second encoded bitstream is encoded based on a second audio encoding configuration such as an intermediate bitrate, intermediate classification layer, and intermediate quality; and the Nth encoded bitstream is encoded based on an Nth audio encoding configuration such as the highest bitrate, highest classification layer, and highest quality.

[0047] In some examples, the adaptive controller (135) considers one or more constraints associated with the media client device (e.g., media processing capability constraints, network condition constraints, user preference constraints, etc.) and selects one of a plurality of encoded bitstreams for the media client device. The selected encoded bitstream is then transmitted to the media client device, for example, via a network. In some examples, one or more of the constraints may change, and in response to a change in constraints, the adaptive controller (135) may decide to switch to another encoded bitstream and transmit that other encoded bitstream to the media client device.

[0048] In one example, the media client device (160A) is the VR device used by user A in a game application. This VR device is configured with sufficient processing power for both video and audio, and the game application prefers high-quality audio for a better user experience. The adaptive controller (135) obtains the configuration of the media client device (160A) and network condition information. The configuration of the media client device (160A) indicates sufficient processing power for audio processing, thus eliminating processing power constraints, and the network condition information indicates sufficient bandwidth and no network connectivity constraints. The adaptive controller (135) can then select the Nth encoded bitstream of the Nth audio encoding configuration to transmit to the media client device (160A).

[0049] In one example, the media client device (160B) is a smartphone used by user B at an airport during a conference call. This smartphone may have limited processing power for video and audio, and for user experience, high-quality audio is not required for the conference call. The adaptive controller (135) obtains the configuration of the media client device (160B) and network condition information. The configuration of the media client device (160B) indicates limited processing power for audio, and the network condition information indicates limited bandwidth at the airport. The adaptive controller (135) can then select a first encoded bitstream with a first audio encoding configuration to transmit to the media client device (160B).

[0050] In some embodiments, the media client device may select a suitable audio encoding configuration based on various constraints and may notify / request the media server device (110) accordingly. The media server device (110) then transmits the encoded bitstream, encoded using the suitable audio encoding configuration, to the media client device. In some examples, when one or more constraints change, the media client device may determine to switch to another audio encoding configuration and notify the media server device (110) accordingly. The media server device (110) then transmits another encoded bitstream, encoded according to the other audio encoding configuration, to the media client device.

[0051] exist Figure 1 In the example, the media client device (160A) includes an adaptive controller (175A) configured to select an appropriate audio encoding configuration based on various constraints associated with the media client device (160A). The media client device (160B) includes an adaptive controller (175B) configured to select an appropriate audio encoding configuration based on various constraints associated with the media client device (160B).

[0052] In one example, the adaptive controller (175A) obtains the configuration of the media client device (160A) and network condition information. The configuration of the media client device (160A) indicates sufficient processing power for audio processing, thus without processing power constraints, and the network condition information indicates sufficient bandwidth and no network connectivity constraints. The adaptive controller (175A) can then select, for example, an Nth audio encoding configuration.

[0053] In one example, the adaptive controller (175B) can obtain the configuration of the media client device (160B) and network condition information. The configuration of the media client device (160B) indicates limited processing power for audio processing, and the network condition information indicates limited bandwidth at the airport. The adaptive controller (175B) can then select, for example, a first audio encoding configuration.

[0054] According to some aspects of this disclosure, the audio input injected into the media client server (110) can have various formats for transmission and reproduction, such as audio channels, audio objects, high-order stereo (HOA) signal sets, or combinations of two or more audio channels, audio objects, and HOA signal sets.

[0055] According to one aspect of this disclosure, the audio content of a scene can be in the format of an audio channel associated with a position in the sound field of the scene. For example, an audio channel can be associated with speakers in a sound system. The sound system can have various multi-channel configurations. In some examples, the speakers in the sound system can be arranged to surround the audience in three vertical layers, referred to as the upper, middle, and lower layers.

[0056] Figure 2 The layout of three vertically tiered speakers surrounding the audience is shown.

[0057] According to one aspect of this disclosure, multi-channel format audio content includes multiple audio channels for positioning in a sound field.

[0058] Figures 3A-3C The speaker arrangement in the upper, middle, and lower layers of an audio system is shown. This audio system is represented by a 22.2-channel audio system capable of playing 22.2-channel audio content. The 22.2-channel audio content comprises 24 audio channels. In one example, the 24 audio channels may correspond to 24 speaker locations in the audio system. The 24 audio channels include two low-frequency effect (LFE) channels. Figures 3A-3C The small squares indicate the speaker locations, and the numbers in the small squares are the indices of the speaker locations. Figure 3A The speaker arrangement on the upper level is shown. Figure 3B The speaker arrangement in the middle layer is shown. Figure 3C The speaker arrangement on the lower layer is shown. In one example, positions 23 and 24 could be used for two LFE channels.

[0059] Some audio systems can have fewer speakers, and 22.2 multi-channel audio content can be down-mixed to form audio content with fewer audio channels.

[0060] In one example, an audio system represented by a 2.0 multi-channel sound system may include two speaker positions, and 22.2 multi-channel audio content may be downmixed to form 2.0 multi-channel audio content, which includes two audio channels corresponding to the two speaker positions. In another example, an audio system represented by a 5.1 multi-channel audio system may include six speaker positions, and 22.2 multi-channel audio content may be downmixed to form 5.1 multi-channel audio content, which includes six audio channels corresponding to the six speaker positions. In yet another example, an audio system represented by a 9.2 multi-channel audio system may include 11 speaker positions, and 22.2 multi-channel audio content may be downmixed to form 9.2 multi-channel audio content, which includes 11 audio channels corresponding to the 11 speaker positions.

[0061] It should be noted that audio content with fewer channels can be represented by fewer bits, and audio content with fewer channels can request fewer transmission and rendering resources.

[0062] According to another aspect of this disclosure, the audio content of a scene can be in the format of multiple audio objects associated with sound sources in the sound field of the scene.

[0063] Figure 4 An example of multiple sound sources (411)-(415) in a sound field of a VR application scenario is shown. The audio content of the scenario may include multiple audio objects, each for a sound source (411)-(415).

[0064] In another example, a hospital audio scene could have a sound field setup similar to that in a doctor's office. This sound field could include a doctor, a patient, a television, a radio, a door, a table, and a chair as sound sources. Therefore, the audio content of this scene could include seven audio objects, each corresponding to a sound source. For example, the first audio object might correspond to the doctor's voice, the second to the patient's voice, the third to the television's voice, the fourth to the radio's voice, the fifth to the door's voice, the sixth to the table's voice, and the seventh to the chair's voice.

[0065] According to another aspect of this disclosure, the audio content of the scene may be in the format of an HOA set.

[0066] Stereo (Ambisonic) is an omnidirectional surround sound format. In addition to the horizontal plane, stereo covers sound sources above and below the listener. The transmission channel for stereo does not carry speaker signals. Instead, the transmission channel includes a speaker-independent representation of the sound field, called B-format, which is then decoded according to the speaker setup. Stereo allows for reproduction to be considered from the perspective of source direction rather than speaker location, and it provides listeners with considerable flexibility in the arrangement and number of speakers used for playback.

[0067] In one example, first-order stereo can be understood as a three-dimensional extension of center (M) stereo / side (S) stereo, adding additional difference channels for height and depth. The resulting signal set is called the B format, which includes four component channels labeled W for sound pressure (M in M / S), X for front-to-back sound pressure gradient, Y for left-to-right (S in M / S), and Z for top-to-bottom.

[0068] The spatial resolution of first-order stereo can be improved by using higher-order stereo. For example, first-order stereo has a slightly blurry source, but a relatively small usable listening area or optimal listening point. By adding a set with more selective directional components to the B-format, the spatial resolution can be increased and the optimal point expanded. The resulting signal set is called second-order stereo, third-order stereo, or collectively higher-order stereo (HOA). Typically, a higher-order stereo set includes more selective directional components than the lower-order stereo set.

[0069] According to some aspects of this disclosure, the audio input of the media server device (110) can be encoded at several different bitrates (corresponding to audio encoding configurations). In some examples, the media client server (110) can select or switch between encoded streams at different bitrates. In some examples, media client devices such as media client devices (160A) and (160B) can select or switch between encoded streams at different bitrates. For example, the selection or switching may depend on available resources (e.g., processing power, network bandwidth) and / or user preferences.

[0070] In some embodiments, the audio input includes audio content in an audio channel format. The audio channels are encoded at different bitrates. For example, an audio channel is encoded at a first bitrate (corresponding to a first audio encoding configuration) to form a first encoded bitstream, an audio channel is encoded at a second bitrate (corresponding to a second audio encoding configuration) to form a second encoded bitstream, and so on. In some examples, a media client server (110) can select or switch between encoded bitstreams at different bitrates (corresponding to different audio encoding configurations). In some examples, media client devices such as media client devices (160A) and (160B) can select or switch between encoded bitstreams at different bitrates (corresponding to different audio encoding configurations). For example, the selection or switching may depend on available resources (e.g., processing power, network bandwidth) and / or user preferences.

[0071] In some embodiments, the audio input includes audio content in an audio object format. The audio object is encoded at several different bitrates. For example, the audio object is encoded at a first bitrate (corresponding to a first audio encoding configuration) to form a first encoded bitstream, and the audio object is encoded at a second bitrate (corresponding to a second audio encoding configuration) to form a second encoded bitstream at the second bitrate. In some examples, the media client server (110) can select or switch between encoded bitstreams at different bitrates (corresponding to different audio encoding configurations). In some examples, media client devices such as media client devices (160A) and (160B) can select or switch between encoded bitstreams at different bitrates (corresponding to different audio encoding configurations). For example, the selection or switching may depend on available resources (e.g., processing power, network bandwidth) and / or user preferences.

[0072] In some embodiments, the audio input includes audio content in HOA signal set format, such as second-order stereo signal set, third-order stereo signal set, fourth-order stereo signal set, etc. The HOA format audio content is encoded at different bitrates. For example, HOA format audio content is encoded at a first bitrate (corresponding to a first audio encoding configuration) to form a first encoded bitstream, and HOA format audio content is encoded at a second bitrate (corresponding to a second audio encoding configuration) to form a second encoded bitstream at the second bitrate. In some examples, the media client server (110) can select or switch between encoded bitstreams with different bitrates (corresponding to different audio encoding configurations). In some examples, media client devices such as media client devices (160A) and (160B) can select or switch between encoded bitstreams with different bitrates (corresponding to different audio encoding configurations). For example, the selection or switching may depend on available resources (e.g., processing power, network bandwidth) and / or user preferences.

[0073] In some embodiments, a quality identifier (ID) is assigned a bitrate. A media server device (110) or content creator can use the quality ID to indicate which bitrate is used to encode the audio input into an encoded stream for transmission. Media client devices, such as media client devices (160A) or (160B), can request a specific quality ID based on available resources (e.g., processing power, network bandwidth) and / or user preferences.

[0074] It's important to note that the audio content in an audio scene can be a mixed format of audio channels, audio objects, HOAs, etc. In some examples, when the audio content is a mixed format of two or more of the audio channels, audio objects, and HOAs, multiple coding bitrates can be applied separately to the audio channels, audio objects, or HOA signals. In other examples, when the audio content is a mixed format of two or more of the audio channels, audio objects, and HOAs, multiple coding bitrates can be applied to the combination of the audio channels, audio objects, and HOA signals.

[0075] According to some aspects of this disclosure, the audio content in the audio input of a media server device (110) can be classified into several classification layers. In some examples, each classification layer may include a portion of the audio content in the audio input. In some examples, a higher classification layer may include an additional portion of the audio content in the lower classification layer and the audio input. Thus, the classification layer can be a parameter in the audio encoding configuration. In some examples, the media client server (110) can select or switch between encoded bitstreams of different classification layers (corresponding to multiple audio encoding configurations). In some examples, media client devices such as media client devices (160A) and (160B) can select or switch between encoded bitstreams of different classification layers (corresponding to multiple audio encoding configurations). For example, the selection or switching may depend on available resources (e.g., processing power, network bandwidth) and / or user preferences, etc.

[0076] In some embodiments, the audio input includes audio content in an audio channel format. Audio channels can be categorized into several classification layers.

[0077] For example, the audio input includes audio content in the 22.2 multichannel audio content format. In one example, the 22.2 multichannel audio content can be categorized into four layers: a first layer of 2.0 multichannel audio content, a second layer of 5.1 multichannel audio content, a third layer of 9.2 multichannel audio content, and a fourth layer of 22.2 multichannel audio content. The 2.0 multichannel audio content can be encoded as a first encoded bitstream (with a first audio encoding configuration), the 5.1 multichannel audio content can be encoded as a second encoded bitstream (with a second audio encoding configuration), the 9.2 multichannel audio content can be encoded as a third encoded bitstream (with a third audio encoding configuration), and the 22.2 multichannel audio content can be encoded as a fourth encoded bitstream (with a fourth audio encoding configuration).

[0078] In some examples, the media client server (110) can select or switch between encoded bitstreams at different classification layers. In some examples, media client devices such as media client devices (160A) and (160B) can select or switch between encoded bitstreams at different classification layers. For example, the selection or switching may depend on available resources (e.g., processing power, network bandwidth) and / or user preferences.

[0079] The above description is an example of audio channel classification. It should be noted that in some examples, 22.2 multi-channel audio content may be classified differently than described above.

[0080] In another embodiment, audio objects are categorized into several classification layers. Taking a hospital audio scene as an example, the audio content of the hospital audio scene may include seven audio objects for multiple sound sources: a first audio object corresponding to the doctor's voice, a second audio object corresponding to the patient's voice, a third audio object corresponding to the television sound, a fourth audio object corresponding to the radio sound, a fifth audio object corresponding to the door sound, a sixth audio object corresponding to the table sound, and a seventh audio object corresponding to the chair sound.

[0081] In one example, seven audio objects can be categorized into three levels. The first level includes a first audio object corresponding to a doctor's voice and a second audio object corresponding to a patient's voice. The second level includes a first audio object corresponding to a doctor's voice, a second audio object corresponding to a patient's voice, a third audio object corresponding to a television sound, and a fourth audio object corresponding to a radio sound. The third level includes a first audio object corresponding to a doctor's voice, a second audio object corresponding to a patient's voice, a third audio object corresponding to a television sound, a fourth audio object corresponding to a radio sound, a fifth audio object corresponding to a sound from the opposite door, a sixth audio object corresponding to a sound from the table, and a seventh audio object corresponding to a sound from the chair.

[0082] The first classification layer can be encoded as a first encoded bitstream (with a first audio encoding configuration), the second classification layer can be encoded as a second encoded bitstream (with a second audio encoding configuration), and the third classification layer can be encoded as a third encoded bitstream (with a third audio encoding configuration). In some examples, the media client server (110) can select or switch between encoded bitstreams of different classification layers. In some examples, media client devices such as media client devices (160A) and (160B) can select or switch between encoded bitstreams of different classification layers. For example, the selection or switching may depend on available resources (e.g., processing power, network bandwidth) and / or user preferences.

[0083] The above description is an example of audio object classification. It should be noted that in some examples, audio scenes containing audio objects may be classified differently than described above.

[0084] In another embodiment, HOA signals are classified into several classification layers according to different orders. In one example, the fourth-order HOA signal set can be classified into four classification layers. The first classification layer includes the first-order HOA signal set. The second classification layer includes the second-order HOA signal set. The third classification layer includes the third-order HOA signal set. The fourth classification layer includes the fourth-order HOA signal set.

[0085] The first classification layer can be encoded as a first encoded bitstream (with a first audio encoding configuration), the second classification layer can be encoded as a second encoded bitstream (with a second audio encoding configuration), the third classification layer can be encoded as a third encoded bitstream (with a third audio encoding configuration), and the fourth classification layer can be encoded as a fourth encoded bitstream (with a fourth audio encoding configuration). In some examples, the media client server (110) can select or switch between encoded bitstreams of different classification layers (corresponding to different audio encoding configurations). In some examples, media client devices such as media client devices (160A) and (160B) can select or switch between encoded bitstreams of different classification layers (corresponding to different audio encoding configurations). For example, the selection or switching may depend on available resources (e.g., processing power, network bandwidth) and / or user preferences, etc.

[0086] The above description is an example of HOA classification. It should be noted that in some examples, HOA signals may be classified differently from those described above.

[0087] In some embodiments, a layer identifier (ID) can be assigned to the classification layer of the audio input. Server devices or content creators can use the layer ID to indicate which layer(s) of the audio input is being transmitted; client devices can request a specific layer ID based on available resources and / or user preferences.

[0088] It's important to note that the audio content in an audio scene can be a mixture of audio channels, audio objects, HOAs, etc. In some examples, when the audio content is a mixture of two or more of these formats, the classification layer can be determined based on the audio channel, audio object, or HOA signal individually. In other examples, when the audio content is a mixture of two or more of these formats, the classification layer can be determined based on a combination of the audio channel, audio object, and HOA signals.

[0089] Figure 5A flowchart of an overview process (500) according to an embodiment of the present disclosure is shown. The process (500) can be used in a client device for audio processing, such as in media client devices (160A) and (160B), and is executed by processing circuits (170A) and (170B), etc. In some embodiments, the process (500) is implemented as software instructions, so that the processing circuit executes the process (500) when the software instructions are executed. The process begins at (S501) and proceeds to (S510).

[0090] At (S510), the client device transmits a selection signal. This selection signal indicates the audio encoding configuration used to encode the audio content in the audio input.

[0091] In some examples, the audio encoding configuration includes a bitrate for encoding the audio content. In some examples, the audio encoding configuration includes a classification layer corresponding to a portion of the audio content in the audio input.

[0092] In one example, the transmission includes identifiers associated with the audio encoding configuration (such as quality identifiers and classification identifiers).

[0093] In one example, the selection signal is determined based on at least one of the client device's media processing capabilities, the client device's network connectivity, and the user's preference input on the client device.

[0094] At (S520), the encoded bitstream is received in response to the transmission of the selection signal. The encoded bitstream includes audio content that has been encoded according to the audio encoding configuration.

[0095] In some examples, the audio encoding configuration includes a bitrate. In one example, the encoded bitstream includes multiple audio channels that have been encoded according to the bitrate. In another example, the encoded bitstream includes multiple audio objects that have been encoded according to the bitrate. In yet another example, the encoded bitstream includes a set of high-order stereo (HOA) audio signals that have been encoded according to the bitrate.

[0096] In some examples, the audio encoding configuration includes a classification layer. In one example, the encoded bitstream includes a subset of audio channels from the audio input content (encoded based on this subset of audio channels). This subset of audio channels corresponds to the classification layer. In another example, the encoded bitstream includes a subset of audio objects from the audio input content (encoded based on this subset of audio objects). This subset of audio objects corresponds to the classification layer. In yet another example, the encoded bitstream includes a reduced-order set of HOA signals from the audio input content (encoded based on this reduced-order set of HOA signals). This reduced-order set of HOA signals corresponds to the classification layer.

[0097] At (S530), the audio signal is rendered based on the encoded bitstream. Then, the process proceeds to (S599) and terminates.

[0098] The process (500) can be adjusted as appropriate. One or more steps in the process (500) can be modified and / or omitted. Additional steps (one or more) can be added. Any suitable implementation order can be used.

[0099] Figure 6 A flowchart of an overview process (600) according to an embodiment of this disclosure is shown. This process (600) can be used in a server device for audio processing, such as a media server device (110), and is executed by processing circuitry (130), etc. In some embodiments, the process (600) is implemented as software instructions, so that the processing circuitry executes the process (600) when the software instructions are executed. The process begins at S601 and proceeds to S610.

[0100] At (S610), the server device determines an audio encoding configuration for the client device (e.g., media client device (160A), media client device (160B), etc.) to encode the audio content in the audio input.

[0101] In some examples, the audio encoding configuration includes a bitrate for encoding the audio content. In some examples, the audio encoding configuration includes a classification layer corresponding to a portion of the audio content in the audio input.

[0102] In some examples, the server device determines the audio encoding configuration based on at least one of the client device's media processing capabilities, the client device's network connectivity, and preference input.

[0103] At (S620), the server device obtains the encoded bitstream, which includes audio content that has been encoded according to the audio encoding configuration.

[0104] In some examples, the audio encoding configuration includes a bitrate. In one example, the encoded bitstream includes multiple audio channels that have been encoded according to the bitrate. In another example, the encoded bitstream includes multiple audio objects that have been encoded according to the bitrate. In yet another example, the encoded bitstream includes a set of high-order stereo (HOA) audio signals that have been encoded according to the bitrate.

[0105] In some examples, the audio encoding configuration includes a classification layer. In one example, the encoded bitstream includes a subset of audio channels from the audio input content (encoded based on this subset of audio channels). This subset of audio channels corresponds to the classification layer. In another example, the encoded bitstream includes a subset of audio objects from the audio input content (encoded based on this subset of audio objects). This subset of audio objects corresponds to the classification layer. In yet another example, the encoded bitstream includes a reduced-order set of HOA signals from the audio input content (encoded based on this reduced-order set of HOA signals). This reduced-order set of HOA signals corresponds to the classification layer.

[0106] At (S630), the encoded bitstream is transmitted to the client device. In some examples, the server device also transmits an identifier (ID) (e.g., a quality identifier, a classification layer identifier, etc.) that indicates the audio encoding configuration used to encode the audio content of the audio input.

[0107] Then, the process proceeds to (S699) and terminates.

[0108] The process (600) can be adjusted as appropriate. One or more steps in the process (600) can be modified and / or omitted. Additional steps (one or more) can be added. Any suitable implementation order can be used.

[0109] The above techniques can be implemented as computer software, which uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 7 A computer system (700) suitable for implementing certain embodiments of the subject matter of this disclosure is shown.

[0110] Computer software can be encoded using any suitable machine code or computer language, which can be assembled, compiled, linked, or similar mechanisms to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through decoding, microcode execution, etc.

[0111] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things (IoT) devices.

[0112] Figure 7The components of the computer system (700) shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any component or combination of components shown in the exemplary embodiments of the computer system (700).

[0113] The computer system (700) may include certain human-machine interface input devices. These devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0114] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard (701), mouse (702), touchpad (703), touch screen (710), data glove (not shown), joystick (705), microphone (706), scanner (707), camera (708).

[0115] The computer system (700) may include certain human-machine interface output devices. These human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, smell / taste. These human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touchscreen (710), data gloves (not shown), or joysticks (705), but may also be tactile feedback devices not used as input devices), audio output devices (e.g., speakers (709), headphones (not depicted)), visual output devices (e.g., screens (710) including CRT screens, LCD screens, plasma screens, OLED screens, each screen may or may not have touchscreen input capability, each screen may or may not have tactile feedback capability, some of which are capable of outputting two-dimensional visual output or more than three-dimensional output through devices such as stereoscopic image output, virtual reality glasses (not described), holographic displays and smoke boxes (not described), and printers (not described).

[0116] The computer system (700) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (720) having media such as CD / DVD (721), finger drives (722), removable hard disk drives or solid-state drives (723), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), and devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted).

[0117] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter currently disclosed does not cover transmission media, carrier waves, or other transient signals.

[0118] The computer system (700) may also include interfaces to one or more communication networks (755). These networks may be, for example, wireless, wired, or optical. Networks may also be local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), vehicle and industrial networks, real-time networks, latency-tolerant networks, etc. Examples of networks include LANs such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless WAN digital networks including cable television, satellite television, and terrestrial broadcast television, and vehicle and industrial television networks including CANBus, etc. Some networks typically require external network interface adapters (e.g., USB ports on the computer system (700)) attached to certain general-purpose data ports or peripheral buses (749); other network interfaces are typically integrated into the core of the computer system (700) by being attached to system buses as described below (e.g., Ethernet interfaces attached to PC computer systems or cellular interfaces to smartphone computer systems). The computer system (700) can use any of these networks to communicate with other entities. This communication can be one-way receiving (e.g., broadcast television), one-way sending (e.g., to a CANbus device), or two-way (e.g., to other computer systems using a local area network or a wide area digital network). As mentioned above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.

[0119] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel (740) of the computer system (700).

[0120] The core (740) may include one or more central processing units (CPU) (741), graphics processing units (GPUs) (742), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs, 743), hardware accelerators (744) for certain tasks, and graphics adapters (750), etc. These devices may be connected to read-only memory (ROM) (745), random access memory (746), and internal mass storage (747) such as internal non-user-accessible hard disk drives, SSDs, etc., via a system bus (748). In some computer systems, the system bus (748) may be accessed in the form of one or more physical plugs to allow for expansion by adding CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (748) or attached to the core's system bus (748) via a peripheral device bus (749). In one example, a screen (710) may be connected to a graphics adapter (750). Peripheral bus architectures include PCI, USB, etc.

[0121] The CPU (741), GPU (742), FPGA (743), and accelerator (744) can execute certain instructions, which, when combined, constitute the aforementioned computer code. The computer code can be stored in ROM (745) or RAM (746). Transient data can also be stored in RAM (746), while permanent data can be stored, for example, in internal mass storage (747). Fast storage and retrieval of any storage device can be achieved using a cache memory, which can be closely associated with one or more CPUs (741), GPUs (742), mass storage (747), ROM (745), RAM (746), etc.

[0122] Computer-readable media may have computer code thereon for performing various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0123] As a non-limiting example, a computer system having architecture 700, particularly kernel 740, can be made functional by the execution of software contained in one or more tangible computer-readable media by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be media associated with user-accessible mass storage as described above, or it can be some memory of the non-transitory kernel (740), such as internal kernel mass storage (747) or ROM (745). Software implementing various embodiments of this disclosure can be stored in such a device and executed by the kernel (740). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the kernel (740), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (746) and modifying these data structures according to processes defined by the software. Additionally or alternatively, the computer system may provide functionality through hard-wired or otherwise embodied logic in circuitry (e.g., accelerator 744), which may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0124] Although several exemplary embodiments have been described in this disclosure, there are changes, substitutions, and various alternative equivalents that fall within the scope of this disclosure. Therefore, it is understood that those skilled in the art will be able to design numerous systems and methods that, although not explicitly shown or described herein, embody the principles of this disclosure and are therefore within its spirit and scope.

Claims

1. A method for audio processing at a client device, characterized in that, include: A selection signal is transmitted to the server device, the selection signal indicating an audio encoding configuration for encoding audio content in the audio input, the audio encoding configuration including a classification layer corresponding to a portion of the audio content in the audio input, a subset of audio channels corresponding to a classification layer of the audio content in the audio input, and each classification layer being assigned a corresponding layer identifier; Receive an encoded bitstream from the server device, the encoded bitstream including audio content encoded according to the audio encoding configuration in response to the transmission of the selection signal; as well as Render the audio signal based on the encoded bitstream.

2. The method according to claim 1, characterized in that, Transmitting the selection signal further includes: The transmission indicator is the selection signal used to specify the bitrate for encoding the audio content.

3. The method according to claim 2, characterized in that, Receiving the encoded bitstream further includes: Receive the encoded bitstream, which includes one or more audio channels encoded according to the bitrate.

4. The method according to claim 2, characterized in that, Receiving the encoded bitstream further includes: Receive the encoded bitstream, which includes one or more audio objects encoded according to the bitrate.

5. The method according to claim 2, characterized in that, Receiving the encoded bitstream further includes: Receive the encoded bitstream comprising an audio high-order stereo HOA signal encoded according to the bitrate.

6. The method according to claim 1, characterized in that, Receiving the encoded bitstream further includes: Receive the encoded bitstream, which is a subset of audio channels in the audio content based on the audio input.

7. The method according to claim 1, characterized in that, Receiving the encoded bitstream further includes: Receive the encoded bitstream, which is a subset of audio objects in the audio content based on the audio input.

8. The method according to claim 1, characterized in that, Receiving the encoded bitstream further includes: The encoded bitstream is received by encoding a reduced set of high-order stereo HOA signals from the audio content based on the audio input.

9. The method according to any one of claims 1 to 8, characterized in that, Transmitting the selection signal further includes: Transmit the identifier associated with the audio encoding configuration.

10. The method according to claim 1, characterized in that, Also includes: The selection signal is determined based on at least one of the following: the media processing capabilities of the client device, the network connectivity of the client device, and preference input.

11. A device for audio processing, characterized in that, Including processor and memory, The memory is used to store program instructions; The processor is used to invoke the program instructions stored in the memory to implement the method of any one of claims 1 to 10.

12. A device for audio processing, characterized in that, The device includes: The transmission module is configured to transmit a selection signal to a server device, the selection signal indicating an audio encoding configuration for encoding audio content in an audio input, the audio encoding configuration including a classification layer corresponding to a portion of the audio content in the audio input, a subset of audio channels corresponding to a classification layer of the audio content in the audio input, and each classification layer being assigned a corresponding layer identifier. A receiving module is configured to receive an encoded bitstream from the server device, the encoded bitstream including audio content encoded according to the audio encoding configuration in response to the transmission of the selection signal; and The rendering module is configured to render audio signals based on the encoded bitstream.

13. A computer-readable storage medium, characterized in that, Includes instructions that, when run on a computer, cause the computer to perform the method of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Transporting coded audio data

    CN107925797A

  • Adaptive audio processing method, device, computer program, and recording medium thereof in wireless communication system

    WO2021015484A1