Audio processing method and apparatus

By adjusting the loudness of the voice signal using audio codecs and binaural rendering tools, the problem of mismatch between audio signals and virtual scenes in virtual reality or augmented reality is solved, thereby improving the user's immersive experience and interactivity.

CN115668369BActive Publication Date: 2026-04-24TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2021-10-07
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In virtual reality or augmented reality applications, existing technologies struggle to match audio signals with virtual scenes, resulting in an unrealistic user experience and an inability to effectively adjust the sound volume in the scene to match the user's real-world experience.

Method used

The audio codec processing circuit decodes the adjusted speech signal and loudness adjustment information. Based on the loudness adjustment of the speech signal, the loudness of multiple sound signals in the scene is adjusted. The binaural rendering tool is used to simulate the audio environment and generate the scene output signal for loudness matching.

Benefits of technology

It enables the matching of audio signals with virtual scenes in virtual reality or augmented reality applications, enhancing the user's immersive experience and making the sound in the virtual scene as if it were in the real world, thus enhancing the interactivity between the user and the scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668369B_ABST
    Figure CN115668369B_ABST
Patent Text Reader

Abstract

Aspects of the disclosure provide methods and apparatuses for audio processing. In some examples, an audio coding apparatus includes processing circuitry. The processing circuitry decodes, from an encoded bitstream, information indicative of an adjusted speech signal and a loudness adjustment to the adjusted speech signal. The adjusted speech signal is indicated in a manner associated with a plurality of speech signals in a scene of an immersive media application. The processing circuitry determines, based on a plurality of loudness adjustments to the adjusted speech signal, a plurality of loudness adjustments to a plurality of sound signals including the plurality of speech signals in the scene, and generates, based on the plurality of loudness adjustments to the plurality of sound signals, the plurality of sound signals in the scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation

[0002] This application claims priority to U.S. Patent Application No. 17 / 450,015, filed October 5, 2021, entitled “METHOD AND APPARATUS IN AUDIOPROCESSING”, which claims priority to U.S. Provisional Application No. 63 / 152,086, filed February 22, 2021, entitled “Scene Loudness Adjustment”. The entire disclosure of both prior applications is incorporated herein by reference. Technical Field

[0003] This disclosure describes embodiments that generally involve audio processing. Background Technology

[0004] The background description provided herein is for the purpose of presenting the general context of this disclosure. Within the scope described in this background section, neither the work of the currently named inventors nor any aspect of this description that does not qualify as prior art at the time of submission is expressly or implicitly acknowledged as prior art to this invention.

[0005] In virtual reality or augmented reality applications, to create an immersive experience for users within the application's virtual world, audio within the application's scene is perceived as if it were in the real world, with sounds originating from relevant virtual objects within the scene. In some examples, a user's physical movements in the real world are perceived as matching movements within the application's virtual scene. Furthermore, and importantly, users can interact with the virtual scene using realistic audio that matches their real-world experience. Summary of the Invention

[0006] This disclosure provides methods and apparatus for audio processing. In some examples, the audio encoding / decoding apparatus includes processing circuitry. The processing circuitry decodes from an encoded bitstream information indicating an adjusted speech signal and loudness adjustments to the adjusted speech signal. The adjusted speech signal is indicated in a manner associated with multiple speech signals in a scene of an immersive media application. The processing circuitry determines multiple loudness adjustments to multiple sound signals including multiple speech signals in the scene based on the loudness adjustments to the adjusted speech signal, and generates multiple sound signals in the scene based on the multiple loudness adjustments to the multiple sound signals.

[0007] In some examples, the processing circuitry decodes from the encoded bitstream an index indicating that one of the multiple speech signals is the adjusted speech signal.

[0008] In one example, this information indicates that the loudest speech signal among multiple speech signals is the adjusted speech signal. In another example, this information indicates that the quietest speech signal among multiple speech signals is the adjusted speech signal.

[0009] In some examples, this information indicates that the adjusted speech signal has the average loudness of multiple speech signals.

[0010] In some examples, this information indicates that the adjusted speech signal has the average loudness of the loudest and quietest speech signals among multiple speech signals.

[0011] In some examples, this information indicates that the adjusted speech signal has the median loudness of multiple speech signals.

[0012] In some examples, this information indicates that the adjusted speech signal has the average loudness of a set of speech signals. This set of speech signals has the loudness of multiple quantiles of the speech signals.

[0013] In some examples, the processing circuitry determines the speech signal associated with a location as an adjusted speech signal. This location is the closest to the center of the locations associated with multiple speech signals.

[0014] In some examples, this information indicates that the adjusted speech signal has a weighted average loudness of multiple speech signals. In one example, the processing circuitry determines the weights of the multiple speech signals based on their location. In another example, the processing circuitry determines the weights of the multiple speech signals based on their respective loudness.

[0015] This disclosure also provides a non-transitory computer-readable medium for storing instructions that, when executed by a computer, cause the computer to perform an audio processing method. Attached Figure Description

[0016] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0017] Figure 1 A block diagram of an immersive media system according to an embodiment of the present disclosure is shown.

[0018] Figure 2 A flowchart illustrating an example of an overview process according to an embodiment of this disclosure is shown.

[0019] Figure 3 A flowchart illustrating another example of a process according to an embodiment of this disclosure is shown.

[0020] Figure 4This is a schematic diagram of a computer system according to one embodiment. Detailed Implementation

[0021] Various aspects of this disclosure provide techniques for adjusting audio loudness associated with scenes in immersive media applications. In immersive media applications, such as interactive virtual reality (VR) or augmented reality (AR), different sound levels in the scene can be set using various techniques, such as technical settings, loudness measurements, and manual settings. According to some aspects of this disclosure, when multiple sound signals associated with a scene in an immersive media application include multiple speech signals, the loudness of an adjusted speech signal can be determined based on the multiple speech signals in the scene of the immersive media application. Then, a loudness adjustment of the adjusted speech signal is determined to match the loudness of the adjusted speech signal to a reference signal. Furthermore, the loudness of the multiple sound signals associated with the scene can be adjusted based on the loudness adjustment of the adjusted speech signal. In some examples, information indicating the adjusted speech signal and the loudness adjustment of the adjusted speech signal can be encoded in a bitstream carrying encoded information for generating the multiple sound signals, such as a bitstream carrying immersive media for the immersive media application. Then, in some examples, when a user device with an immersive media player receives a bitstream, the user device can determine an adjusted speech signal for a scene based on information in the bitstream. Furthermore, based on loudness adjustments to the adjusted speech signal, the user device can adjust multiple sound signals associated with the scene.

[0022] Figure 1 A block diagram of an immersive media system (100) according to an embodiment of the present disclosure is shown. The immersive media system (100) can be used in various applications, such as augmented reality (AR) applications, virtual reality applications, video game goggles applications, sports game animation applications, etc.

[0023] The immersive media system (100) includes an immersive media encoding subsystem (101) and an immersive media decoding subsystem (102) that can be connected via a network (not shown). In one example, the immersive media encoding subsystem (101) may include one or more devices with audio encoding / decoding and video encoding / decoding capabilities. In one example, the immersive media encoding subsystem (101) includes a single computing device, such as a desktop computer, laptop computer, server computer, tablet computer, etc. In another example, the immersive media encoding subsystem (101) includes a data center, server cluster, etc. The immersive media encoding subsystem (101) can receive video and audio content and compress the video and audio content into a coded bitstream according to a suitable media encoding / decoding standard. The coded bitstream can be transmitted to the immersive media decoding subsystem (102) via the network.

[0024] The immersive media decoding subsystem (102) includes one or more devices with video and audio encoding / decoding capabilities for immersive media applications. In one example, the immersive media decoding subsystem (102) includes computing devices such as desktop computers, laptop computers, server computers, tablet computers, wearable computing devices, head-mounted display (HMD) devices, etc. The immersive media decoding subsystem (102) can decode the encoded bitstream according to a suitable media encoding / decoding standard. The decoded video and audio content can be used for immersive media playback.

[0025] The immersive media encoding subsystem (101) can be implemented using any suitable technology. Figure 1 In the example, the immersive media encoding subsystem (101) includes processing circuitry (120) and interface circuitry (111) coupled together.

[0026] The processing circuitry (120) may include any suitable processing circuitry, such as one or more central processing units (CPUs), one or more graphics processing units (GPUs), application-specific integrated circuits (ASICs), etc. Figure 1 In one example, the processing circuitry (120) can be configured to include various encoders, such as an audio encoder (130), a video encoder (not shown), etc. In one example, one or more CPUs and / or GPUs can execute software to function as the audio encoder (130). In another example, the audio encoder (130) can be implemented using an application-specific integrated circuit (ASIC).

[0027] In some examples, the audio encoder (130) participates in a listening test setup for determining multiple loudness adjustments of multiple sound signals. Furthermore, the audio encoder (130) can appropriately encode information about the multiple loudness adjustments of the multiple sound signals in an encoded bitstream, such as in metadata. For example, the audio encoder (130) may include a loudness controller (140) that determines the loudness adjustment based on the loudness of the adjusted speech signal. The loudness of the adjusted speech signal is a function of multiple speech signals associated with a scene. A scene may have multiple speech signals among multiple sound signals associated with the scene. Metadata indicating the adjusted speech signal and the loudness adjustment of the adjusted speech signal can then be included in the encoded bitstream.

[0028] The interface circuit (111) can connect the immersive media encoding subsystem (101) to a network. The interface circuit (111) may include a receiving section for receiving signals from the network and a transmitting section for sending signals to the network. For example, the interface circuit (111) can transmit signals carrying encoded bitstreams to other devices, such as the immersive media decoding subsystem (102), via the network.

[0029] The network is appropriately coupled to the immersive media encoding subsystem (101) and the immersive media decoding subsystem (102) via wired and / or wireless connections (e.g., Ethernet, fiber optic, WiFi, cellular, etc.). The network may include network server devices, storage devices, network devices, etc. The components of the network are appropriately coupled together via wired and / or wireless connections.

[0030] The immersive media decoding subsystem (102) is configured to decode the encoded bitstream. In one example, the immersive media decoding subsystem (102) may perform video decoding to reconstruct a sequence of video frames that can be displayed, and perform audio decoding to reconstruct an audio signal for playback.

[0031] The immersive media decoding subsystem (102) can be implemented using any suitable technology. Figure 1 In the example, an immersive media decoding subsystem (102) is shown, but not limited to, a head-mounted display (HMD) with headphones that can be used by a user device. The immersive media decoding subsystem (102) includes, for example... Figure 1 The interface circuit (161) and processing circuit (170) are shown coupled together.

[0032] The interface circuit (161) can connect the immersive media decoding subsystem (102) to a network. The interface circuit (161) may include a receiving section for receiving signals from the network and a transmitting section for sending signals to the network. For example, the interface circuit (161) can receive signals carrying data from the network, such as signals carrying coded bit streams.

[0033] The processing circuit (170) may include suitable processing circuitry, such as a CPU, GPU, application-specific integrated circuit, etc. The processing circuit (170) may be configured to include various decoders, such as an audio decoder (180), a video decoder (not shown), etc.

[0034] In some examples, the audio decoder (180) can decode scene-related audio content, as well as metadata indicating the adjusted speech signal and its loudness adjustment. Furthermore, the audio decoder (180) includes a loudness controller (190) that can adjust the sound levels of multiple sound signals associated with the scene based on the adjusted speech signal and its loudness adjustment.

[0035] According to some aspects of this disclosure, the immersive media system (100) can be implemented according to immersive media standards, such as the Moving Image Experts Group Immersive (MPEG-I) standard suite, which includes "immersive audio," "immersive video," and "system support." Immersive media standards can support VR or AR presentations in which users can navigate and interact with the environment using six degrees of freedom (6DoF) (including spatial navigation (x, y, z) and user head orientation (yaw, pitch, roll)).

[0036] Immersive media systems (100) can give users the feeling of actually existing in a virtual world. In some examples, the audio of the scene is perceived as if it were in the real world, with the sound originating from relevant visual objects. For example, sound is perceived at the correct location and distance within the scene. The user's physical movement in the real world is perceived as matching movement within the virtual scene. Furthermore, the user can interact with the scene and emit sounds that are perceived as realistic and match the user experience in the real world.

[0037] Typically, content providers and / or technology providers may use hearing test setups to determine the sound level of a sound signal to achieve an immersive user experience. In some relevant examples, the sound level (also known as loudness) of a sound signal in a scene is adjusted based on speech signals in the scene. In some examples, multiple speech signals exist within multiple sound signals in a scene. Some aspects of this disclosure provide techniques for loudness adjustment based on an adjusted speech signal when multiple sound signals associated with a scene include multiple speech signals. The loudness of the adjusted speech signal is determined based on multiple speech signals.

[0038] According to one aspect of this disclosure, a loudness adjustment process can be performed by a content creator or technology provider to determine the loudness adjustment of a scene relative to a reference signal (also known as an anchor signal). In one example, the reference signal is a specific speech signal, such as male English speech on track 50 of a Sound Quality Assessment Material (SQAM) disc in a WAV file. In some examples, the loudness adjustment process is performed against a Pulse-Code Modulation (PCM) audio signal used in an Encoder Input Format (EIF). In some examples, a binaural rendering tool, such as a General Binaural Renderer (GBR) with a Dirac head Related Transfer Function (HRTF), can be used during the loudness adjustment process. The binaural rendering tool can simulate the audio environment of the scene and generate the audio signal in the WAV file based on the audio content of the scene.

[0039] In some examples, for instance, one or two measurement points in the scene may be determined by the content creator or technology provider. These measurement points may represent locations on the scene task path with a “normal” loudness for that scene.

[0040] In some examples, binaural rendering tools can be used to define the spatial relationship between the sound source location and the measurement point, and output the scene output signal (e.g., sound signal) at the measurement point based on the audio content at the sound source location.

[0041] In some examples, the scene output signal (e.g., an audio signal) is a WAV file, which can be compared with a reference signal to determine the necessary adjustments to the sound level.

[0042] In one example, the audio content of the scene includes speech. In a binaural rendering tool, the source location and measurement location of the speech content can be defined as being approximately a distance apart, such as a predefined distance (e.g., 1.5 meters), or a distance specific to the scene. Other suitable configurations for the scene can be set in the binaural rendering tool, which simulates the audio environment of the scene and generates a scene output signal at the measurement location based on the speech content at the source, such as a speech signal in a WAV file. The speech signal can then be compared with a reference signal to determine the loudness adjustment of the speech signal, which can be used to match the loudness of the speech signal to the reference signal. In one example, loudness can be measured as a function of the average signal strength over a time range. After determining the loudness adjustment of the speech signal, the sound levels of other sound signals in the scene can be adjusted based on this loudness adjustment.

[0043] According to some aspects of this disclosure, there may be two or more speech signals in a scene, and an adjusted speech signal can be determined based on these two or more speech signals. Then, exemplarily, a loudness adjustment of the adjusted speech signal is determined so that the loudness of the adjusted speech signal matches that of a reference signal. Other sound signals in the scene (e.g., speech signals, non-speech signals, etc.) can then be adjusted in an appropriate manner based on the loudness adjustment of the adjusted speech signal.

[0044] Furthermore, in some examples, content creators or technology providers can identify the loudest points on the scene's task path. In one example, they check whether the sound loudness at the loudest point is not clipped (e.g., below clipping limits). Additionally, in some examples, they can identify and check whether some very soft points or areas in the scene are excessively quiet.

[0045] It is worth noting that the adjusted speech signal can be determined using various techniques based on multiple speech signals in a scene, and the loudness of the adjusted speech signal can also be determined using various techniques. Assuming there are M speech signals in the scene (M is an integer greater than 1), the loudness of the speech signals can be represented by S1, S2, S3, ..., S... M express.

[0046] In some embodiments, the adjusted speech signal may be one of multiple speech signals presented in the scene. In one example, the content creator or technology provider may determine the selection of one of the multiple speech signals. The selection of one of the multiple speech signals may be indicated in the encoded bitstream or as part of the metadata associated with the audio content.

[0047] Specifically, in one example, the measurement location and sound source location of the selected speech signal can be defined in the binaural rendering tool. Other suitable configurations for the scene can be set in the binaural rendering tool, which simulates the audio environment of the scene and generates a scene output signal in a WAV file based on the audio content of the selected speech signal. In this example, the scene output signal is the adjusted speech signal. The adjusted speech signal can be compared with a reference signal to determine the loudness adjustment of the adjusted speech signal. The loudness adjustment of the adjusted speech signal can be used to match the loudness of the adjusted speech signal with the reference signal. For example, when i is the index of the selected speech signal, S i This is the loudness of the adjusted speech signal. Then, S... i The loudness of the adjusted speech signal in the scene is determined by comparing it with the loudness of a reference signal to match the loudness of the reference signal.

[0048] In some embodiments, the adjusted speech signal may be the loudest speech signal present in the scene.

[0049] Specifically, in one example, to generate each speech signal in a scene, the measurement location and sound source location of the speech signal can be defined in the binaural rendering tool. Other suitable configurations for the scene can be set in the binaural rendering tool. The binaural rendering tool can simulate the audio environment of the scene and generate the scene output signal in a WAV file, i.e., the speech signal that can be perceived at the measurement location. Then, the loudest speech signal among multiple speech signals can be selected as the adjusted speech signal. The adjusted speech signal can be compared with a reference signal to determine the loudness adjustment of the loudest speech signal. Loudness adjustment can be used to match the loudness of the adjusted speech signal to the reference signal. For example, S max S1, S2, S3, ..., S M The maximum loudness in S. max The loudness is compared with that of a reference signal to determine the loudness adjustment of the loudest speech signal in the scene.

[0050] In some embodiments, the adjusted speech signal corresponds to the quietest speech signal presented in the scene.

[0051] Specifically, in one example, to generate each speech signal in a scene, the measurement location and sound source location of the speech signal can be defined in the binaural rendering tool. Other suitable configurations for the scene can be set in the binaural rendering tool. The binaural rendering tool can simulate the audio environment of the scene and generate the scene output signal in a WAV file, i.e., the speech signal perceived at the measurement location. Then, the quietest speech signal among multiple speech signals is determined as the adjusted speech signal. The adjusted speech signal can be compared with a reference signal to determine the loudness adjustment of the adjusted speech signal. Loudness adjustment can be used to match the loudness of the adjusted speech signal to the reference signal. For example, S min S1, S2, S3, ..., S M The minimum loudness in S. min The loudness of the speech signal is compared with that of a reference signal to determine the loudness adjustment of the quietest speech signal in the scene.

[0052] In some embodiments, the adjusted speech signal may be the average of all speech signals presented in the scene.

[0053] Specifically, in one example, to generate each speech signal in a scene, the measurement location and sound source location of the speech signal can be defined in the binaural rendering tool. Other suitable configurations for the scene can be set in the binaural rendering tool. The binaural rendering tool can simulate the audio environment of the scene and generate the scene output signal in a WAV file, i.e., the speech signal perceived at the measurement location. Then, the average loudness of multiple speech signals can be determined as the loudness of the adjusted speech signal, which can be considered as a virtual signal. The average loudness can be compared with the loudness of a reference signal to determine the loudness adjustment. Loudness adjustment can be used to match the loudness of the adjusted speech signal to the reference signal. For example, S 平均 S1, S2, S3, ..., S M The average loudness can be calculated using formula (1).

[0054] S 平均 = (S1+S2+S3+…+S M Formula (1) / M

[0055] S 平均 The loudness of the adjusted speech signal is determined by comparing it with the loudness of a reference signal.

[0056] In some embodiments, the adjusted speech signal may be the average of the loudest and quietest speech signals that occur in the scene.

[0057] Specifically, in one example, to generate each speech signal in a scene, the measurement location and sound source location of the speech signal can be defined in the binaural rendering tool. Other suitable configurations for the scene can be set in the binaural rendering tool. The binaural rendering tool can simulate the audio environment of the scene and generate the scene output signal in a WAV file, i.e., the speech signal perceived at the measurement location. Then, the loudest and quietest speech signals among multiple speech signals can be determined. The loudness of the adjusted speech signal is calculated as the average loudness of the loudest and quietest speech signals. The loudness of the adjusted speech signal is compared with the loudness of a reference signal to determine the loudness adjustment of the adjusted speech signal. Loudness adjustment can be used to match the loudness of the adjusted speech signal to the reference signal. For example, S max S1, S2, S3, ..., S M The maximum loudness in, S min S1, S2, S3, ..., S M The minimum loudness in, S a The average loudness, representing the maximum and minimum loudness, can be calculated using formula (2).

[0058] S a =(S max +S min ) / 2 formula (2)

[0059] S a The loudness of the adjusted speech signal is determined by comparing it with the loudness of a reference signal.

[0060] In some embodiments, the adjusted speech signal may be the median of all speech signals presented in the scene.

[0061] Specifically, in one example, to generate each speech signal in a scene, the measurement location and sound source location of the speech signal can be defined in the binaural rendering tool. Other suitable configurations for the scene can be set in the binaural rendering tool. The binaural rendering tool can simulate the audio environment of the scene and generate the scene output signal in a WAV file, i.e., the speech signal perceived at the measurement location. The median loudness among multiple speech signals can then be determined as the loudness of the adjusted speech signal. The loudness of the adjusted speech signal can be compared with a reference signal to determine the loudness adjustment of the adjusted speech signal. Loudness adjustment can be used to match the loudness of the adjusted speech signal to the reference signal. For example, S 中值 S1, S2, S3, ..., S M The median loudness can be expressed by formula (3).

[0062] S 中值 = median {S1,S2,S3,…,S}M} Formula (3)

[0063] S 中值 The loudness of the adjusted speech signal is determined by comparing it with the loudness of a reference signal.

[0064] In some embodiments, the adjusted speech signal corresponds to the average of the quantiles of all speech signals presented in the scene, such as the 25th to 75th percentiles.

[0065] Specifically, in one example, to generate each speech signal in a scene, the measurement location and sound source location of the speech signal can be defined in the binaural rendering tool. Other suitable configurations for the scene can be set in the binaural rendering tool. The binaural rendering tool can simulate the audio environment of the scene and generate the scene output signal in a WAV file, i.e., the speech signal perceived at the measurement location. The speech signals can then be sorted according to loudness to determine a set of speech signals in the quantiles of the speech signals. The loudness of the adjusted speech signal can then be calculated as the average loudness of that set of speech signals. The loudness of the adjusted speech signal can be compared with a reference signal to determine the loudness adjustment of the adjusted speech signal. Loudness adjustment can be used to match the loudness of the adjusted speech signal to the reference signal. For example, S qa-b S1, S2, S3, ..., S M The average loudness of a subset (quantiles from a% to b%) can be expressed by formula (4).

[0066] S qa-b = Mean (quantiles) a%,b% {S1,S2,S3,…,S M}) Formula (4)

[0067] S qa-b The loudness of the adjusted speech signal is determined by comparing it with the loudness of a reference signal.

[0068] In another example, S q25-75 S1, S2, S3, ..., S M The average loudness of a subset (from 25% to 75% quantiles) can be expressed by formula (5).

[0069] S q25-75 = Mean (quantiles) 25%,75% {S1,S2,S3,…,S M}) Formula (5)

[0070] S q25-75 The loudness of the adjusted speech signal is determined by comparing it with the loudness of a reference signal.

[0071] In some embodiments, the adjusted speech signal may be the speech signal that is closest to the cluster center of all speech signals presented in the scene.

[0072] Specifically, in one example, the source location of the speech signal closest to the cluster center of all speech signals can be determined based on the source locations of multiple speech signals; this speech signal is referred to as the center speech signal. In a binaural rendering tool, the measurement location and source location of the center speech signal can be defined. Other suitable scene configurations can be set in the binaural rendering tool. The binaural rendering tool can simulate the audio environment of the scene and generate the scene output signal in a WAV file, i.e., the center speech signal perceived at the measurement location. Then, in this example, the center speech signal is the adjusted speech signal. The loudness of the adjusted speech signal can be compared with a reference signal to determine the loudness adjustment of the adjusted speech signal. The loudness adjustment of the center speech signal can be used to match the loudness of the adjusted speech signal to the reference signal. For example, S 中心 S1, S2, S3, ..., S M One of them, whose corresponding speech signal is the central speech signal, can be represented by formula (6).

[0073] S 中心 = Cluster_centers {S1,S2,S3,…,S} M} Formula (6)

[0074] In some embodiments, the adjusted speech signal may be a weighted average of all speech signals presented in the scene, wherein the weights may be based on distance or on loudness.

[0075] Specifically, in one example, to generate each speech signal in a scene, the measurement location and sound source location of the speech signal can be defined in the binaural rendering tool. Other suitable configurations for the scene can be set in the binaural rendering tool. The binaural rendering tool can simulate the audio environment of the scene and generate the scene output signal in a WAV file, i.e., the speech signal perceived at the measurement location. Then, a weighted average loudness of multiple speech signals can be calculated and used as the loudness of the adjusted speech signal. The adjusted speech signal can be considered as a virtual signal. The weighted average loudness can be compared with the loudness of a reference signal to determine the loudness adjustment. For example, S 加权 Indicates the weighted average loudness; w1, w2, w3, ..., w M S1, S2, S3, ..., S M The weights, S 加权 It can be calculated according to formula (7).

[0076] S 加权=S1×w1+S2×w2+S3×w3+…+S M ×w M Formula (7)

[0077] In one example, the weights w1, w2, w3, ..., w M The sum of S equals 1. 加权 The loudness of the adjusted speech signal is determined by comparing it with the loudness of a reference signal. In some examples, weights w1, w2, w3, ..., w... M The weights are determined based on the distances from each sound source location to the measurement location. In some examples, the weights w1, w2, w3, ..., w... M Based on loudness S1, S2, S3, ..., S M To determine.

[0078] Figure 2 A flowchart of an overview process (200) according to an embodiment of this disclosure is shown. The process (200) can be used for audio encoding / decoding, such as for an immersive media encoding subsystem (101), and is executed by processing circuitry (120), etc. In some embodiments, the process (200) is implemented as software instructions, so that when the processing circuitry executes these software instructions, the processing circuitry performs the process (200). The process begins at (S201) and proceeds to (S210).

[0079] Step (S210) determines the loudness of the adjusted speech signal based on multiple speech signals related to the scene in the immersive media application.

[0080] Step (S220) determines the loudness adjustment that matches the loudness of the adjusted speech signal to the reference signal.

[0081] In step (S230), loudness adjustment is encoded in a bitstream carrying audio content associated with the scene.

[0082] In some examples, the adjusted speech signal is one of multiple speech signals, and the index used to indicate the selection of the adjusted speech signal from the multiple speech signals can be encoded in the bitstream.

[0083] In some examples, the loudest or quietest of multiple speech signals can be selected as the adjusted speech signal.

[0084] In some examples, the average loudness of multiple speech signals is determined as the loudness of the adjusted speech signal.

[0085] In some examples, the average loudness of the loudest and quietest speech signals among multiple speech signals is determined as the loudness of the adjusted speech signal.

[0086] In some examples, the median loudness of multiple speech signals is determined as the loudness of the adjusted speech signal.

[0087] In some examples, the loudness of a set of speech signals is determined as the loudness of the adjusted speech signal. This set of speech signals consists of quantiles of multiple speech signals, such as the 20th to 75th percentiles.

[0088] In some examples, the speech signal associated with a location in the scene is determined to be the adjusted speech signal. This location is the closest point in the scene to the center of the location associated with multiple speech signals.

[0089] In some examples, the weighted average loudness of multiple speech signals is determined as the loudness of the adjusted speech signal. In one example, the weights of the multiple speech signals are determined based on their location. In another example, the weights of the multiple speech signals are determined based on their respective loudness.

[0090] Then, the process proceeds to (S299) and ends.

[0091] Figure 3 A flowchart of an overview process (300) according to an embodiment of the present disclosure is shown. The process (300) can be used for audio encoding / decoding, such as for an immersive media decoding subsystem (102), and is executed by processing circuitry (170), etc. In some embodiments, the process (300) is implemented as software instructions, so that the processing circuitry executes the process (300) when the software instructions are executed. The process begins at (S301) and proceeds to (S310).

[0092] Step (S310) decodes information indicating the adjusted speech signal and the loudness adjustment of the adjusted speech signal from the encoded bitstream. The adjusted speech signal is indicated in a manner that associates it with multiple speech signals in the scenario of the immersive media application.

[0093] Step (S320) determines multiple loudness adjustments for multiple sound signals, including multiple speech signals in the scene, based on the loudness adjustment of the adjusted speech signal.

[0094] Step (S330) generates multiple sound signals in the scene based on multiple loudness adjustments of multiple sound signals.

[0095] In some examples, an index is decoded from the encoded bitstream to indicate that one of the multiple speech signals is the adjusted speech signal.

[0096] In one example, this information indicates that the loudest speech signal among multiple speech signals is the adjusted speech signal. In another example, this information indicates that the quietest speech signal among multiple speech signals is the adjusted speech signal.

[0097] In some examples, this information indicates that the adjusted speech signal has the average loudness of multiple speech signals.

[0098] In some examples, this information indicates that the adjusted speech signal has the average loudness of the loudest and quietest speech signals among multiple speech signals.

[0099] In some examples, this information indicates that the adjusted speech signal has the median loudness of multiple speech signals.

[0100] In some examples, this information indicates that the adjusted speech signal has the average loudness of a set of speech signals. This set of speech signals has the loudness of multiple speech signal quantiles (e.g., 25% to 75% quantiles).

[0101] In some examples, the speech signal associated with a location is determined to be an adjusted speech signal. For example, the location is the sound source location of the speech signal. This location is the closest to the center of the locations associated with multiple speech signals.

[0102] In some examples, this information indicates that the adjusted speech signal has a weighted average loudness of multiple speech signals. In one example, the weights of the multiple speech signals are determined based on their respective positions. In another example, the weights of the multiple speech signals are determined based on their respective loudnesses.

[0103] Then, the process proceeds to (S399) and ends.

[0104] The above-described techniques can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 4 A computer system (400) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0105] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode, etc.

[0106] The instructions can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things devices.

[0107] Figure 4 The components of the computer system (400) shown are exemplary in nature and are not intended to impose any limitation on the scope or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any one or a combination of components shown in the exemplary embodiments of the computer system (400).

[0108] The computer system (400) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users, for example, through input such as: tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images acquired from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0109] The input human-machine interface device may include one or more of the following (only one of each is shown): keyboard (401), mouse (402), touchpad (403), touch screen (410), data glove (not shown), joystick (405), microphone (406), scanner (407), camera (408).

[0110] The computer system (400) may also include certain human-machine interface output devices. Such human-machine interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touchscreen (410), a data glove (not shown), or a joystick (405), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (409), headphones (not shown)), visual output devices (e.g., screens (410) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input functionality, each with or without tactile feedback functionality - some of these screens are capable of outputting two-dimensional or more three-dimensional visual outputs through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).

[0111] The computer system (400) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (420) with media such as CD / DVD (421), finger drives (422), removable hard disk drives or solid-state drives (423), conventional magnetic media such as magnetic tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown), etc.

[0112] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0113] The computer system (400) may also include an interface (454) to one or more communication networks (455). The network may be, for example, a wireless network, a wired network, or an optical network. The network may further be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CANBus, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (400)) to connect to some general-purpose data port or peripheral bus (449); other network interfaces are typically integrated into the core of the computer system (400) by connecting to a system bus (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system (400) can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANbus connected to certain CANbus devices), or bidirectional, such as connecting to other computer systems using a local area network (LAN) or wide area network (WAN) digital network. As mentioned above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0114] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel (440) of the computer system (400).

[0115] The core (440) may include one or more central processing units (CPU) (441), graphics processing units (GPUs) (442), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (443), hardware accelerators (444) for certain tasks, graphics adapters (450), etc. These devices, as well as read-only memory (ROM) (445), random access memory (446), and internal mass storage (447) such as internal non-user-accessible hard disk drives, SSDs, etc., may be connected via a system bus (448). In some computer systems, the system bus (448) may be accessed in the form of one or more physical plugs to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus (448) or connected to the core's system bus via a peripheral bus (449). In one example, a screen (410) may be connected to a graphics adapter (450). Peripheral bus architectures include PCI, USB, etc.

[0116] The CPU (441), GPU (442), FPGA (443), and accelerator (444) can execute certain instructions, which can be combined to form the aforementioned computer code. This computer code can be stored in ROM (445) or RAM (446). Transient data can also be stored in RAM (446), while permanent data can be stored, for example, in internal mass storage (447). Fast storage and retrieval to any storage device can be achieved using a cache, which can be closely associated with one or more CPUs (441), GPUs (442), mass storage (447), ROM (445), RAM (446), etc.

[0117] Computer-readable media may have computer code thereon for performing various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0118] As a non-limiting example, a computer system having an architecture (400), particularly a kernel (440), can provide functionality by having one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, and some non-transitory memory of the kernel (440), such as internal kernel mass storage (447) or ROM (445). Software implementing various embodiments of this disclosure can be stored in such means and executed by the kernel (440). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the kernel (440), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures (446) stored in RAM and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system may provide functionality through hard-wired or otherwise embodied logic in circuitry (e.g., accelerator (444)) that may replace or operate with the software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0119] Although several exemplary embodiments have been described in this disclosure, modifications, substitutions, and various equivalent alternatives that fall within the scope of this disclosure exist. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.

Claims

1. An audio processing method, comprising: The processor decodes from the encoded bitstream information indicating an adjusted speech signal and information on loudness adjustment of the adjusted speech signal, wherein the adjusted speech signal is indicated in a manner associated with multiple speech signals in a scenario of an immersive media application, and the information indicates that the adjusted speech signal has a weighted average loudness of multiple speech signals, the weights of which are determined based on the distances from each sound source location to the measurement location. The processor determines the sound source location of the central speech signal that is closest to the cluster center of all speech signals based on the sound source locations of the plurality of speech signals, and the central speech signal is the adjusted speech signal; The processor determines multiple loudness adjustments for multiple sound signals, including the plurality of speech signals in the scene, based on the loudness adjustment of the adjusted speech signal; and The processor generates the plurality of sound signals in the scene based on multiple loudness adjustments of the plurality of sound signals.

2. The method according to claim 1, further comprising: The index indicating that one of the plurality of speech signals is the adjusted speech signal is decoded from the encoded bitstream.

3. The method according to claim 1, wherein, The information indicates at least one of the following: The loudest of the multiple speech signals is the adjusted speech signal; or The quietest of the multiple speech signals is the adjusted speech signal.

4. The method according to claim 1, wherein, The information indicates that the adjusted speech signal has the average loudness of the plurality of speech signals.

5. The method according to claim 1, wherein, The information indicates that the adjusted speech signal has the average loudness of the loudest and quietest speech signals among the plurality of speech signals.

6. The method according to claim 1, wherein, The information indicates that the adjusted speech signal has the median loudness of the plurality of speech signals.

7. The method according to claim 1, wherein, The information indicates that the adjusted speech signal has an average loudness of a set of speech signals, wherein the set of speech signals has the loudness of the quantiles of the plurality of speech signals.

8. The method according to claim 1, further comprising: The speech signal associated with a location is determined to be the adjusted speech signal, wherein the location is the location closest to the center of the locations associated with the plurality of speech signals.

9. The method of claim 1, further comprising at least one of the following: The weights of the multiple speech signals are determined based on their positions; or The weights of the multiple speech signals are determined based on their respective loudness.

10. An audio processing apparatus, comprising processing circuitry, the processing circuitry being configured to: Decode from the encoded bitstream information indicating the adjusted speech signal and information on the loudness adjustment of the adjusted speech signal, wherein... The adjusted speech signal is indicated in a manner that it is associated with multiple speech signals in the scenario of an immersive media application. The information indicates that the adjusted speech signal has a weighted average loudness of multiple speech signals, the weights of which are determined based on the distance from each sound source location to the measurement location. The processor determines the sound source location of the center speech signal that is closest to the cluster center of all speech signals based on the sound source locations of the plurality of speech signals, and the center speech signal is the adjusted speech signal; Based on the loudness adjustment of the adjusted speech signal, multiple loudness adjustments are determined for multiple sound signals including the multiple speech signals in the scene; as well as The plurality of sound signals are generated in the scene based on multiple loudness adjustments to the plurality of sound signals.

11. The apparatus according to claim 10, wherein, The processing circuit is further configured to: The index indicating that one of the plurality of speech signals is the adjusted speech signal is decoded from the encoded bitstream.

12. The apparatus according to claim 10, wherein, The information indicates at least one of the following: The loudest of the multiple speech signals is the adjusted speech signal; or The quietest of the multiple speech signals is the adjusted speech signal.

13. The apparatus according to claim 10, wherein, The information indicates that the adjusted speech signal has the average loudness of the plurality of speech signals.

14. The apparatus according to claim 10, wherein, The information indicates that the adjusted speech signal has the average loudness of the loudest and quietest speech signals among the plurality of speech signals.

15. The apparatus according to claim 10, wherein, The information indicates that the adjusted speech signal has the median loudness of the plurality of speech signals.

16. The apparatus according to claim 10, wherein, The information indicates that the adjusted speech signal has an average loudness of a set of speech signals, wherein the set of speech signals has the loudness of the quantiles of the plurality of speech signals.

17. The apparatus according to claim 10, wherein, The processing circuit is further configured to: The speech signal associated with a location is determined to be the adjusted speech signal, wherein the location is the location closest to the center of the locations associated with the plurality of speech signals.

18. The apparatus according to claim 10, wherein, The processing circuit is also configured to determine the weights of the plurality of speech signals based on at least one of the following: The location of the multiple voice signals; or The loudness of each of the multiple speech signals.

19. A method for storing an encoded bit stream, characterized in that, The encoded bitstream is processed using the audio processing method according to any one of claims 1 to 9.

20. A method for transmitting coded bit streams, characterized in that, The encoded bitstream is processed using the audio processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Decoder, encoder and method for informed loudness estimation employing by-pass audio object signals in object-based audio coding systems

    US20150348564A1