Method and apparatus for loudness adjustment of audio scenes associated with MPEG-I immersive audio streams

By receiving and determining the voice signals in the audio scene and adjusting the loudness of the audio scene using the anchor voice signal, the lack of audio signal intensity adjustment in the MPEG-I immersive audio stream is solved, and the precise matching of the loudness of the audio scene is achieved and the user experience is improved.

CN115486096BActive Publication Date: 2025-08-22TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180032268.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-10-14
Filing Date
2021-10-15
Publication Date
2025-08-22
Estimated Expiration
2041-10-15

AI Technical Summary

Technical Problem

There are shortcomings in the prior art in how to adjust the audio signal strength in the virtual world, especially in MPEG-I immersive audio streams, which are difficult to effectively adjust the loudness of the audio scene to match the user's real-world experience.

Method used

The receiving module receives the sound signal information in the audio scene, uses the first determination module to determine whether there is a voice signal, and adjusts the loudness level of the audio scene based on the anchor voice signal, uses the second adjustment module to adjust the loudness level of other voice signals according to the reference voice signal, and transmits the adjustment information in combination with the signaling method to achieve precise control of loudness.

Benefits of technology

It realizes accurate adjustment of the loudness of the audio scene in the MPEG-I immersive audio stream, ensuring that the audio signal matches the user experience, avoiding the problem of excessive or low loudness, and improving the quality of the immersive audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115486096B_ABST
    Figure CN115486096B_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure include a method, apparatus, and non-transitory computer-readable storage medium for loudness adjustment in an audio scene. The method includes receiving a first syntax element indicating the number of sound signals included in the audio scene; determining whether the sound signals indicated by the first syntax element include one or more speech signals; determining a reference speech signal from the one or more speech signals based on the one or more speech signals included in the sound signals; adjusting the loudness level of the reference speech signal of the audio scene based on an anchor speech signal; and adjusting the loudness level of the sound signal based on the adjusted loudness level of the reference speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by reference

[0002] This application claims the benefit of priority to U.S. Patent Application No. 17 / 501,749, filed on October 14, 2021, entitled "SIGNALINGLOUDNESS ADJUSTMENT FOR AN AUDIO SCENE," which claims the benefit of priority to U.S. Provisional Application No. 63 / 158,261, filed on March 8, 2021, entitled "SIGNALING LOUDNESS ADJUSTMENT FOR AUDIO SCENE." The disclosures of both of the foregoing applications are incorporated herein by reference in their entireties. Technical Field

[0003] The present application relates generally to virtual reality technology, and more particularly to a method and apparatus for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work described in this background section, the work of the presently named inventors and aspects of the description that may not have otherwise been qualified as prior art at the time of filing are not admitted, either explicitly or implicitly, to be prior art with respect to the present disclosure.

[0005] The Moving Picture Experts Group (MPEG) has proposed a set of standards for immersive audio, immersive video, and system support that can support virtual reality (VR) or augmented reality (AR) presentations, in which users can navigate and interact with the environment using six degrees of freedom (6DoF).

[0006] The goal of MPEG-I presentation is to give the user the feeling of actually being in a virtual world. Audio signals in the virtual world (or virtual scene) are perceived as real-world audio signals, with the sounds coming from the associated visual image. That is, the sounds are perceived at the correct position and distance. The user's physical movements in the real world are perceived as matching movements in the virtual world. Furthermore, and importantly, the user can interact with the virtual scene, so the sounds should be perceived as real and match the user's experience in the real world.

[0007] In interactive VR / AR testing, different sound levels are involved in the listening test setup. The relationship between these sound levels can be given by technical settings, normalized by loudness measurements, or set manually. The process of scene loudness adjustment is described as part of the MPEG-I Immersive Audio call of proposals (CfP).

[0008] However, how to adjust the intensity of the audio signal in the virtual world is a problem that needs to be solved urgently in the existing technology. Summary of the Invention

[0009] Aspects of the present disclosure provide an apparatus for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream, comprising a receiving module, a first determining module, a second determining module, a first adjusting module, and a second adjusting module. The receiving module is configured to receive a first syntax element indicating the number of sound signals included in the audio scene. The first determining module is configured to determine whether the sound signal indicated by the first syntax element includes one or more speech signals. The second determining module is configured to determine a reference speech signal from the one or more speech signals based on the fact that the sound signal includes one or more speech signals. The first adjusting module is configured to adjust the loudness level of the reference speech signal of the audio scene based on an anchor speech signal. The second adjusting module is configured to adjust the loudness level of the sound signal based on the adjusted loudness level of the reference speech signal.

[0010] In an embodiment, the receiving module is further configured to receive a second grammar element indicating whether the sound signal includes one or more speech signals. The first determining module is further configured to determine whether the sound signal includes one or more speech signals based on the second grammar element indicating that the sound signal includes one or more speech signals.

[0011] In an embodiment, the receiving module is further configured to receive a plurality of third syntax elements, each of the third syntax elements indicating whether a corresponding one of the sound signals is a speech signal. The first determining module is further configured to determine that the sound signal includes one or more speech signals based on at least one of the third syntax elements indicating that the corresponding one of the sound signals is a speech signal.

[0012] In an embodiment, the receiving module is further configured to receive a fourth grammar element indicating the number of the one or more speech signals included in the sound signal. The first determining module is further configured to determine that the sound signal includes one or more speech signals based on the number of the one or more speech signals indicated by the fourth grammar element being greater than zero.

[0013] In an embodiment, the receiving module is further configured to receive a fifth syntax element indicating a reference speech signal based on the number of the one or more speech signals being greater than one.

[0014] In an embodiment, the receiving module is further configured to receive a plurality of sixth syntax elements, each of the sixth syntax elements indicating an identification index of a corresponding one of the sound signals.

[0015] In an embodiment, the first determining module is further configured to determine that the sound signal does not include a speech signal. The apparatus further comprises a third adjusting module configured to adjust the loudness level of the sound signal based on a default reference signal.

[0016] Aspects of the present disclosure provide a method for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream. In one method, a first syntax element indicating the number of sound signals included in the audio scene is received. A determination is made as to whether the sound signal indicated by the first syntax element includes one or more speech signals. A reference speech signal is determined from the one or more speech signals based on the inclusion of the one or more speech signals in the sound signal. The loudness level of the reference speech signal of the audio scene is adjusted based on an anchor speech signal. The loudness level of the sound signal is adjusted based on the adjusted loudness level of the reference speech signal.

[0017] Aspects of the present disclosure provide methods for loudness adjustment signaling of an audio scene associated with an MPEG-I immersive audio stream. In one method, a first syntax element indicating the number of sound signals included in the audio scene is included in loudness adjustment information, wherein, in response to determining that the sound signals indicated by the first syntax element include one or more speech signals, a reference speech signal is determined from the one or more speech signals, the loudness level of the reference speech signal of the audio scene is adjusted based on the anchor speech signal, and the loudness level of the sound signal is adjusted based on the adjusted loudness level of the reference speech signal.

[0018] Aspects of the present disclosure further provide a computer device comprising a processor and a memory. The memory is configured to store program code and transmit the program code to the processor. The processor is configured to execute any one or a combination of methods for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream according to instructions in the program code.

[0019] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform any one or a combination of methods for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream.

[0020] Embodiments of the present disclosure provide a method and apparatus for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream. The method for loudness adjustment of an audio scene includes receiving a first syntax element indicating the number of sound signals included in an audio scene; determining whether one or more speech signals are included in the sound signal indicated by the first syntax element; determining a reference speech signal from one or more speech signals based on the inclusion of one or more speech signals in the sound signal; adjusting the loudness level of the reference speech signal of the audio scene based on the anchor speech signal; and adjusting the loudness level of the sound signal based on the adjusted loudness level of the reference speech signal. The method for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream of the present disclosure indicates signaling information for adjustment, which can be transmitted between parties, such as a sender and a receiver, as part of a bitstream or part of metadata. After receiving the signaling information, the receiver can use such information to determine whether and how to adjust the signal level of the received sound signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0022] Figure 1 An example of 6 degrees of freedom according to an embodiment of the present disclosure is shown;

[0023] Figure 2 An exemplary flow chart illustrating an embodiment according to the present disclosure; and

[0024] Figure 3 is a schematic diagram of a computer system according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] MPEG has proposed a set of standards including immersive audio, immersive video, and system support that can support VR or AR presentations where users can navigate and interact with the environment using 6 degrees of freedom (6DoF). Figure 1 An example of 6 degrees of freedom according to an embodiment of the present disclosure is shown. Figure 1 In

[15] , 6 degrees of freedom can be represented by spatial navigation (x, y, z) and user head orientation (yaw, pitch, roll).

[0026] I. Loudness Adjustment of Audio Scenes

[0027] The present disclosure includes a signaling method for scene loudness adjustment.

[0028] According to various aspects of the present disclosure, a scene creator may provide an anchor speech signal as a reference signal to adjust scene loudness. For a sound signal in an audio scene, the process of scene loudness adjustment may be described as follows.

[0029] Loudness adjustments between scene sounds and specified anchor signals should be made by the scene creator (or content creator). In an example, the scene sound can be a pulse-code modulation (PCM) audio signal for encoder input format (EIF). Pulse code modulation (PCM) is a method for digitally representing sampled analog signals. EIF describes the structure and representation of scene metadata information read and compressed by the MPEG-I immersive audio encoder. Content creators can use a general binaural renderer (GBR) with a Dirac head-related transfer function (HRTF) for loudness adjustments.

[0030] One or more (eg, one or two) measurement points may be defined in a scene. These measurement points should represent positions on the scene task path that represent the typical loudness of the scene.

[0031] The scene creator can record the scene output signal using GBR with Dirac HRTF at these locations and use the resulting audio file (e.g., wav file) to compare with the reference signal and determine the necessary adjustments to the scene loudness level.

[0032] If there is a speech signal in the scene, in one example, a measurement position may be about 1.5 m away from the speech source, and the loudness level of the speech signal at the measurement position may be adjusted to be the same as that of the anchor speech signal.

[0033] The loudness levels of all other sound signals in the scene can be adjusted based on the loudness level of the speech signal. For example, each of the loudness levels of all other sound signals can be multiplied by a corresponding scaling factor based on the loudness level of the refined speech signal.

[0034] If no speech signal exists in the scene, the loudness level of the sound signal in the scene may be adjusted compared to the anchor speech signal.

[0035] Additionally, the loudest points along the scenario's task path should be identified by the scenario creator. The loudness level at these loudest points should be checked to avoid clipping. For example, when the listener is unusually close to a sound source, edge clipping should be prevented. In embodiments, adjusting the sound level for unusual proximity is the renderer's job.

[0036] Next, you should check whether there are soft spots or areas on the scenario task path that are too quiet. For example, there should not be long periods of silence on the scenario task path.

[0037] In some embodiments, it is important to determine a reference signal based on the sound signal in the audio scene and adjust the reference signal to the same loudness level as the anchor signal. Without determining the reference signal, the scaling factor of the sound signal may not be determined. For example, if there are two sound signals A (loudness of 5) and B (loudness of 20) in the audio scene and the loudness of the anchor voice signal is 10, then without determining the reference signal, it may not be clear whether to amplify sound signal A to 10 or reduce sound signal B to 10. In this case, a possible solution is to adjust both sound signal A and sound signal B to the same loudness level as the anchor voice signal (e.g., 10). This solution may be undesirable in some applications. Therefore, if a reference signal is determined based on the sound signal in the audio scene, the scaling factor of the sound signal can be determined. For example, if sound signal A is selected as the reference signal, sound signal A can be amplified to 10 using a scaling factor of 2, and sound signal B can be amplified to 40 using the same scaling factor of 2. In addition, due to the anchor voice signal, the voice signal in the audio scene can be selected as the reference signal.

[0038] According to aspects of the present disclosure, when there are two or more speech signals in an audio scene, scene loudness adjustment may be performed as follows.

[0039] The loudness adjustment between the scene sound and the designated anchor signal can be performed by the scene creator (or content creator). In an example, the scene sound can be a PCM audio signal used in EIF. The content creator can use GBR with Dirac HRTF for loudness adjustment.

[0040] One or more (eg, one or two) measurement points may be defined in a scene. These measurement points should represent positions on the scene task path that represent the typical loudness of the scene.

[0041] The scene creator can record the scene output signal using GBR and Dirac HRTF at these locations, and use the resulting audio files (eg, wav files) to compare with the reference signal and determine the necessary adjustments to the scene loudness level.

[0042] If there are two or more speech signals in the scene, an adjusted speech signal can be created. The loudness level of the adjusted speech signal can then be further adjusted to the same loudness as the anchor speech signal. Thereafter, the adjusted speech signal can be used as the refined speech signal.

[0043] The loudness levels of all other sound signals in the scene may be adjusted based on the loudness level of the refined speech signal. For example, each of the loudness levels of all other sound signals may be multiplied by a corresponding scaling factor based on the loudness level of the refined speech signal.

[0044] Additionally, the loudest point on the scenario's task path can be identified by the scenario creator. The loudness level at the loudest point should be checked to avoid clipping. For example, when the listener is unusually close to the sound source, edge clipping should be prevented. In embodiments, it is the renderer's job to adjust the sound level for unusual proximity.

[0045] Next, you should check whether there are soft spots or areas on the scenario task path that are too quiet. For example, there should not be long periods of silence on the scenario task path.

[0046] According to aspects of the present disclosure, when two or more speech signals exist, an adjusted speech signal may be generated from the two or more speech signals present in a scene.

[0047] In an embodiment, the adjusted speech signal may be one of the speech signals present in the scene, with the scene creator making a selection. The selection may be indicated to the user. For example, the selection may be indicated in the bitstream or as part of metadata associated with the audio signal.

[0048] The adjusted speech signal can be selected based on various criteria. For example, the adjusted speech signal can be selected based on at least one characteristic of one or more of the speech signals, or at least one mathematical relationship between one or more of the speech signals. For example, the adjusted speech signal can be determined based on sound level or volume. In an embodiment, the adjusted speech signal can be the loudest speech signal present in the scene. In an embodiment, the adjusted speech signal can be the quietest speech signal present in the scene.

[0049] In some embodiments, the adjusted speech signal may be determined based on a subset of the speech signals or an average or median of the speech signals. Furthermore, in some embodiments, the average may be weighted. In an embodiment, the adjusted speech signal may be the average of all speech signals present in the scene. In an embodiment, the adjusted speech signal may be the average of the loudest and quietest speech signals present in the scene. In an embodiment, the adjusted signal may be the median of all speech signals present in the scene. In an embodiment, the adjusted signal may be the average of a quantile of all speech signals present in the scene, such as a quantile of 25% to 75%. In an embodiment, the adjusted signal may be a weighted average of all speech signals present in the scene, where the weights may be distance-based or loudness-based.

[0050] In some embodiments, the adjusted speech signal may be determined based on clustering of speech signals. For example, the adjusted signal may be a speech signal that is located closest to the cluster center of all speech signals present in the scene.

[0051] Note that the methods included in this disclosure can be used individually or in any combination. The methods can be used in part or as a whole.

[0052] The present disclosure includes a signaling method for scene loudness adjustment. In the signaling method, necessary information for adjustment can be indicated. The signaling information can be part of a bitstream or part of metadata. The signaling information can be transmitted between parties, such as a sender and a receiver. After receiving the signaling information, the receiver can use this information to determine whether and how to adjust the signal level of the received sound signal.

[0053] In some embodiments, the signaling information may specify whether a voice signal is present in the scene. For example, when a voice signal is present in the scene, the signaling information specifies the presence of the voice signal. When a voice signal is present in the scene, the signaling information may specify whether two or more voice signals are present in the scene. Furthermore, if necessary, the signaling information may specify the number of the two or more voice signals.

[0054] In some embodiments, the signaling information may specify whether and how to use a speech signal (when present in the scene) as a reference signal for loudness adjustment, or whether and how to use a default signal level as a reference signal level for loudness adjustment.

[0055] In an embodiment, the signaling information may specify whether to use one of the voice signals (when present in the scene) and adjust it to the same loudness as the anchor voice signal for loudness adjustment. If the voice signal is not used, a default signal level (e.g., the loudness level of the anchor voice signal) may be used as a reference level for adjusting other sound signals.

[0056] In an embodiment, when determining to adopt one of the speech signals for loudness adjustment, the signaling information may specify which speech signal present in the scene is adopted and adjusted to the same loudness as the anchor speech signal.

[0057] In an embodiment, the signaling information may specify whether to use one of the speech signals (when present in the scene) for loudness adjustment. If it is determined that one of the speech signals is to be used for loudness adjustment, the speech signal to be used and adjusted to the same loudness as the anchor speech signal may be determined based on characteristics of the speech signal (e.g., sound level or volume). For example, the loudest speech signal in the scene may be used and adjusted to the same loudness as the anchor speech signal. In another example, the quietest speech signal in the scene may be used and adjusted to the same loudness as the anchor speech signal.

[0058] In an embodiment, the signaling information may specify whether to use one of the speech signals (when present in the scene) for loudness adjustment. If it is determined that one of the speech signals is to be used for loudness adjustment, a speech signal to be used and adjusted to the same loudness as the anchor speech signal may be determined based on clustering of the speech signals. For example, the speech signal located closest to the cluster center of all speech signals present in the scene may be used and adjusted to the same loudness as the anchor speech signal. The cluster center may be derived based on the positions of all speech signals.

[0059] In an embodiment, the signaling information may specify whether to use one of the speech signals (when present in the scene) for loudness adjustment. If it is determined that one of the speech signals is to be used for loudness adjustment, a speech signal to be used and adjusted to the same loudness as the anchor speech signal may be determined based on the adjusted speech signal. For example, the adjusted speech signal may be generated based on the available speech signals in the scene and adjusted to the same loudness as the anchor speech signal.

[0060] In some embodiments, the signaling information may specify how to generate the adjusted speech signal based on the available speech signals in the scene. The adjusted speech signal may be determined based on a subset of the speech signals or an average or median of the speech signals. Furthermore, in some embodiments, the average may be weighted.

[0061] In an embodiment, the signaling information may specify whether to use an adjusted speech signal generated from available speech signals (when present in the scene) as a reference signal for loudness adjustment. If it is determined that the generated adjusted speech signal is used as the reference signal for loudness adjustment, the adjusted speech signal may be an average of all speech signals present in the scene.

[0062] In an embodiment, the signaling information may specify whether to use an adjusted speech signal generated from an available speech signal (when present in the scene) as a reference signal for loudness adjustment. If it is determined that the generated adjusted speech signal is used as the reference signal for loudness adjustment, the adjusted speech signal may be an average of the loudest speech signal and the quietest speech signal present in the scene.

[0063] In an embodiment, the signaling information may specify whether to use an adjusted speech signal generated from an available speech signal (when present in the scene) as a reference signal for loudness adjustment. If it is determined that the generated adjusted speech signal is used as a reference signal for loudness adjustment, the adjusted speech signal may be the median of all speech signals present in the scene.

[0064] In an embodiment, the signaling information may specify whether to use an adjusted speech signal generated from an available speech signal (when present in the scene) as a reference signal for loudness adjustment. If it is determined that the generated adjusted speech signal is used as the reference signal for loudness adjustment, the adjusted speech signal may be the average of the quantiles of all speech signals present in the scene.

[0065] In an embodiment, the signaling information may specify whether to use an adjusted speech signal generated from available speech signals (when present in the scene) as a reference signal for loudness adjustment. If it is determined that the generated adjusted speech signal is used as the reference signal for loudness adjustment, the adjusted speech signal may be a weighted average of all speech signals present in the scene.

[0066] In an embodiment, the signaling information may specify that the weight is based on distance. For example, the farther away from the assumed center, the lower the level of weight that may be assigned.

[0067] In an embodiment, the signaling information may specify that the weighting is based on loudness. For example, the quieter the speech signal, the lower the level of weighting that may be assigned.

[0068] An exemplary syntax table of signaling information is shown in Table 1.

[0069] Table 1

[0070] name Bit length describe num_sound 2 or more The number of sound signals in the scene sound_id 2 or more Sound signal identification index is_speech_flag 1 Is the sound signal speech? speech_present_flag 1 Is there a voice signal in the scene? num_speech_signals 2 or more The number of speech signals in the scene adjusted_speech_signal_method 3 or more How to create a conditioned speech signal

[0071] In Table 1, the syntax element num_sound (e.g., 2 or more bits) indicates the number of sound signals in the audio scene. For each sound signal in the audio scene, the signaling information may include a corresponding syntax element sound_id (e.g., 2 or more bits) that specifies the identification index of the corresponding sound signal. For each sound signal in the audio scene, the signaling information may include a corresponding one-bit flag is_speech_flag that specifies whether the corresponding sound signal is a speech signal.

[0072] In an embodiment, the signaling information may include a one-bit flag speech_present_flag, which specifies whether a speech signal exists in the scene.

[0073] In an embodiment, it may be determined whether there is a speech signal in the scene by checking whether there is a sound signal with an associated syntax element is_speech_flag equal to 1.

[0074] In an embodiment, if it is determined that a speech signal exists in the scene, the signaling information may include a syntax element num_speech_signals (eg, 2 bits or more) that specifies the number of speech signals existing in the scene.

[0075] In an embodiment, the number of speech signals present in a scene may be obtained by counting the number of sound signals in which the respective associated syntax element is_speech_flag is equal to 1.

[0076] In an embodiment, multiple loudness adjustment methods may be supported. The multiple loudness adjustment methods may include one or more methods described in this disclosure. In an example, a subset of these methods may be allowed.

[0077] In an embodiment, if the number of speech signals present in a scene is more than one, the signaling information may include a syntax element adjusted_speech_signal_method (eg, 3 bits or more) that specifies how to generate an adjusted speech signal for loudness adjustment.

[0078] Table 2 shows an exemplary signaling method for loudness adjustment.

[0079] Table 2

[0080]

[0081]

[0082]

[0083] This disclosure includes a data structure for loudness adjustment signaling of an audio scene associated with an MPEG-I immersive audio stream. The data structure includes, in loudness adjustment information, a first syntax element indicating the number of sound signals included in the audio scene. In response to determining that one or more speech signals are included in the sound signal based on the first syntax element, a reference speech signal is determined from the one or more speech signals. The loudness level of the reference speech signal of the audio scene is adjusted based on the anchor speech signal. The loudness level of the sound signal is adjusted based on the adjusted loudness level of the reference speech signal.

[0084] In an embodiment, the data structure includes a second syntax element indicating whether the sound signal includes one or more speech signals in the loudness adjustment information. Based on the second syntax element indicating that the sound signal includes one or more speech signals, it is determined that the sound signal includes one or more speech signals.

[0085] In an embodiment, the data structure includes a plurality of third syntax elements in the loudness adjustment information. Each of the third syntax elements indicates whether a corresponding one of the sound signals is a speech signal. Based on at least one of the third syntax elements indicating that the corresponding one of the sound signals is a speech signal, it is determined that the sound signals include one or more speech signals.

[0086] In an embodiment, the data structure includes a fourth syntax element indicating the number of the one or more speech signals included in the sound signal in the loudness adjustment information. Based on the number of the one or more speech signals indicated by the fourth syntax element being greater than zero, it is determined that the sound signal includes the one or more speech signals.

[0087] In an embodiment, based on the number of the one or more speech signals being greater than one, the data structure includes a fifth syntax element indicating the reference speech signal in the loudness adjustment information.

[0088] In an embodiment, the data structure includes a plurality of sixth syntax elements in the loudness adjustment information. Each of the sixth syntax elements indicates an identification index of a corresponding one of the sound signals.

[0089] II. Flowchart

[0090] Figure 2 A flow chart outlining an exemplary process (200) according to an embodiment of the present disclosure is shown. In various embodiments, the process (200) is performed by processing circuitry (e.g., Figure 3 In some embodiments, the process (200) is implemented as software instructions, so when the processing circuitry executes the software instructions, the processing circuitry performs the process (200).

[0091] The process (200) may generally begin at step (S210) where the process (200) receives a first syntax element indicating the number of sound signals included in the audio scene. The process (200) then proceeds to step (S220).

[0092] At step (S220), the process (200) determines whether the sound signal indicated by the first syntax element includes one or more speech signals.Then, the process (200) proceeds to step (S230).

[0093] At step (S230), the process (200) determines a reference voice signal from one or more voice signals based on the inclusion of one or more voice signals in the sound signal.Then, the process (200) proceeds to step (S240).

[0094] At step (S240), the process (200) adjusts the loudness level of the reference speech signal of the audio scene based on the anchor speech signal.Then, the process (200) proceeds to step (S250).

[0095] At step (S250), the process (200) adjusts the loudness level of the sound signal based on the adjusted loudness level of the reference speech signal.Then, the process (200) terminates.

[0096] In an embodiment, the process (200) receives a second syntax element indicating whether the sound signal includes one or more speech signals. The process (200) determines that the sound signal includes one or more speech signals based on the second syntax element indicating that the sound signal includes one or more speech signals.

[0097] In an embodiment, the process (200) receives a plurality of third syntax elements, each of the third syntax elements indicating whether a corresponding one of the sound signals is a speech signal. The process (200) determines that the sound signals include one or more speech signals based on at least one of the third syntax elements indicating that the corresponding one of the sound signals is a speech signal.

[0098] In an embodiment, the process (200) receives a fourth syntax element indicating the number of one or more speech signals included in the sound signal. The process (200) determines that the sound signal includes one or more speech signals based on the number of the one or more speech signals indicated by the fourth syntax element being greater than zero.

[0099] In an embodiment, the process (200) receives a fifth syntax element indicating a reference speech signal based on the number of the one or more speech signals being greater than one.

[0100] In an embodiment, the process (200) receives a plurality of sixth syntax elements, each of the sixth syntax elements indicating an identification index of a corresponding one of the sound signals.

[0101] In an embodiment, the process (200) determines that the sound signal does not include a speech signal.The process (200) adjusts the loudness level of the sound signal based on a default reference signal.

[0102] III. Loudness Adjustment Device, Computer Device, and Non-Transitory Computer-Readable Storage Medium

[0103] In some embodiments, a loudness adjustment apparatus for an audio scene associated with an MPEG-I immersive audio stream includes: a receiving module for receiving a first syntax element indicating the number of sound signals included in the audio scene; a first determining module for determining whether the sound signal indicated by the first syntax element includes one or more speech signals; a second determining module for determining a reference speech signal from the one or more speech signals based on the inclusion of the one or more speech signals in the sound signal; a first adjusting module for adjusting the loudness level of the reference speech signal of the audio scene based on an anchor speech signal; and a second adjusting module for adjusting the loudness level of the sound signal based on the adjusted loudness level of the reference speech signal.

[0104] In an embodiment, the receiving module is used to: receive a second grammar element indicating whether the sound signal includes one or more speech signals; and the first determining module is used to determine whether the sound signal includes one or more speech signals based on the second grammar element indicating that the sound signal includes one or more speech signals.

[0105] In an embodiment, the receiving module is configured to: receive a plurality of third grammar elements, each of the third grammar elements indicating whether a corresponding one of the sound signals is a speech signal; and the first determining module is configured to determine that the sound signal includes one or more speech signals based on at least one of the third grammar elements indicating that the corresponding one of the sound signals is a speech signal.

[0106] In an embodiment, the receiving module is configured to: receive a fourth grammar element indicating the number of one or more speech signals included in the sound signal; and the first determining module is configured to determine that the sound signal includes one or more speech signals based on the number of the one or more speech signals indicated by the fourth grammar element being greater than zero.

[0107] In an embodiment, the receiving module is configured to receive a fifth syntax element indicating a reference speech signal based on the number of the one or more speech signals being greater than one.

[0108] In an embodiment, the receiving module is configured to: receive a plurality of sixth syntax elements, each of the sixth syntax elements indicating an identification index of a corresponding one of the sound signals.

[0109] In an embodiment, the first determining module is configured to: determine that the sound signal does not include a speech signal; and the apparatus further comprises a third adjusting module configured to adjust the loudness level of the sound signal based on a default reference signal.

[0110] In some embodiments, a method for loudness adjustment signaling of an audio scene associated with an MPEG-I immersive audio stream includes including a first syntax element indicating the number of sound signals included in the audio scene in loudness adjustment information, wherein, in response to determining that the sound signals indicated by the first syntax element include one or more speech signals, a reference speech signal is determined from the one or more speech signals, a loudness level of the reference speech signal of the audio scene is adjusted based on the anchor speech signal, and a loudness level of the sound signal is adjusted based on the adjusted loudness level of the reference speech signal.

[0111] In an embodiment, the method further includes: including a second syntax element indicating whether the sound signal includes one or more speech signals in the loudness adjustment information, wherein the inclusion of the one or more speech signals in the sound signal is determined based on the second syntax element indicating that the sound signal includes one or more speech signals.

[0112] In an embodiment, the method further includes: including a plurality of third syntax elements in the loudness adjustment information, each of the third syntax elements indicating whether a corresponding one of the sound signals is a speech signal, wherein determining that the sound signals include one or more speech signals is based on at least one of the third syntax elements indicating that the corresponding one of the sound signals is a speech signal.

[0113] In an embodiment, the method further includes: including a fourth syntax element indicating the number of the one or more speech signals included in the sound signal in the loudness adjustment information, wherein it is determined that the sound signal includes the one or more speech signals based on the number of the one or more speech signals indicated by the fourth syntax element being greater than zero.

[0114] In an embodiment, the method further comprises: including a fifth syntax element indicating the reference speech signal in the loudness adjustment information based on the number of the one or more speech signals being greater than one.

[0115] In an embodiment, the method further comprises: including a plurality of sixth syntax elements in the loudness adjustment information, each of the sixth syntax elements indicating an identification index of a corresponding one of the sound signals.

[0116] In some embodiments, a computer device includes a processor and a memory. The memory is configured to store program code and transmit the program code to the processor; the processor is configured to execute the above-described method for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream according to instructions in the program code.

[0117] In some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by a computer, cause the computer to perform the above-described method for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream.

[0118] IV. Computer Systems

[0119] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 3 A computer system (300) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0120] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to mechanisms such as assembly, compilation, and linking to create code comprising instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or through interpretation, microcode execution, and the like.

[0121] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, IoT devices, and the like.

[0122] Figure 3 The components shown for the computer system (300) are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Neither should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components shown in the exemplary embodiment of the computer system (300).

[0123] The computer system (300) may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, hand clapping), visual input (e.g., gestures), or olfactory input (not depicted). The human interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0124] Input human interface devices may include one or more of the following (only one of each is depicted): keyboard (301), mouse (302), touchpad (303), touch screen (310), data gloves (not shown), joystick (305), microphone (306), scanner (307) and camera (308).

[0125] The computer system (300) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more senses of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback via a touch screen (310), a data glove (not shown), or a joystick (305), although there may also be tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (309), headphones (not depicted)), visual output devices (e.g., screens (310), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities—some of which may be capable of outputting two-dimensional visual output or more than three-dimensional output through means such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and cigarette cans (not depicted)), and printers (not depicted). These visual output devices (e.g., screens (310)) may be connected to the system bus (348) via a graphics adapter (350).

[0126] The computer system (300) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (320) with CD / DVD etc. media (321), thumb drives (322), removable hard drives or solid-state drives (323), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD based devices such as security dongles (not depicted), and the like.

[0127] Those skilled in the art will also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0128] The computer system (300) may also include a network interface (354) to one or more communication networks (355). The one or more communication networks (355) may be, for example, wireless, wired, or optical. The one or more communication networks (355) may also be local, wide, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of the one or more communication networks (355) include: a local area network such as Ethernet, a wireless LAN, a cellular network including GSM, 3G, 4G, 5G, LTE, etc., a television wired or wireless wide area digital network including cable television, satellite television, and terrestrial broadcast television, a vehicular and industrial network including CANBus, etc. Some networks typically require an external network interface adapter attached to some general-purpose data port or peripheral bus (349) (such as, for example, a USB port of the computer system (300)); other networks are typically integrated into the core of the computer system (300) by attaching to a system bus as described below (e.g., to an Ethernet interface in a PC computer system or to a cellular network interface in a smartphone computer system). Using any of these networks, the computer system (300) can communicate with other entities. Such communication can be one-way, receive-only (e.g., broadcast television), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way (e.g., to other computer systems using a local area digital network or a wide area digital network). Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0129] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core ( 340 ) of the computer system ( 300 ).

[0130] The core (340) may include one or more central processing units (CPUs) (341), graphics processing units (GPUs) (342), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (343), hardware accelerators (344) for certain tasks, etc. These devices, along with read-only memory (ROM) (345), random access memory (346), and internal mass storage devices (347) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (348). In some computer systems, the system bus (348) may be accessed in the form of one or more physical plugs to enable expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus (348) directly or via a peripheral bus (349). Peripheral bus architectures include PCI, USB, etc.

[0131] The CPU (341), GPU (342), FPGA (343), and accelerator (344) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (345) or RAM (346). Transient data can also be stored in RAM (346), while permanent data can be stored in, for example, an internal mass storage device (347). Fast storage and retrieval of any of the memory devices can be achieved by using cache memory, which can be closely associated with one or more CPUs (341), GPUs (342), mass storage devices (347), ROM (345), RAM (346), etc.

[0132] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of a type well known and available to those skilled in the computer software arts.

[0133] By way of example and not limitation, a computer system having architecture (300) and in particular core (340) can provide functionality resulting from the execution of software implemented in one or more tangible computer-readable media by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media associated with a user-accessible mass storage device as described above, as well as certain storage devices of the core (340) having non-transitory properties, such as a core-internal mass storage device (347) or ROM (345). Software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (340). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the core (340) and in particular the processor therein (including a CPU, GPU, FPGA, etc.) to perform specific processing or specific parts of specific processing described herein, including defining data structures stored in RAM (346) and modifying such data structures according to processing defined by the software. Additionally or alternatively, the computer system may provide functionality resulting from logic hardwired or otherwise implemented in circuitry (e.g., accelerator (344)) that may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, reference to software may include logic, and conversely, reference to logic may include software. Where appropriate, reference to a computer-readable medium may include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry implementing logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0134] Although the present disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of the present disclosure. It will therefore be appreciated that those skilled in the art will be able to devise many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. A method for loudness adjustment of an audio scene associated with an MPEG-I immersive audio stream, characterized in that The method comprises: Receive a data structure comprising: a first syntax element indicating the number of sound signals included in the audio scene, a fourth syntax element indicating the number of multiple speech signals included in the sound signal, and a fifth syntax element indicating a reference speech signal; determining, based on a fourth syntax element in the data structure, a plurality of speech signals included in the sound signal indicated by the first syntax element; In response to determining a plurality of speech signals included in the sound signal, determining a reference speech signal from the plurality of speech signals based on a fifth grammar element in the data structure; adjusting the loudness level of the reference speech signal of the audio scene based on an anchor speech signal; and The loudness level of the sound signal is adjusted based on the adjusted loudness level of the reference speech signal.

2. The method according to claim 1, characterized in that Determining the plurality of speech signals included in the sound signal includes determining the plurality of speech signals included in the sound signal based on the number of speech signals indicated by the fourth syntax element being greater than zero.

3. The method according to claim 1, characterized in that The method further comprises: The received data structure includes a plurality of sixth syntax elements, each of the sixth syntax elements indicating an identification index of a corresponding one of the sound signals.

4. The method according to any one of claims 1 to 3, characterized in that determining that the sound signal does not include a speech signal, and Adjusting the loudness level of the sound signal includes adjusting the loudness level of the sound signal based on a default reference signal.

5. A loudness adjustment device for an audio scene associated with an MPEG-I immersive audio stream, characterized in that The device comprises: A receiving module, configured to receive a data structure comprising: a first syntax element indicating the number of sound signals included in an audio scene, a fourth syntax element indicating the number of multiple speech signals included in the sound signal, and a fifth syntax element indicating a reference speech signal; a first determining module, configured to determine, based on a fourth grammar element in the data structure, a plurality of speech signals included in the sound signal indicated by the first grammar element; a second determining module configured to determine, in response to determining a plurality of speech signals included in the sound signal, a reference speech signal from the plurality of speech signals based on a fifth grammar element in the data structure; a first adjustment module, configured to adjust the loudness level of the reference speech signal of the audio scene based on an anchor speech signal; and The second adjustment module is configured to adjust the loudness level of the sound signal based on the adjusted loudness level of the reference speech signal.

6. The device according to claim 5, characterized in that The receiving module is used for: The first determining module is configured to determine a plurality of speech signals included in the sound signal based on the number of speech signals indicated by the fourth grammar element being greater than zero.

7. The device according to claim 5, characterized in that The receiving module is used for: The received data structure includes a plurality of sixth syntax elements, each of the sixth syntax elements indicating an identification index of a corresponding one of the sound signals.

8. The device according to any one of claims 5 to 7, characterized in that: The first determining module is used for: Determining that the sound signal does not include a speech signal; and the device further includes a third adjustment module for The loudness level of the sound signal is adjusted based on a default reference signal.

9. A method for loudness adjustment signaling of an audio scene associated with an MPEG-I immersive audio stream, characterized in that The method comprises: A data structure is received, wherein the data structure includes a first syntax element indicating the number of sound signals included in the audio scene in loudness adjustment information, the data structure includes a fourth syntax element indicating the number of multiple speech signals included in the sound signal in the loudness adjustment information, and the data structure includes a fifth syntax element indicating a reference speech signal in the loudness adjustment information, wherein: determining, based on a fourth syntax element in the data structure, a plurality of speech signals included in the sound signal indicated by the first syntax element, In response to determining a plurality of speech signals included in the sound signal, determining a reference speech signal from the plurality of speech signals based on a fifth grammar element in the data structure, adjusting the loudness level of the reference speech signal of the audio scene based on an anchor speech signal, and The loudness level of the sound signal is adjusted based on the adjusted loudness level of the reference speech signal.

10. The method according to claim 9, characterized in that The method also include: The plurality of speech signals included in the sound signal is determined based on the number of speech signals indicated by the fourth syntax element being greater than zero.

11. The method according to claim 9, characterized in that The method further comprises: The data structure indicates that a plurality of sixth syntax elements are included in the loudness adjustment information, each of the sixth syntax elements indicating an identification index of a corresponding one of the sound signals.

12. An apparatus for loudness adjustment signaling of an audio scene associated with an MPEG-I immersive audio stream, characterized in that The device comprises: a memory storing instructions; and A processor in communication with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 9 to 11.

13. A computer device, characterized in that: The computer device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method according to any one of claims 1 to 4 or the method according to any one of claims 9 to 11 according to the instructions in the program code.

14. A non-transitory computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by a computer, the computer executes the method according to any one of claims 1 to 4 or the method according to any one of claims 9 to 11.

Citation Information

Patent Citations

  • Speech enhancement in entertainment audio

    CN101647059A

  • Loudness level control for audio reception and decoding equipment

    US20170302240A1