Audio level metering for listener position and object position
By simulating the playback process of audio signals in the audio system and using the model of listening areas, the problem of unbalanced audio signal loudness in different locations and environments is solved, and the function of real-time adjustment of loudness is realized, which improves the flexibility and quality of audio creation.
Patent Information
- Application Number
- CN202210452598.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-24
- Filing Date
- 2022-04-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-27
AI Technical Summary
The prior art is difficult to adjust the loudness of the audio signal in real time during audio production to adapt to the changing listener position and sound source position, resulting in the audio work sounding uncomfortable or unbalanced in different locations and environments.
By simulating the playback process of the audio signal from the play position to the listening position in the audio system, the loudness of the audio signal is determined using the model of the listening area, and loudness information is presented on the display for the user to adjust the level of the audio signal and the sound source position.
It realizes real-time adjustment of the loudness of the audio signal during the audio production process, ensuring that the audio works sound balanced and suitable in different locations and environments, and improving the flexibility and quality of audio creation.
Smart Images

Figure CN115250417B_ABST
Abstract
Description
Technical Field
[0001] One aspect of the present disclosure relates to audio level metering. Background Art
[0002] Humans can estimate the position of a sound by analyzing the sound at both of their ears. This is known as binaural hearing, and the human auditory system can use the way sound diffracts and reflects around our bodies and interacts with our pinnae to estimate the direction of the sound. These spatial cues can be artificially generated using spatial filters.
[0003] Spatial filters can be utilized to render audio for playback such that the audio is perceived to have a spatial quality, e.g., originating from a position above, below, or to one side of the listener. The spatial filters can artificially impart spatial cues to the audio that are similar to the diffraction, delay, and reflection naturally caused by our ergonomics and pinnae. The spatially filtered audio can be produced by a spatial audio reproduction system (renderer) and output through speakers (e.g., on headphones).
[0004] In a spatial audio environment, the position of the sound source and / or the listener can vary. The same sound that is further from the listener can be played less loudly than when the sound is closer to the listener. This mimics the actual situation because sound energy attenuates as it travels a certain distance. Thus, as the distance the sound travels increases, the sound pressure level (SPL) of the sound naturally decreases. Summary of the Invention
[0005] In some aspects of the present disclosure, a method for creating an audio level meter can provide loudness relative to changing listener positions and sound source positions. An audio signal is received for measurement. The playback of the audio signal is simulated. The playback is simulated based on the playback position of the audio signal from the viewpoint of the listening position, both within a model of the listening area. Thus, the playback loudness of the audio signal is determined because the loudness perceived by the listener at the listening position is determined.
[0006] The perceived loudness can be affected by the acoustic properties defined by the model of the listening area. Such acoustic properties can include the room geometry, reverberation, acoustic damping of surfaces, furniture, and other objects (e.g., furniture, people, etc.) in the listening area. For example, a small room can have a different reverberation quality than a large room, and the same is true for open spaces. Soft surface materials can absorb more sound energy than hard surface materials. A room with furniture will sound different than a room without furniture. Thus, if the sound source (playback position of the audio signal) is intended to sound as if it is located within the room, the model of the listening area can define different parameters of the listening area as well as the room geometry.
[0007] The perceived loudness can be presented on a display to indicate to a user the loudness of an audio signal at a particular listening position in a listening area. The display can be a standard computer monitor, a tablet, a phone, or the display of other mobile electronic devices, a head-up display, a touchscreen display, and / or other known display types.
[0008] In this way, a user can create a three-dimensional audio or audiovisual experience and view whether the placement of the sound source or listening position is ideal (not too loud or too quiet). The loudness is displayed on a level meter, which can provide guidance to the user when creating an audio or audiovisual work. Then, the user can modify the placement during the creation process and / or increase or decrease the level of the audio signal. The resulting audio or audiovisual work can be produced with a better understanding of how the loudness of the sound source is perceived by the listener, even if its position may change during playback.
[0009] In some aspects, an audio system having a display and one or more processors can perform such processes as described above. The audio system can be integrated with audio or audiovisual production tools (e.g., plugins), such as, for example, digital audio workstations (DAWs), 3D media developer tools, video game developer / editors, movie editors, and / or other content creation tools, to provide an improved audio level meter to the user that takes into account the position of the sound source, the listener, and / or the listening environment. Thus, each audio signal can be perceived by the listener at one location (or an area around that location) in a controlled manner. In some aspects, the audio system is a stand-alone tool. As a stand-alone tool, the system can receive input generated by other tools and / or from the user that specifies a model of the sound source position, the user position, and the listening area.
[0010] The above summary does not include an exhaustive list of all aspects of the present disclosure. It is contemplated that the present disclosure includes all systems and methods that can be practiced with all suitable combinations of the various aspects outlined above and those disclosed in the detailed description below and particularly pointed out in the claims section. Such combinations may have specific advantages not specifically set forth in the above summary of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Aspects of the present disclosure are illustrated by way of example and not limited to the diagrams of the various figures, in which like reference numerals indicate like elements. It should be noted that reference to "an" or "one" aspect in the present disclosure is not necessarily to the same aspect and means at least one. Additionally, for the sake of brevity and to reduce the total number of figures, a given figure may illustrate features of more than one aspect of the present disclosure, and not all elements in the figure may be required for a given aspect.
[0012] Figure 1Shows an audio production process with metering.
[0013] Figure 2 Shows a method for metering according to some aspects.
[0014] Figure 3 Shows an exemplary audio meter with known sound source positions according to some aspects.
[0015] Figure 4 Shows an exemplary audio meter with maximum and minimum loudness indications according to some aspects.
[0016] Figure 5 Shows an exemplary audio meter with varying listener positions and sound source positions according to some aspects.
[0017] Figure 6 Shows an audio system for generating an audio meter on a display according to some aspects.
[0018] Figure 7 Shows an example of the hardware of an audio system according to some aspects. Detailed Description
[0019] Aspects of the present disclosure will now be explained with reference to the accompanying drawings. Whenever the shape, relative position, and other aspects of the described components are not explicitly defined, the scope of the present invention is not limited only to the components shown, and the components shown are only for illustrative purposes. Additionally, although many details are set forth, it should be understood that some aspects of the present disclosure may be practiced without these details. In other instances, well-known circuits, algorithms, structures, and techniques are not shown in detail so as not to obscure the understanding of the description.
[0020] Audio level metering measures the loudness of an audio signal. Traditionally, audio level metering of an audio signal is performed at different stages of an audio production process, but in a static manner that does not take into account a virtual (computer-generated) playback environment, playback location, or listener position.
[0021] Figure 1 Shows an exemplary audio production process with metering. The audio production may include one or more audio signals A - N. Depending on the format of the audio or audiovisual work, an audio signal may be an audio channel dedicated to a speaker (e.g., the left front speaker in a stereo format). If the audio or audiovisual work has an object-based audio format, each audio signal may represent a sound produced by an "object".
[0022] Typically, the levels of audio signals A-C can each be measured separately at the pre-fader metering 4 (before the fader and panning stages 5-7). The levels can be measured again at the post-fader metering 8. The mixing bus 9 can mix the audio signals together according to the output format (in this case, stereo). The output level of the mixing bus can be measured separately again at the output metering 6. Although this example shows the audio signals being mixed into a two-channel stereo, different output formats are possible, such as 5.1, 7.1.4, or object-based audio, such as Dolby Atmos or MPEG-H for example.
[0023] Thus, the levels of audio signals can be measured at different stages to help content creators determine whether the levels fall within a predetermined listening range (e.g., a range that is comfortable, safe, and / or audible for a human listener). However, such meters do not describe the audio levels relative to changing listener positions and / or changing sound source positions. For example, when measuring the audio signals at the different stages described above, a user can create an audio work. However, such individual measurements cannot indicate to the user the levels of the audio signals during playback relative to the listener position or relative to the listening environment.
[0024] The sound pressure level decreases with increasing distance. Nominally, the relationship between distance and sound is 1 / r 2 . This relationship describes how the sound pressure level attenuates based on the distance the sound travels from the source to the listener. Thus, the sound intensity, loudness, or SPL decreases inversely with the square of the distance the sound travels. This distance can be measured from the sound source. Doubling the distance reduces the sound intensity to one-fourth of its initial value.
[0025] Additionally, in a typical listening environment such as a room, there is additional energy caused by the reflection of sound energy off the walls and objects in the room (e.g., furniture, other people). This sound energy is typically described as early reflections and late reflections (e.g., reverberation). The behavior of acoustic reflections can vary based on the room geometry (such as the size and shape of the room). Additionally, the behavior of reflections can be different at different frequencies (even in the same listening area), resulting in different distance attenuation curves at different frequencies.
[0026] Therefore, the acoustic effects of the listening area and the relative positions between the sound source and the listener can greatly affect the perceived loudness of the sound source. If the listener position is different from what the audio work was designed for, or if the listener position changes dynamically during playback, the loudness of the audio signal perceived by the listener may be uncomfortably loud or too small to hear.
[0027] In addition, in some extended reality (XR) environments, the sound source location and the listener location can also change dynamically. Extended reality includes augmented reality, virtual reality, mixed reality, or other immersive technologies that combine the physical world and the virtual (computer-generated) world. Accordingly, it is desirable to indicate to a user that during the production of an audio work, the perceived loudness varies according to the sound source location, the listener location, and the listening environment. The user can then adjust the level of the audio signal, the listener location, the source location, and / or the listening environment based on the meter readings. The adjustments can be made during production rather than learning about undesirable levels after the audio or audiovisual work is completed.
[0028] Figure 2 Method 10 for providing the level of an audio signal in accordance with some aspects is shown. At operation 11, the method includes receiving an audio signal. Depending on the formatting of the audio or audiovisual work, the audio signal can be associated with an object or object-based audio format such as, for example, Dolby Atmos or MPEG-H. In other aspects, the audio signal can be an audio channel for a speaker corresponding to a multi-speaker output format (e.g., 5.1, 7.2, etc.). The audio or audiovisual work can include movies, XR media, video games, teleconferencing applications, songs, and more.
[0029] At operation 12, the method includes simulating the playback of the audio signal from a playback location to a listening location in a model of the listening area. The simulation yields the loudness of the playback of the audio signal that will be perceived at the listening location. The loudness can be expressed as loudness K-weighted full scale loudness (LKFS), root mean square (RMS) loudness, peak loudness, or other loudness metric that reflects the sound pressure level (SPL) perceived by the listener. The loudness can be expressed in decibels (dB).
[0030] At operation 13, the loudness is presented (as a computer image or animation) to a display. In some aspects, the loudness is presented as a line or bar, such as those shown in a meter of Figures 3 to 5 . Additionally or alternatively, the loudness can be presented as a numerical value. The meter / loudness can be presented in different ways without departing from the scope of the present disclosure. In some examples, the loudness or loudness range (e.g., minimum loudness or maximum loudness) can be displayed relative to the frequency of the audio signal. Multiple loudnesses or loudness ranges can be displayed for various frequency bands of the audio signal. The loudness of one frequency component can be different from that of another frequency component in the same audio signal at a given location.
[0031] It should be understood that the method 10 can be repeated for one or more additional audio signals. For example, the sound scene of an object-based audio work can include an airplane flying above the head, a barking dog on the right, and a person speaking in front of the listener's position. In a multi-speaker format, different audio signals can represent different speaker channels. Regardless of the audio format, each audio signal can be associated with a different playback position. Thus, with respect to the listener position (and / or other listener positions) in the listening area, the method can be repeated for each of one or more additional audio signals at the corresponding playback positions. The position of the sound source or the listener can include its position (e.g., in 2D or 3D coordinates) and orientation (e.g., spherical coordinates).
[0032] Figure 3 An exemplary audio meter with known sound source positions according to some aspects is shown. The audio mixer can be played on physical speakers. The amplification of the physical speakers can be calibrated for the optimal listening position and / or average level in a series of seats. Different surround sound speaker formats have predefined optimal speaker positions.
[0033] For example, for a surround sound speaker format (e.g., 5.1, 7.2, etc.), the optimal speaker positions of the right speaker 20, center speaker 22, left speaker 24, subwoofer 24, right surround speaker 28, and left surround speaker 26 are known or predefined. The amplification (e.g., gain value) of each audio channel driving the corresponding speaker can be defined for the optimal listening position (e.g., listening position 34). Then, other listening positions (e.g., 38 and 36) will be subject to the speaker output that has been customized for position 34, and thus they will not be optimal. For example, relative to position 34, position 38 is closer to the right speaker 20 and farther from the left surround speaker 26. Thus, for position 38, the right speaker 20 may sound stronger and the left surround speaker 26 may sound weaker than at position 34. The meter 32 can be presented to the content producer on the display 30. The meter can show the loudness of one of the listening positions. In some embodiments, multiple meters can be presented on the display, each showing the loudness of one of the listening positions. For example, a first meter can show the loudness measured at position 34, a second meter can show the loudness measured at position 36, and a third meter can show the loudness measured at position 38. The content producer can analyze the loudness levels of each speaker for multiple listening positions and adjust the levels accordingly. The loudness of each channel can be balanced for one or more listening positions.
[0034] Figure 4An exemplary audio meter with maximum and / or minimum loudness indication is shown in accordance with some aspects. The audio level meter 34 may show the loudness of the sound source 34 at a defined listening position. The minimum loudness may be determined based on a second listening position. The maximum loudness may be determined based on a third listening position. The minimum loudness and / or the maximum loudness may be presented to a display. The maximum loudness and the minimum loudness may show the range of loudness shown relative to a series of listening positions. The audio meter may include or be integrated with a graphical user interface (GUI) that employs user input (e.g., via a mouse, touch screen, keypad, or other user input device).
[0035] In some aspects, a listening position defining area 33 may be defined around the listening position. The area may be circular, square, triangular, or irregular in shape. The minimum loudness may represent the listening position where the sound source is weakest in the area. The maximum loudness may represent the listening position where the sound source is strongest (e.g., loudest) in the area. In some examples, the audio meter may adjust the listening position, the second listening position, and / or the third listening position in response to a change in the level of the audio signal being played. For example, the audio meter may be part of a user interface that receives user input to increase or decrease the level of the audio signal being played. The second listening position and / or the third listening position may be automatically adjusted in response to a change in level. For example, if the level increases, the second and third positions may move away from the sound source 34. Similarly, if the level decreases, the second and third positions may move closer to the sound source 34. More generally, the audio meter may determine one or more loudnesses that would be perceived at various different listening positions and display those loudnesses simultaneously. For example, the audio meter may obtain user input specifying a plurality of different listening positions within a given listening area. The audio meter may simulate the playing of the audio signal at each of the different listening positions and present each loudness to the display. Additionally, the audio meter may obtain user input to adjust any one of the listening positions. In response to the user input, the audio meter may adjust the loudness based on the adjusted listening position. Additionally, the audio meter may obtain user input to adjust the level of the signal. In response to the user input, the audio meter may adjust any or all of the listening positions according to the adjusted level of the signal, as discussed. Additionally, the audio meter may determine a range (e.g., maximum loudness and minimum loudness and the corresponding listening positions) for each of the listening positions. Thus, relative to a given audio signal, a user may audition and use the audio meter to compare different listening positions. The audio meter may dynamically adjust one or more listening positions or their loudness based on the level of the audio signal or the frequency content of the audio signal.
[0036] In some aspects, the loudness of a sound source at a listening position is indicated by a line or bar. The length of the bar indicates the loudness of the audio signal. In some aspects, a minimum loudness and a maximum loudness are presented as a second bar or line having a start based on the minimum loudness and an end based on the maximum loudness.
[0037] For example, a range indicator 35 starts at a minimum loudness (e.g., 65 dB) and ends at a maximum loudness (e.g., 80 dB). Thus, the meter shows the listening position loudness, as well as the range of loudness heard at positions surrounding the listening position. Such a meter can show the level at which an audio signal will be perceived by a listener at a given position, as well as the maximum / minimum levels in a given area around the listener. Thus, in a surround sound speaker environment, when a listener changes position during playback, a user (e.g., a content creator) will not be surprised by the loudness of the content, nor by the loudness of the content heard by the listener at various positions within the intended listening area 33.
[0038] In some aspects, the meter can indicate a range based on the frequency composition of the measured signal. For example, different levels or ranges of levels can be shown for one or more frequencies or frequency bands.
[0039] Figure 5 An exemplary audio meter with varying listener positions and sound source positions is shown in accordance with some aspects. In a spatial audio environment, one or more sound sources are spatially filtered to give the appearance of sound at positions around the listener. These sound sources (e.g., sound source 37) can change position over time. For example, an audio production can include a sound that moves from the left of the listener to the right. Thus, the loudness of the sound source perceived by the listener at the listener position 38 can vary over time.
[0040] In addition, in an immersive environment such as extended reality (XR), the listener's position can also change over time. For example, sensors (e.g., cameras, gyroscopes, and / or accelerometers) can be used with tracking algorithms (e.g., visual odometry, SLAM) to track the listener's head position. The sound source level perceived by the listener is dynamic. The level can depend on the relative positions of the sound source and the listener, the listener's environment 40, and the audio signal associated with the sound source. Thus, content creators can use 3D content creation tools (such as, for example, Unity, Unreal Engine, CRYENGINE, or other equivalent technologies) to "move" the assumed sound source position or the listener position around. In response, a level meter 43 can be presented to show the level at the listening position 38 and / or the maximum / minimum level relative to the area 39 surrounding the listening position. Content creators can see these levels during production and make necessary adjustments during production, rather than making corrections after completion. Content creators can adjust different parameters, such as, for example, the position of the sound source and / or the listener, the amplification (e.g., gain) associated with the sound source, and / or the model of the listening environment.
[0041] The model of the listening area can include a geometric definition of the listening area. If the listening area is in a room (e.g., a virtual room), the model of the listening area can include a room model 42. The room model can include the room shape, the length and width of the walls, and / or the overall volume. The room model can include the surface materials present in the listening area (e.g., cloth, stone, carpet, cement, hardwood, etc.), the acoustic attenuation coefficient (describing the absorption or scattering of sound through the propagation path), the sound absorption coefficient, the reverberation time, and / or the objects in the listening area. The room model can include a CAD model of the room or other computer-defined 3D model. The room model can include the objects located in the room, such as doorways, furniture, and / or other people.
[0042] In some aspects, one or more parameters, such as the sound source position, the listener position, the listening environment model, and / or the level of the audio signal, can be automatically adjusted based on the desired perceived loudness of the playback of the audio signal at the listening position. A threshold can represent the maximum level and / or the minimum level. Adjustments can be made to keep the level below the maximum level and / or above the minimum level.
[0043] For example, if the perceived loudness of the playback of the audio signal at the listening position meets a threshold (e.g., 44), the loudness of the audio signal associated with the sound source 37 can be automatically adjusted by the system. The sound source can be moved away from the listener to decrease the loudness, or it can be moved toward the listening position to increase the loudness.
[0044] Additionally or alternatively, if the playback loudness of the audio signal at the listening position meets a threshold, the listener position can be automatically adjusted. The listener position can be moved away from the sound source to reduce the loudness, or towards the sound source to increase the loudness.
[0045] Additionally or alternatively, if the perceived loudness of the playback of the audio signal at the listening position meets a threshold, the model of the listening environment (e.g., the room model) can be adjusted. For example, the room can be made larger to reduce the loudness caused by reflections, or the room can be made smaller to increase the loudness caused by reflections. The sound absorption and / or acoustic attenuation coefficient can be increased or decreased to reduce or increase the loudness caused by reflections.
[0046] Settings that can be configured by the user by default and / or modified can control whether the automatic adjustment is to be performed by the system. Additionally or alternatively, the settings can determine which parameters should be automatically adjusted in response to the threshold being met. In some aspects, the settings can define a hierarchy, e.g., first adjust the sound source position, then the listening environment, then the gain associated with the audio signal, and then the listener position.
[0047] Figure 6 An audio system for generating an audio meter for a display is shown in accordance with some aspects. The audio system can perform the operations and methods described herein, such as with respect to Figure 2 Method 10 described.
[0048] The audio signal 50 represents the audio of the sound source or the audio channel for driving the speaker. The audio signal can be one or more audio signals for each part of a common audio or audiovisual work. The audio signal can vary over time and frequency bands. The audio signal can be a time-domain or frequency-domain (e.g., STFT) audio signal.
[0049] The meter signal data generator 56 can determine or select an appropriate impulse response 58 based on the sound source position 52 relative to the listener position 54. When applied to the audio signal, each impulse response is capable of imparting spatial cues (e.g., frequency-dependent gain and delay) to the audio signal that simulate how the human body and ears affect the audio, thus simulating the natural physics when sound waves propagate from the sound source to the human ear. The impulse response can include one or more room impulse responses determined or selected based on the model 62 of the listening area.
[0050] For example, room impulse responses can be used to simulate small rooms, large rooms, concert halls, open spaces, etc. by characterizing the sound energy caused by the reflection and / or scattering of sound in the corresponding environment. The impulse response is determined or selected to be incorporated into an audio signal, representing a model with respect to the listening area of how sound from a playback location is perceived at a listening location. The audio signal can be convolved with the impulse response to simulate the playback of the audio signal at the listening location and includes the acoustic characteristics of the listening area (e.g., early acoustic reflections and late acoustic reflections).
[0051] The model of the listening area can be stored as metadata describing the reverberation time, scattering parameters, absorption parameters, surface materials, and / or a full geometric model of the listening area.
[0052] The meter signal data generator can measure the resulting audio signal to determine its level and / or level range at the listening location. As discussed, the audio source location 52 can be one or more audio source locations. Similarly, the listener location can be one or more listener locations. For each different location of the audio source relative to the listener, the corresponding impulse response can be determined of how the simulated sound travels from the source (directly and / or indirectly) to the listener.
[0053] The meter signal provides one or more levels to the renderer and display 60. The renderer and display can produce one or more graphics representing the level meter. The level meter can use visual indicators to represent the loudness and / or loudness range at the listening location, which can include one or more shapes (e.g., bars, pointers, circles, curves, etc.), symbols (e.g., numbers, letters, etc.), and / or other visual indicators. In some aspects, the visual indicator can indicate loudness based on color, light intensity, or a combination thereof. In some aspects, as shown, the level meter can be shown as a symbol (e.g., a numerical value). In some aspects, the level meter can be presented as a rotating pointer. For example, the pointer can rotate around a pivot to indicate the loudness value. Without departing from the scope of the present disclosure, other visual shapes or symbols indicating the audio level can be rendered. The renderer and display can include various electronic display systems, as described in other sections.
[0054] Figure 7 An example of an audio processing system 150 is shown according to some aspects. The audio processing system can be a computing device, such as, for example, a desktop computer, a tablet, a smartphone, a laptop computer, a smart speaker, a media player, headphones, a head-mounted display (HMD), smart glasses, an infotainment system for an automobile or other vehicle, or an electronic device configured to present XR. The system can be configured to perform the methods and processes described in the present disclosure. In some aspects, such as Figure 6 the system shown is implemented as one or more audio processing systems.
[0055] Although various components of an audio processing system that can be incorporated into headphones, speaker systems, microphone arrays, and entertainment systems are shown, this illustration is merely an example of a particular specific implementation of the types of components that can exist in an audio processing system. This example is not intended to represent any particular architecture or manner of interconnecting these components, as such details are not closely related to the aspects described herein. It should also be understood that other types of audio processing systems with fewer or more components than those shown may also be used. Thus, the processes described herein are not limited to use with the hardware and software shown.
[0056] The audio processing system 150 may include one or more buses 162 for interconnecting the various components of the system. As is well known in the art, one or more processors 152 are coupled to the bus 162. The one or more processors may be microprocessors or specialized processors, system-on-a-chip (SOC), central processing units, graphics processing units, processors created by application-specific integrated circuits (ASICs), or combinations thereof. The memory 151 may include read-only memory (ROM), volatile memory, and non-volatile memory or combinations thereof coupled to the bus using techniques known in the art. The sensor / head tracking unit 158 may include an IMU and / or one or more cameras (e.g., RGB cameras, RGBD cameras, depth cameras, etc.) or other sensors described herein. The audio processing system may also include a display 160 (e.g., an HMD or a touchscreen display).
[0057] The memory 151 may be connected to the bus and may include DRAM, a hard disk drive, or flash memory, or a magneto-optical drive or magnetic memory, or an optical drive or other type of memory system that maintains data even after the system is powered off. In one aspect, the processor 152 retrieves computer program instructions stored in a machine-readable storage medium (the memory) and executes these instructions to perform the methods and other operations described herein.
[0058] Although not shown, audio hardware may be coupled to one or more buses 162 to receive audio signals to be processed and output by the speakers 156. The audio hardware may include a digital-to-analog converter and / or an analog-to-digital converter. The audio hardware may also include an audio amplifier and filters. The audio hardware may also be connected to a microphone 154 (e.g., a microphone array) to receive audio signals (whether analog or digital), digitize them if necessary, and transmit the signal to the bus 162.
[0059] The communication module 164 can communicate with remote devices and networks. For example, the communication module 164 can communicate via known technologies such as Wi-Fi, 3G, 4G, 5G, Bluetooth, ZigBee, or other equivalent technologies. The communication module can include a wired or wireless transmitter and receiver that can communicate (e.g., receive and send data) with networked devices such as servers (e.g., in the cloud) and / or other devices such as remote speakers and remote microphones.
[0060] It should be understood that the aspects disclosed herein can utilize a memory remote from the system, such as a network storage device coupled to the audio processing system via a network interface such as a modem or an Ethernet interface. As is well known in the art, the buses 162 can be connected to each other through various bridges, controllers, and / or adapters. In one aspect, one or more network devices can be coupled to the bus 162. The one or more network devices can be wired network devices (e.g., Ethernet) or wireless network devices (e.g., WI-FI, Bluetooth). In some aspects, the various aspects (e.g., analog, analysis, estimation, modeling, object detection, etc.) can be performed by a networked server in communication with the capture device.
[0061] The various aspects described herein can be embodied, at least in part, in software. That is, these techniques can be implemented in the audio processing system in response to its processor executing a sequence of instructions contained in a storage medium such as a non-transitory machine-readable storage medium (such as DRAM or flash memory). In various aspects, hardwired circuitry can be used in combination with software instructions to implement the techniques described herein. Thus, these techniques are not limited to any specific combination of hardware circuitry and software, nor to any particular source of the instructions executed by the audio processing system.
[0062] In this specification, certain terms are used to describe features of the various aspects. For example, in some cases, the terms "module", "processor", "unit", "renderer", "system", "device", "filter", "sensor", "display", and "component" denote hardware and / or software configured to perform one or more processes or functions. For example, examples of "hardware" include, but are not limited to, integrated circuits such as processors (e.g., digital signal processors, microprocessors, application-specific integrated circuits, microcontrollers, etc.). Thus, as understood by those skilled in the art, different combinations of hardware and / or software can be implemented to perform the processes or functions described by the above terms. Of course, the hardware can alternatively be implemented as a finite state machine or even combinational logic components. Examples of "software" include applications, applets, routines, or even executable code in the form of a sequence of instructions. As described above, the software can be stored in any type of machine-readable medium.
[0063] Certain portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the audio processing art to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, a self-consistent sequence of operations leading to a desired result. These operations are those requiring physical manipulation of physical quantities. It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, it is evident from the above discussion that throughout the specification, discussions using terms such as those set forth in the claims below are referring to the actions and processes of an audio processing system or like electronic device that manipulates data represented as physical (electronic) quantities within the registers and memories of the system and transforms them into other data similarly represented as physical quantities in the system memory or registers or other such information storage, transmission, or display devices.
[0064] The processes and blocks described herein are not limited to the specific examples described and are not limited to the specific order used as examples herein. Instead, any processing block may be reordered, combined, or removed as needed and executed in parallel or serially to achieve the above results. The processing blocks associated with implementing an audio processing system may be executed by one or more programmable processors executing one or more computer programs stored on a non-transitory computer-readable storage medium to perform the functions of the system. All or part of the audio processing system may be implemented as dedicated logic circuitry (e.g., FPGA (Field Programmable Gate Array) and / or ASIC (Application Specific Integrated Circuit)). All or part of the audio system may be implemented using electronic hardware circuitry including at least one of electronic devices such as, for example, a processor, a memory, a programmable logic device, or a logic gate. Additionally, the processes may be implemented in any combination of hardware devices and software components.
[0065] Although certain aspects have been described and illustrated in the drawings, it should be understood that these aspects are merely illustrative of the invention and not limiting, and the invention is not limited to the specific structures and arrangements shown and described, as various other modifications may occur to those of ordinary skill in the art.
[0066] To assist the Patent Office and any readers of any patent issued on this application in interpreting the appended claims, the applicant wishes to note that they do not intend any of the appended claims or claim elements to invoke 35 U.S.C. 112(f) unless the words "means for" or "step for" are expressly used in a particular claim.
[0067] As is well known, the use of personally identifiable information should comply with privacy policies and practices that are recognized as meeting or exceeding industry or government requirements for maintaining user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of inadvertent or unauthorized access or use, and the nature of authorized use should be clearly explained to users.
[0068] Cross - Reference to Related Applications
[0069] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 180,354, filed on April 27, 2021, which is hereby incorporated by reference in its entirety.
Claims
1. A computer-implemented method for audio level metering, comprising: Receiving an audio signal; Simulating the playback of the audio signal from a playback position to a listening position in a model of a listening area, generating the loudness of the playback of the audio signal that would be perceived at the listening position; And Presenting the loudness of the playback of the audio signal to a display, Wherein the method further comprises the operation of: modifying the listening position if the loudness of the playback of the audio signal at the listening position meets a threshold.
2. The method according to claim 1, wherein the loudness is presented as a line or bar having a length depending on the loudness.
3. The method according to claim 1, further comprising determining a minimum loudness based on a second listening position, and determining a maximum loudness at a third listening position, and presenting the minimum loudness and the maximum loudness to the display.
4. The method according to claim 3, wherein the loudness, the minimum loudness or the maximum loudness is determined and presented relative to the frequency of the audio signal.
5. The method according to claim 3, wherein the minimum loudness and the maximum loudness are presented as a second bar or line having a start based on the minimum loudness and an end based on the maximum loudness.
6. The method according to claim 5, further comprising adjusting the listening position, the second listening position or the third listening position in response to a change in the level of the playback of the audio signal.
7. The method according to claim 1, wherein simulating the playback includes convolving the audio signal with a room impulse response that incorporates how sound from the playback position is perceived at the listening position relative to the model of the listening area.
8. The method according to claim 1, wherein the model of the listening area includes room geometry.
9. The method according to claim 1, wherein the model of the listening area includes at least one of the following: surface material, acoustic damping coefficient, reverberation time, or objects in the listening area.
10. The method according to claim 1, wherein the loudness includes at least one of loudness K-weighted full scale (LKFS), root mean square (RMS) loudness, or peak loudness.
11. The method according to claim 1, wherein the loudness is presented as a numerical value.
12. The method according to claim 1, wherein the performance of the method is integrated as part of a digital audio workstation (DAWS), a game developer application, or a 3D media developer tool.
13. The method according to claim 1, wherein the method further comprises at least one of the following operations: (i) modifying the loudness of the audio signal if the loudness of the playback of the audio signal at the listening position meets a threshold; or (ii) modifying the room geometry of the listening area if the loudness of the playback of the audio signal at the listening position meets a threshold.
14. The method according to claim 1, wherein the listening position or the playback position changes over time, and the method is repeated for one or more additional audio signals.
15. The method according to claim 1, wherein the audio signal is an audio channel of a speaker corresponding to a multi-speaker output format or is associated with a sound source object of an object-based audio format.
16. The method according to claim 1, further comprising simulating the playback of the audio signal from the playback position to one or more second listening positions, generating one or more second loudness levels of the playback of the audio signal that would be perceived at the one or more second listening positions, and presenting the one or more second loudness levels of the playback of the audio signal to the display.
17. The method according to claim 16, wherein the playback position or the one or more second listening positions are determined or adjusted based on a user input, and the loudness or the one or more second loudness levels are updated based on the adjustment of the playback position or the one or more second listening positions.
18. An audio system, comprising: a display; and a processor configured to perform the operations of the method according to any one of claims 1 to 17.
19. A computer-readable storage medium having computer program instructions stored thereon, the computer program instructions, when executed by a processor of an audio system, cause the processor to perform the operations of the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Real-time sound propagation for dynamic sources
US20110081023A1