Determine the virtual listening environment
By determining the acoustic environment parameters of the user and the audio file through the audio processing system and applying spatial filters to render the audio signal, the problem of acoustic effect mismatch in the virtual listening environment is solved, resulting in a better listening experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2022-09-21
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies struggle to achieve an effective compromise between preserving the original acoustic effects of audio files and the user's current acoustic environment in virtual listening environments, resulting in a mismatch between virtual and real acoustic effects and impacting the listening experience.
The audio processing system determines the user's current acoustic environment parameters based on sensor signals, and combines this with the acoustic environment of the audio file to select or generate preset acoustic parameters. It then applies spatial filters to render the audio signal, thereby achieving a compromise between the virtual environment and the real environment.
It enhances the realism of the virtual listening environment while maintaining the original acoustic effects of the audio files, thus improving the realism and fidelity of the listening experience.
Smart Images

Figure CN115842984B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 246484, filed September 21, 2021, which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to audio processing. Background Technology
[0004] Content creators can create audio or audiovisual works. Audio can be precisely fine-tuned to suit the creator's taste, delivering a specific experience to the listener. Creators can meticulously craft audio to evoke perceptible cues associated with a specific setting, such as an echoing outdoor hillside, a stadium, or a small enclosed space. Audio recorded outdoors can have perceptible acoustic cues that create the feeling of being outdoors for the listener. Similarly, if an audio work is recorded indoors, it can virtually create the feeling of being in that room for the listener.
[0005] Users can listen to audio works in various locations. Each location can have a different acoustic environment. For example, a user can listen to audio or audiovisual works in a car, on a grassy field, in a classroom, on a train, or in a living room. Each acoustic environment around the user can carry expectations of how the sound should be heard, even if the sound is produced by the headphones the user is wearing. Summary of the Invention
[0006] In one aspect, the method executed by the processor includes: determining one or more acoustic parameters of the user's current acoustic environment based on sensor signals captured by one or more sensors of the device; determining one or more preset acoustic parameters based on the one or more acoustic parameters of the user's current acoustic environment and the acoustic environment of an audio file including an audio signal, the acoustic environment of which is determined based on the audio signal or metadata of the audio file; generating a binaural audio signal by spatially rendering the audio signal by applying the one or more preset acoustic parameters to the audio signal; and using the binaural audio signal as input to drive a loudspeaker. In this way, a trade-off can be achieved between the user's current acoustic environment and the acoustic environment of the audio file.
[0007] In one aspect, the method performed by the device's processor includes: determining whether an audio file or audiovisual file includes metadata containing an acoustic environment for playback. In response to the audio file or audiovisual file including metadata containing the acoustic environment, the processor spatially renders an audio signal associated with the audio file or audiovisual file based on the acoustic environment of the metadata. In response to the audio file or audiovisual file not containing metadata, the processor spatially renders the audio signal based on one or more acoustic parameters of the user's current environment, the current scene of the audio file or audiovisual file, and / or the content type of the audio file or audiovisual file. In this way, content creators can precisely control the acoustic environment through metadata; however, if metadata is absent, a trade-off can be struck between the user's current acoustic environment and the acoustic environment of the audio file.
[0008] The above overview does not constitute an exhaustive list of all aspects of this disclosure. It is contemplated that this disclosure encompasses all systems and methods that can be practiced by all suitable combinations of the aspects outlined above and those disclosed in the detailed descriptions below and specifically pointed out in the claims section. Such combinations may have specific advantages not specifically set forth in the foregoing summary. Attached Figure Description
[0009] To facilitate identification of any particular element or action being discussed, one or more of the most significant digits in the reference numerals refer to the drawing number in which the element was first introduced.
[0010] Figure 1 This demonstrates audio processing based on various aspects of audio or audio-visual files.
[0011] Figure 2 The selection of acoustic parameters for audio processing is shown based on several aspects.
[0012] Figure 3 The workflow for determining and applying preset acoustic parameters is shown based on several aspects.
[0013] Figure 4 A method for determining preset acoustic parameters based on several aspects is shown.
[0014] Figure 5 A method for determining preset acoustic parameters based on several aspects is shown.
[0015] Figure 6 The audio processing operations used to determine preset acoustic parameters are illustrated based on several aspects.
[0016] Figure 7 Metadata with acoustic parameters is shown based on some aspects.
[0017] Figure 8 An audio processing system based on some aspects is shown. Detailed Implementation
[0018] Humans can estimate the location of a sound by analyzing the sound at both of their ears. This is called binaural hearing, and the human auditory system estimates the direction of a sound by how it diffracts and reflects away from our body and interacts with our auricles.
[0019] Microphones sense sound by converting changes in sound pressure into electrical signals using an electroacoustic transducer. These electrical signals can be digitized using an analog-to-digital converter (ADC). Spatial filters can be used to render audio for playback, making it perceived as having spatial characteristics. Spatial filters artificially imbue audio with spatial cues, similar to diffraction, delay, and reflection naturally caused by our ergonomics and the shape of our ears. Spatially filtered audio can be produced by a spatial audio reproduction system and output through headphones.
[0020] Spatial audio reproduction systems with headphones can track a user's head movements. Binaural filters can be selected based on the user's head position and continuously updated as the head position changes. These filters are applied to the audio to maintain the illusion that the sound is coming from some desired location in space. These spatial binaural filters are called head-related impulse response (HRIR).
[0021] A listener's ability to estimate distance (not just relative angles), especially in an indoor space, is related to the level of the direct portion of the signal (i.e., without reflections) relative to the level of reverberation (with reflections). This relationship is known as the direct-to-reverberation ratio (DRR). In a listening environment, reflections are caused by sound energy bouncing off one or more surfaces (e.g., walls or objects) before reaching the listener's ear. Within a room, a single sound source can produce numerous reflections from different surfaces at different times. The sound energy from these reflections (which can be understood as reverberation) may gradually increase and then decrease over time.
[0022] Reverberation helps create the strong illusion that the sound originates from a source within the room. Thus, the spatial filter and the binaural cues assigned to the left and right output audio channels should include some reverberation. This reverberation can be shaped by the presence of people and the nature of the room, and can be described by a set of binaural room impulse responses (or BRIRs).
[0023] Robust virtual acoustic simulations (e.g., spatial audio) greatly benefit from virtualizing a room to induce a sense of externalization of sound, which can be understood as feeling that the sound is not coming from headphones, but from the outside world. Determining the acoustic parameters of a virtual room is crucial for providing a convincing spatial audio experience.
[0024] Generally speaking, the more a virtual room's acoustics closely resemble those of a real room in which a person operates, the more realistic the externalization will be. However, when reproducing pre-recorded audio content such as movies, podcasts, music, or other content, using spatial audio to simulate a real room can be detrimental to the experience because the virtual room's acoustics may be stronger or perceptually different from the recorded content's. A typical example of this is an outdoor movie scene, where users might expect little to no reverberation, but due to virtualization, they might hear a lot of reverberation from the virtual room. In such cases, a trade-off or compromise can be struck between reproducible realism (which contributes to externalization and a sense of immersion) and reproducible fidelity (which maintains the viewing experience as the content creator intended).
[0025] In some respects, the system can select the parameters of the best virtual room preset or reverberation algorithm based on the analysis and / or prior knowledge of the acoustic effects of the real room and the acoustic effects of the content being played back.
[0026] Figure 1 Audio processing of audio or audiovisual files is illustrated according to some aspects. The audio processing system 116 may include sensors 122, such as microphones, cameras, or other sensors that can characterize the acoustic environment 120 of the user 112. The audio processing system may be integrated within devices such as headsets 104 and / or other computing devices 110. In some aspects, the computing device may be a laptop, mobile phone, tablet, smart speaker, media player, or other computing device. The computing device may include a display. In some aspects, the audio processing system may be distributed among more than one computing device. The audio processing system can sense the user's acoustic environment. For example, the audio processing system may be integrated within device 110 or 104, which is located in an outdoor field or in a living room with a user. Sensor 122 may generate microphone signals of sensed sound and acoustic properties along with the space in which it is carried.
[0027] Audio file 102 may contain audio signal 106 and metadata 114. In some respects, the audio file may also be an audiovisual file that includes video signal 108. Audio processing system 116 may determine preset acoustic parameters 118 to be applied to spatially render audio signal 106. The preset acoustic parameters may be determined or selected as a trade-off between the user's acoustic environment and the acoustic environment of the audio file.
[0028] Audio processing system 116 can determine one or more acoustic parameters of a user's current acoustic environment based on microphone signals, video signals, or other sensed data. The audio processing system can apply one or more audio processing algorithms to the microphone signals to extract acoustic parameters of the user's environment. These acoustic parameters may include at least one of the following specific to the user's current environment: reverberation time, direct-to-reverberation ratio (DRR), reflection density, immersion, or speech intelligibility. Therefore, if the user repositions to a new space, these acoustic parameters can be updated to characterize the user's new acoustic environment. Such acoustic parameters can be determined repeatedly, such as periodically or in response to changes in the environment, to update how the audio is spatialized in response to real-time changes in the user's environment.
[0029] The audio processing system can determine or select one or more preset acoustic parameters 118 based on various factors, such as one or more acoustic parameters of the user's current acoustic environment and the acoustic environment of the audio file 102. In some cases, the acoustic environment of the audio file can be artificially created, for example, using software. The audio file or audiovisual file 102 may include, for example, songs, podcasts, radio programs, movies, performances, recorded concerts, video games, or other audio or audiovisual works. Different acoustic environments can have different acoustic parameters, such as reverberation, DRR, echo, reflection density, immersion, or speech intelligibility. For example, audio signals 106 can be recorded in acoustic environments such as outdoor sets, indoor sets, cathedrals, storage rooms, and other acoustic environments, each with unique acoustic characteristics.
[0030] The acoustic environment of an audio file can be determined based on its audio signal or metadata; this can be referred to as a content-based acoustic environment. The audio processing system 116 can apply one or more algorithms to the audio signal 106 to extract its acoustic parameters. The audio processing system can classify the acoustic environment in which the audio signal 106 is recorded (or artificially created). For example, the audio processing system can classify the file's acoustic environment as "outdoor" or "indoor." The acoustic environment can be classified at a greater granularity; for example, the audio file's acoustic environment can be classified as "large room," "medium room," or "small room." The acoustic environment of an audio file can change from one scene to another. For example, at the beginning of the audio file, the scene can be outdoor. The scene can change to indoor in the middle of the audio file and then return to outdoor. Similarly, the scene can change from an indoor room to different rooms with different geometries, sizes, and / or damped surfaces.
[0031] Audio or audiovisual files may contain metadata 114 that describes one or more scenes of a work. These scenes may change throughout the work. For example, the metadata may specify a first scene as "outdoor" and a second scene as "indoor." Audio processing system 116 may read the metadata to classify the acoustic environment of the audio file. For example, if the metadata indicates the current scene is "outdoor," the content-based acoustic environment may be classified as "outdoor." If the metadata indicates the current scene is "indoor," the content-based acoustic environment may be classified as "indoor." If the metadata indicates the current scene is "bedroom," the content-based acoustic environment may be classified as "bedroom." In some aspects, the metadata may specify that the current scene will be "the acoustic environment of user 112," in which case the audio processing system can determine the user's acoustic environment from microphone signals or other sensor data, as described. In some aspects, the metadata may include one or more preset acoustic parameters 118 to be used by the audio processing system, which may also change depending on the scene.
[0032] In some respects, the audio processing system can analyze video signal 108 to determine the content-based acoustic environment. For example, the audio processing system can use computer vision (e.g., trained machine learning algorithms) to analyze the video to determine whether the content is displaying an outdoor or indoor scene, a large room, a medium-sized room, a small room, a stadium, or other acoustic environment. Similarly, the audio processing system can apply computer vision algorithms to images to determine the acoustic environment of user 112.
[0033] Audio processing can select or determine preset acoustic parameters 118 based on the acoustic environment of user 112, and the content-based acoustic environment, as described, can be categorized based on metadata 114, audio signal 106, and / or video signal 108. In some cases, the audio processing system may use preset acoustic parameters that are more closely similar to the content-based acoustic environment, while in others, it may use preset acoustic parameters that are more closely similar to the acoustic environment of user 120. This can depend on various factors, as further described in other sections.
[0034] In some respects, the audio processing system may first scan metadata indicating the content-based acoustic environment. If such metadata exists, the audio processing can determine preset acoustic parameters based on the metadata. If it does not exist, the audio processing system can backtrack by analyzing the audio signal 106 and / or the video signal 108 to determine the preset acoustic parameters.
[0035] An audio processing system can spatially render one or more audio signals, which may include applying one or more preset acoustic parameters to the audio signals. For example, the audio processing system may convolve the audio signals using spatial filters characterizing the head correlation transfer function (HRTF). These spatial filters may include preset acoustic parameters such as reverberation time, DRR, or other preset acoustic parameters. The audio processing system can then use the resulting binaural audio signals to drive the speakers of the headset 104.
[0036] Figure 2 A system for selecting acoustic parameters for audio processing is illustrated, based on several aspects. As discussed, a trade-off or compromise can be made between reproducibility and fidelity. The more externalized and surrounded the user in the audio, the more realistic or convincing the spatial reproduction appears to the user. However, such rendering may not meet the intended acoustic requirements. Preset selector 212 can determine the trade-off between making the spatial audio realistic and maintaining the original viewing experience as intended by the content creator.
[0037] The preset selector 212 can determine one or more preset acoustic parameters 216. The preset acoustic parameters can be associated with each of a plurality of preset acoustic environments 214. The preset acoustic environments can be stored and accessed in a computer-readable medium. Such a library of preset acoustic environments can include various environments, such as, for example, large rooms, small rooms, rooms with different geometries, rooms with different surface-absorbing surfaces, rooms with various arrangements of objects (e.g., furniture), outdoor spaces and echoing outdoor spaces, cathedrals, stadiums, libraries, living rooms, bedrooms, or other acoustic environments. Preset acoustic environments can be determined at an earlier time, stored in memory, and recalled when processing audio files. Each preset acoustic environment in the preset acoustic environment can include corresponding preset acoustic parameters. For example, a large room may have a long reverberation time, while a small room may have a short reverberation time. The preset selector can select a preset acoustic environment 214 or directly select a preset acoustic parameter 216. In some respects, the preset selector does not choose a preset acoustic environment, but instead generates a room model that models the desired acoustic environment based on the analysis of the user's acoustic environment and / or the analysis of the audio or video signals of the content. This desired acoustic environment can be used in place of the preset acoustic environment, or it can be used to inform the user of the selection of a preset acoustic environment.
[0038] The preset selector can determine or select preset acoustic parameters to increase the similarity to the acoustic environment of the audio file if the audio file's acoustic environment is outdoor. On the other hand, the preset selector can select preset parameters to increase the similarity to the user's current acoustic environment if the audio file's acoustic environment is indoor or nonexistent.
[0039] For example, if a user is listening to an audio file recorded outdoors, the creator's intention might be to make the user experience the sound as if the scene were outdoors. In this case, the preset selector can reduce the 'room effect' applied to the audio file. However, if the audio file's scene becomes an indoor setting, the preset selector can increase the 'room effect' to improve the realism of the spatial rendering. This can be done with less consideration for the intended audio scene, as the indoor scene of the audio file can be perceptibly similar to the user's acoustic environment (e.g., a living room). As shown in the line graph, the preset selector can select preset acoustic parameters for a more realistic indoor movie scene 206 (further to the left) and a more fidelity-oriented outdoor movie scene 208 (further to the right). The 'room effect' can be understood as the artificial application of acoustic parameters to a virtual room, where the virtual room can resemble the user's actual acoustic environment.
[0040] The preset selector can select a preset acoustic parameter 216 in response to an indication 210 in the metadata of the audio file to increase the similarity to the acoustic environment of the audio file. For example, the metadata may include a control such as a value that completely disables the "room effect" at one end, allowing the audio to be spatially rendered with the recorded acoustic environment without added reverberation or other artificially added acoustic environment. At the other end of this value, a preset acoustic environment can be selected to increase the similarity to the user's current environment. In some aspects, the control may be a binary value that also completely disables the "room effect," allowing the recorded acoustic environment to be spatialized without any additional acoustic environment applied to the audio file.
[0041] In some respects, acoustic parameters can be specified in the metadata of an audio file. If acoustic parameters exist in the metadata, the preset selector can apply these metadata acoustic parameters as preset acoustic parameters 216 to the audio file. If acoustic parameters and controls exist in the metadata, the controls can indicate conditions such as applying these acoustic parameters to a specific scenario at a specific time.
[0042] In some respects, the preset selector can choose preset parameters to increase the similarity to the acoustic environment of the audio file in response to its association with the visual work. On the other hand, the preset selector can choose preset parameters to increase the similarity to the user's current acoustic environment in response to an audio file not being associated with the visual work.
[0043] For example, as the bar chart shows, if the audio file is an audiovisual file such as a movie, this can cause the preset selector to bias towards fidelity. The preset selector can choose preset parameters where the "room effect" for the indoor movie scene 206 is less than that for the podcast 204 or the virtual assistant 202. Although podcasts can be recorded in an acoustic environment, this is usually an artifact of the recording process, not the creator's intended acoustic effect. Therefore, the preset selector can choose preset parameters where fidelity is emphasized more, with minimal impact on the podcast experience. On the other hand, virtual assistants can include artificially generated voices without an acoustic environment. Therefore, the preset selector can choose preset acoustic parameters where fidelity is fully emphasized. In some aspects, one or more acoustic parameters of the user's current acoustic environment can be determined, and those acoustic parameters can be applied to produce a "room effect."
[0044] The preset selector can apply or adjust weights or other control parameters to bias the selection of preset acoustic parameters toward an acoustic environment similar to that of an audio file or a user's acoustic environment. For example, increasing or decreasing the control parameters can bias the selection of preset acoustic parameter 216 or preset acoustic environment 214 toward realism. Increasing or decreasing the control parameters can bias the selection toward fidelity. The control parameters can be applied linearly or non-linearly to bias the selection as needed.
[0045] Figure 3 The diagram illustrates a workflow for determining and applying preset acoustic parameters based on several aspects. One or more microphones 316 can generate corresponding microphone signals. One or more microphones can be integrated within a computing device. In some aspects, one or more microphones can be integrated into a common device having a speaker 314. The speaker 314 can be a headset speaker or one or more speakers.
[0046] The acoustic parameter estimator 302 can determine one or more acoustic parameters of the user's current acoustic environment based on microphone signals captured by one or more microphones. Acoustic parameters may include one or more of the following: reverberation time (e.g., T60, T30, etc.), direct-to-reverberation ratio (DRR), reflection density, sense of immersion, speech intelligibility, or other acoustic parameters.
[0047] In some respects, acoustic parameter estimators can apply machine learning models (e.g., neural networks or other machine learning algorithms) to microphone signals to determine the acoustic parameters of a user's current acoustic environment. Neural networks or other machine learning models can be trained using existing datasets, enabling the models to extract acoustic parameters with minimal error.
[0048] Additionally or alternatively, the acoustic parameter estimator 302 may use digital signal processing algorithms, such as blind-room estimation algorithms, beamforming, or frequency-domain adaptive filters (FDAF), to determine one or more acoustic parameters of the user's current acoustic environment. Blind-room estimation can be understood as using a recording of a reverberant signal to estimate the acoustic parameters of a space, such as without using the original emitted signal and without generating artificial test stimuli to analyze the space's response. Therefore, a blind-room estimation algorithm can be applied to the microphone signal to determine the acoustic parameters of the space in which the microphone is located, which can be assumed to be the user's space. Beamforming may include applying phase shifts to the microphone or audio signal to create constructive and deconstructive interference, thereby emphasizing acoustic pickup in some directions and de-emphasizing it in others. Frequency-domain adaptive filters may include filtering of the microphone signal, error estimation, and tap weight self-adjustment based on error estimation. Other digital signal processing algorithms may be used to estimate the acoustic parameters of the user's environment.
[0049] Similarly, acoustic parameter estimator 308 can apply digital signal processing algorithms or machine learning models (as described relative to box 302) to one or more audio signals of audio file 312 to determine content-based acoustic parameters. It should be understood that, for the purposes of this disclosure, audio files are interchangeable with audiovisual files. Content-based acoustic parameters can be understood as the acoustic parameters of the acoustic environment in which the audio signals are recorded. In some respects, the acoustic environment of the content can be artificially altered by the creator, for example, in post-production. In any case, the audio signals can carry acoustic parameters that serve as a perceptible cue to the acoustic environment of the content. For example, if the scene in the audio file is a concert hall, there may be a long reverberation time and strong acoustic energy in many directions. In this case, estimator 308 can determine RT60 (which will be relatively long), DRR (which will be relatively low), surround sound (which will be relatively high), or other content-based acoustic parameters from the audio signals.
[0050] In some respects, metadata can include content-based acoustic parameters of the scene. The acoustic environment classifier 310 can scan the metadata, or the preset selector can directly select these content-based acoustic parameters and use them as preset acoustic parameters to be applied to the audio signal.
[0051] The acoustic environment classifier 310 can classify the environment of an audio file based on acoustic parameters or metadata. The acoustic environment of an audio file can be classified by room volume (e.g., large room, medium room, small room), whether it is an open space (e.g., outdoors), or an enclosed space (e.g., indoors). Different levels of granularity can be used to classify the acoustic environment. In some respects, the environment can be classified based on spatial type, such as, for example, a room, a library, a cathedral, a stadium, a forest, an open field, a hillside, a valley, etc.
[0052] Metadata can indicate whether a scene is "outdoor," "indoor," a large room, a medium-sized room, a small room, a reverberant room, a concert hall, a library, or another acoustic environment. An acoustic environment classifier can use the environment indicated in the metadata to classify an acoustic environment. If the environment is not present in the metadata, the classifier can determine the acoustic environment based on content-based acoustic parameters. For example, if RT60 is a quantity "x" and DRR is a quantity "y," the acoustic environment can be classified as a concert hall. If RT60 is "a" and / or DRR is "b," the acoustic environment can be classified as outdoor.
[0053] Preset selector 304 may determine one or more preset acoustic parameters based on: one or more acoustic parameters of the user's current acoustic environment, determined at box 302; and / or the acoustic environment of an audio file, as categorized at box 310, determined based on the audio signal or metadata of the audio file. In some aspects, the preset selector may use rule-based algorithms that determine the preset acoustic parameters (or select preset acoustic environments that include the preset acoustic parameters) based on content type, scene type, and the user's acoustic environment. For example, the preset selector may enforce a rule stating that if content type = "movie", scene = "outdoor", and the acoustic parameter of the user's acoustic environment = "reverberation", then the preset acoustic parameter is set to "low reverberation". In some aspects, the preset selector may generate a room model to form a desired virtual acoustic environment. The room model may include parameters, algorithms, and / or mathematical relationships that define acoustic behavior, such as, for example, reverberation time, impulse response, or acoustic parameters. The estimation results of the user's acoustic environment (from box 302) and the estimation results of the content (from box 308) can include parameters of the room model, such as, for example, room size and / or the absorptivity of the simulated surfaces. Reverberation time and / or other acoustic parameters can be derived from the relationship between room size (e.g., volume), absorptivity, and reverberation time. For example, T = .16V / A, where T represents the reverberation time, V represents the room volume, and A represents the total absorptivity of the room. The room model can include other relationships from which acoustic parameters are derived based on control parameters. These acoustic parameters can be used as preset acoustic parameters.
[0054] Additionally or alternatively, the preset selector may use data-driven algorithms, such as trained neural networks or other trained machine learning models. Data-driven algorithms can select one or more preset acoustic parameters from a large pool of data. Machine learning models can be trained to output preset acoustic parameters with minimal error when applied to one or more acoustic parameters related to content type, audio scene type, and / or user environment.
[0055] In this way, the system can classify the acoustic environment of audio files (in box 310) and select preset acoustic parameters (in box 304) as a balance or trade-off between the audio file's acoustic environment and one or more real acoustic parameters. A user's "room effect" can be added in some cases, such as in an indoor scene, strictly for audio content, or when the audio does not have its own acoustic environment. A user's "room effect" can be reduced or turned off in other cases, such as in movies, or when the content creator has specified it in the metadata. The system can take the audio's acoustic environment and the user's acoustic environment from different parameters (e.g., metadata) and determine or select the optimal acoustic scenario.
[0056] In some respects, the acoustic environment classifier 310 can classify a space based on the video signal from the audio file or audiovisual file 312. For example, the classifier may include a computer vision algorithm that can determine whether a scene is an outdoor scene or an indoor scene.
[0057] In some respects, one or more acoustic parameters of a user's current acoustic environment can be stored and reused later. For example, a user can watch a performance (e.g., an audiovisual file), thereby triggering... Figure 3 The workflow is illustrated. The reverberation time and / or DRR of the user's living room can be stored in a computer-readable medium at box 302. The next day, when the user returns to the living room and listens to a podcast, a device such as a smartphone or speaker 314 can sense that the user is in the same acoustic environment (i.e., the living room). The stored reverberation time and / or DRR can be reused to spatialize the podcast so that they do not need to be recalculated.
[0058] Spatial renderer 306 can apply spatial filters to one or more audio signals 318. The spatial renderer can convolve the spatial filters on the one or more audio signals to produce a resulting spatialized audio channel. The resulting spatialized audio channel can be used to drive a speaker 314. The speaker 314 can include a left and a right speaker worn in or on a user's ear. In some aspects, the speaker 314 can include one or more speaker arrays that can be integrated with one or more speaker enclosures. The spatial renderer can select spatial filters based on preset acoustic parameters, such that the spatial filters include the desired effects of the preset acoustic parameters, such as, for example, desired reverberation time, DRR, surround effect, reflection density, and / or speech intelligibility.
[0059] It should be understood that although the processing boxes shown are grouped into individual boxes to illustrate the workflow, each of them can be executed by an audio processing system or distributed among multiple audio processing systems that can communicate via a network. Some boxes or all boxes can be combined into one or more other boxes.
[0060] Figure 4 An audio processing method 400 according to some aspects is illustrated. Method 400 can be performed using the aspects described. The method can be performed by a device, hardware (e.g., circuitry, dedicated logic components, programmable logic components, processors, processing devices, central processing units (CPUs), system-on-a-chip (SoCs), etc.), software (e.g., instructions running / executing on the processing device), firmware (e.g., microcode), or a combination thereof. Although specific functional blocks (“blocks”) are described in the method, such blocks are examples. That is, the aspects are well-suited for performing the various other blocks or variations of the blocks described in the method. It should be understood that the blocks in the method may be performed in a different order than presented, and not all blocks in the method can be performed.
[0061] At box 402, the processor can determine one or more acoustic parameters of the user's current acoustic environment based on sensor signals captured by one or more sensors of the device. For example, the processor can apply digital signal processing algorithms or machine learning algorithms to microphone signals captured by a microphone and / or camera images captured by a camera, as described in other sections.
[0062] At box 404, the processor can determine one or more preset acoustic parameters based on one or more acoustic parameters of the user’s current acoustic environment and the acoustic environment of an audio file including an audio signal, the acoustic environment of which is determined based on the audio signal of the audio file or the metadata of the audio file.
[0063] For example, a processor can determine content-based acoustic parameters from the audio signal of an audio file. The processor can categorize the environment of the audio file based on metadata or content-based acoustic parameters. A preset selector can select preset acoustic parameters based on the categorization of the audio file's acoustic environment and the acoustic parameters of the user's acoustic environment. Other aspects are also described.
[0064] At box 406, the processor can spatially render the audio signal by applying one or more preset acoustic parameters, thereby generating a spatialized audio signal. At box 408, the processor can use the spatial audio signal to drive a speaker.
[0065] Figure 5 A method for determining preset acoustic parameters based on several aspects is illustrated. Method 500 can be performed using the aspects described. The method can be performed by hardware (e.g., circuitry, dedicated logic components, programmable logic components, processors, processing devices, central processing units (CPUs), system-on-a-chip (SoCs), etc.), software (e.g., instructions running / executing on the processing device), firmware (e.g., microcode), or combinations thereof. Although specific functional blocks (“blocks”) are described in the method, such blocks are examples. That is, the aspects are well-suited for performing various other blocks or variations of the blocks described in the method. It should be understood that the blocks in the method can be performed in a different order than presented, and not all blocks in the method can be performed.
[0066] In box 502, the processor can determine whether the audio or audiovisual file includes metadata containing acoustic environment or acoustic parameters for playback. For example, the processor can scan the metadata to determine if an acoustic environment or acoustic parameters exist.
[0067] In box 504, in response to an audio file or audiovisual file including metadata containing acoustic environment or acoustic parameters, the processor can spatially render the audio signal associated with the audio file or audiovisual file based on the acoustic environment or acoustic parameters of the metadata.
[0068] In box 506, in response to the fact that the audio or audio file does not include metadata, the processor may render the audio signal spatially based on one or more acoustic parameters of the user’s current environment, the current scene of the audio or audio file, or the content type of the audio or audio file.
[0069] For example, if the current scene of an audio or audio file is an outdoor scene, the audio signal can be rendered spatially if its similarity to the user's current environment is relatively small. On the other hand, if the current scene of an audio or audio file is an indoor scene, the audio signal can be rendered spatially if its similarity to the user's current environment is relatively large.
[0070] In response to the content type of an audio or audiovisual file being a movie, the audio signal can be rendered spatially where its similarity to the user's current environment is relatively small, and its similarity to acoustic parameters extracted from the audio or video signal of the audio or audiovisual file is relatively large. On the other hand, in response to the content type of an audio or audiovisual file being a podcast or talk show, the audio signal can be rendered spatially where its similarity to the user's current environment is relatively large.
[0071] By increasing or decreasing weights or other control parameters and / or by selecting preset acoustic parameters based on a specific order, the rendering of an audio signal can be biased towards being more similar to the current environment (e.g., an increased "room effect") and less similar to the current environment (e.g., a smaller "room effect"). For example, preset acoustic parameters can be arranged in a specific order such as... Figure 2 The sliding scale shown sorts and groups the data from realism to fidelity.
[0072] Figure 6 Audio processing operations are illustrated according to several aspects. These operations can be performed by an audio processing system, with each aspect described in other sections.
[0073] At box 602, the audio processing system can read the metadata of the audio or audiovisual file and determine whether there are controls specifying whether the original acoustic environment should be preserved. These controls may specify that no additional reverberation or other "room effects" will be added. In some cases, the controls may define acoustic parameters to be applied during spatialization. Therefore, at box 602, if the metadata includes controls or other indications, the audio processing system can proceed to box 612, where it can use the acoustic environment or acoustic parameters specified in the metadata or inherent in the audio signal of the audio file.
[0074] If the metadata does not provide any such indication, the audio processing system may proceed to box 604. If the audio file is merely audio without visual components, such as podcasts, talk shows, music, or virtual assistants, the audio processing system may proceed to box 610 and amplify the effect of the user's acoustic environment. The audio processing system may proceed to box 608 and determine whether the audio file has an indoor or outdoor setting. This may be performed based on techniques such as, for example, digital signal processing, machine learning-based techniques, or metadata, as described in other sections. If the audio file is determined to have an outdoor setting, the audio processing system may proceed to box 614 and reduce the effect of the user's acoustic environment. If the audio file has an indoor setting, the audio processing system may revisit box 610 and further amplify the effect of the user's acoustic environment. Thus, a plain audio file that may be intended to emit sound outdoors may have a smaller applied room effect, while those audio files recorded indoors may have a larger applied room effect.
[0075] If the audio file is not purely audio, the audio processing system can proceed to box 614 and reduce the impact of the user's acoustic environment. The audio processing system can then proceed to box 606. If the audiovisual file is a movie, the audio processing system can revisit box 614 to further reduce the impact of the user's acoustic environment. If the audiovisual file is not a movie, the audio processing system can proceed to box 610 and increase the impact of the user's acoustic environment.
[0076] The audio processing system can proceed to box 608. If the audiovisual scene is an outdoor scene, the audio processing system can revisit box 614 and further reduce the impact of the user's acoustic environment. Otherwise, if it is an indoor movie scene, the audio processing system can proceed to box 610 to increase the impact of the user's acoustic environment. As discussed, one or more weights or other control parameters can be adjusted to increase or decrease the impact of the user's acoustic environment.
[0077] At box 608, considering the extent to which the user's acoustic environment should influence the determination of preset acoustic parameters (e.g., as a result of operation), the audio processing system can use preset acoustic parameters determined based on metadata or the user's acoustic environment to spatialize the audio.
[0078] Figure 7 Metadata 704 is shown according to some aspects. Metadata 704 may be integrated with or associated with audio or audiovisual files. As mentioned, audio or audiovisual files may include static files or streaming data.
[0079] Metadata may include timestamps 716 describing the start and end of a scene. Each scene may have its own set of fields describing the acoustic environment of the scene. Fade-in / fade-out areas 78 may be specified for transitions between scenes. Additionally, metadata may include virtualization indicator 724, which may include controls indicating whether room effects are applied to the audio. Such indicator may be a binary value or other value that provides a sliding scale showing how much the user's environment may influence the current audio or audiovisual work.
[0080] Metadata 704 may include various fields that can indicate the acoustic environment or acoustic parameters. For example, metadata may have a field 710 indicating whether the scene is indoors or outdoors. Metadata may specify the room size 706, room geometry 708, and / or surface absorptivity 712 of various surfaces in the acoustic environment of the scene. Metadata may include air density and / or humidity 714 within the acoustic environment. Metadata may include acoustic parameters 726, such as reverberation time, DRR, reflection density, immersion, speech intelligibility, or other acoustic parameters.
[0081] In some aspects, the audio processing system 720 can create metadata and embed it into or associate it with audio or audiovisual files. The audio processing system can obtain audio or audiovisual files from a source 702, which can be a capture device (e.g., a microphone and / or camera). In some aspects, the source 702 can be a downstream device, such as a computer used for post-production of audio or audiovisual data.
[0082] In some aspects, the audio processing system can create metadata in real time or as the scene is captured by the capture device. In some aspects, the audio processing system 720 can be integrated with the capture device. The audio processing system may include sensors 722, such as, for example, one or more microphones, barometers, and / or cameras. The audio processing system applies digital signal processing algorithms and / or machine learning models to the audio signal or video of the audio file or to the sensor data to determine metadata fields such as 706, 708, 710, 712, 714, and 726. The user can set the virtualization indicator 722, or the audio processing system can apply rule-based algorithms or machine learning-based algorithms to set the virtualization indicator fields, as described in other sections.
[0083] Thus, the audio processing system can generate metadata 704 for use by downstream devices (e.g., audio processing systems described in other sections) to determine when and which acoustic parameters to apply to an audio file. The metadata can explicitly indicate the desired acoustic environment, the content-based acoustic environment, and the user's acoustic environment, and / or other acoustic data (e.g., room size, geometry, surface parameters, etc.) that can be inferred from it downstream.
[0084] In some aspects, a method includes: determining whether an audio file or audiovisual file includes metadata containing an acoustic environment for playback; in response to the audio file or audiovisual file including metadata (which contains an acoustic environment for playback), spatially rendering an audio signal associated with the audio file or audiovisual file based on the acoustic environment of the metadata; and in response to the audio file or audiovisual file not including metadata, spatially rendering the audio signal based on one or more acoustic parameters of the user's current environment, the current scene of the audio file or audiovisual file, or based on the content type of the audio file or audiovisual file. In some aspects, in response to the current scene of the audio file or audiovisual file being an outdoor scene, the audio signal is spatially rendered if its similarity to the user's current environment is low. In some aspects, in response to the current scene of the audio file or audiovisual file being an indoor scene, the audio signal is spatially rendered if its similarity to the user's current environment is high. In some respects, in response to the content type of an audio or audiovisual file being a movie, the audio signal is rendered spatially where its similarity to the user's current environment is relatively small, and where its similarity to acoustic parameters extracted from the audio or video signal of the audio or audiovisual file is relatively large. In other respects, in response to the content type of an audio or audiovisual file being a podcast or talk show, the audio signal is rendered spatially with a high degree of similarity to the user's current environment.
[0085] Figure 8 An audio processing system is illustrated according to some aspects. The audio processing system can be a computing device, such as, for example, a desktop computer, tablet computer, smartphone, laptop computer, smart speaker, media player, home appliance, headphone assembly, head-mounted display (HMD), smart glasses, infotainment system for automobiles or other vehicles, or other computing devices. The system can be configured to perform the methods and processes described in this disclosure.
[0086] While various components of an audio processing system that can be incorporated into headphones, speaker systems, microphone arrays, and entertainment systems are shown, this illustration is merely one example of a specific implementation of the types of components that can exist in an audio processing system. This example is not intended to represent any particular architecture or manner in which these components are interconnected, as such details are not closely related to the aspects described herein. It should also be understood that other types of audio processing systems with fewer or more components than those shown can also be used. Therefore, the processes described herein are not limited to use with the hardware and software shown.
[0087] The audio processing system may include one or more buses 818 for interconnecting various components of the system. As is known in the art, one or more processors 804 are coupled to the buses. The one or more processors may be microprocessors or dedicated processors, system-on-a-chip (SOC), central processing units, graphics processing units, processors created using application-specific integrated circuits (ASICs), or combinations thereof. Memory 810 may include read-only memory (ROM), volatile memory, and non-volatile memory, or combinations thereof, coupled to the buses using techniques known in the art. Sensor 816 may include an IMU and / or one or more cameras (e.g., RGB cameras, RGBD cameras, depth cameras, etc.) or other sensors described herein. The audio processing system may also include a display 814 (e.g., an HMD or touchscreen display).
[0088] The memory 810 may be connected to a bus and may include DRAM, hard disk drive or flash memory, or magnetic optical drive or magnetic storage, or optical drive or other types of memory systems that maintain data even after the system is powered off. In one aspect, the processor 804 retrieves computer program instructions stored in a machine-readable storage medium (memory) and executes those instructions to perform the operations described herein.
[0089] Although not shown, the audio hardware may be coupled to one or more buses to receive audio signals to be processed and output by the speaker 808. The audio hardware may include digital-to-analog converters and / or analog-to-digital converters. The audio hardware may also include audio amplifiers and filters. The audio hardware may also be connected to a microphone 806 (e.g., a microphone array) to receive audio signals (whether analog or digital), digitize them as appropriate, and transmit the signals to the buses.
[0090] The communication module 812 can communicate with remote devices and networks via wired or wireless interfaces. For example, the communication module can communicate using known technologies such as TCP / IP, Ethernet, Wi-Fi, 3G, 4G, 5G, Bluetooth, ZigBee, or other equivalent technologies. The communication module may include wired or wireless transmitters and receivers capable of communicating (e.g., receiving and sending data) with networked devices such as servers (e.g., the cloud) and / or other devices such as remote speakers and remote microphones.
[0091] It should be understood that the aspects disclosed herein can utilize memory located remotely to the system, such as network storage devices coupled to the audio processing system via network interfaces such as modems or Ethernet interfaces. As is well known in the art, buses can be interconnected via various bridges, controllers, and / or adapters. In one aspect, one or more network devices can be coupled to the bus. Network devices can be wired network devices (e.g., Ethernet) or wireless network devices (e.g., Wi-Fi, Bluetooth). In some aspects, the aforementioned aspects (e.g., simulation, analysis, estimation, modeling, object detection, etc.) can be performed by a networked server communicating with the capture device.
[0092] The various aspects described herein can be embodied, at least in part, in software. That is, these techniques can be implemented in an audio processing system in response to its processor executing a sequence of instructions contained in a storage medium, such as a non-transitory machine-readable storage medium (such as DRAM or flash memory). In each aspect, hard-wired circuitry can be used in combination with software instructions to implement the techniques described herein. Therefore, these techniques are not limited to any specific combination of hardware circuitry and software, nor to any particular source of instructions executed by the audio processing system.
[0093] In this specification, certain terms are used to describe the characteristics of various aspects. For example, in some cases, the terms “module,” “processor,” “unit,” “renderer,” “system,” “device,” “filter,” “reverberator,” “estimator,” “classifier,” “box,” “selector,” “simulation,” “model,” and “component” refer to hardware and / or software configured to perform one or more processes or functions. For example, examples of “hardware” include, but are not limited to, integrated circuits such as processors (e.g., digital signal processors, microprocessors, application-specific integrated circuits, microcontrollers, etc.). Thus, as those skilled in the art will understand, different combinations of hardware and / or software can be implemented to perform the processes or functions described by the foregoing terms. Of course, hardware may alternatively be implemented as a finite state machine or even a combinational logic component. Examples of “software” include application programs, applets, routines, and even executable code in the form of a series of instructions. As mentioned above, software can be stored on any type of machine-readable medium.
[0094] Certain portions of the foregoing detailed description have been presented in accordance with algorithms and symbolic representations for manipulating data bits in computer memory. These algorithmic descriptions and representations are methods used by those skilled in the art of audio processing, and these methods are also the most effective way to communicate the substance of their work to others skilled in the art. An algorithm herein and generally refers to a self-consistent sequence of operations that leads to a desired result. These operations are those that require physical manipulation of physical quantities. However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specifically stated, it is apparent from the foregoing discussion that throughout the specification, the use of terms such as those given in the claims below refers to the actions and processes of an audio processing system or similar electronic device that manipulate data represented as physical (electronic) quantities in the system's registers and memories, and convert them into other data similarly represented as physical quantities in the system's memory or registers or other such information storage, transmission, or display devices.
[0095] The processes and blocks described herein are not limited to the specific examples described, nor are they limited to the specific order in which they are used as examples herein. Rather, any processing blocks can be reordered, combined, or removed, executed in parallel or serially, as needed to achieve the above results. Processing blocks associated with implementing an audio processing system can be executed by one or more programmable processors executing one or more computer programs stored on a non-transitory computer-readable storage medium to perform the functions of the system. All or part of the audio processing system can be implemented as special-purpose logic circuitry (e.g., FPGA (Field-Programmable Gate Array) and / or ASIC (Application-Specific Integrated Circuit)). All or part of the audio system can be implemented using electronic hardware circuitry including at least one of electronic devices such as, for example, processors, memories, programmable logic devices, or logic gates. Additionally, processes can be implemented in any combination of hardware devices and software components.
[0096] In some aspects, this disclosure may include the language "[element A] and [element B] at least one". This language may refer to one or more of these elements. For example, "at least one of A and B" may refer to "A", "B", or "A and B". Specifically, "at least one of A and B" may refer to "at least one of A and at least one of B" or "at least either A or B". In some aspects, this disclosure may include the language "[element A], [element B], and / or [element C]". This language may refer to any of these elements or any combination thereof. For example, "A, B, and / or C" may refer to "A", "B", "C", "A and B", "A and C", "B and C", or "A, B, and C".
[0097] While certain aspects have been described and shown in the accompanying drawings, it should be understood that such aspects are illustrative only and not restrictive, and this disclosure is not limited to the specific constructions and arrangements shown and described, as various other modifications will be apparent to those skilled in the art.
[0098] In order to assist the Patent Office and any reader of any patent published in this application in interpreting the appended claims, the applicant wishes to note that they do not intend any of the appended claims or claim elements to invoke 35 USC112(f) unless the words “means for…” or “steps for…” are expressly used in a particular claim.
[0099] As is widely recognized, the use of personally identifiable information should comply with privacy policies and practices that are generally accepted to meet or exceed industry or governmental requirements for protecting user privacy. Specifically, personally identifiable information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use, and the nature of authorized use should be clearly explained to users.
Claims
1. A method executed by a processor of a device, the method comprising: Based on sensor signals captured by one or more sensors of the device, one or more acoustic parameters of the user's current acoustic environment are determined; Based on one or more acoustic parameters of the user's current acoustic environment and the acoustic environment of an audio file including an audio signal, one or more preset acoustic parameters are determined, wherein the acoustic environment of the audio file is determined based on the audio signal of the audio file or the metadata of the audio file; By applying one or more preset acoustic parameters to the audio signal, the audio signal is rendered spatially, thereby generating a binaural audio signal; as well as The speaker is driven by the binaural audio signal.
2. The method of claim 1, wherein determining the one or more preset acoustic parameters comprises: In response to the acoustic environment of the audio file being an outdoor scene, one or more preset acoustic parameters are selected to increase the similarity to the acoustic environment of the audio file.
3. The method of claim 1, wherein determining the one or more preset acoustic parameters comprises: In response to an indication in the metadata, one or more preset acoustic parameters are selected to increase the similarity to the acoustic environment of the audio file.
4. The method of claim 1, wherein determining the one or more preset acoustic parameters comprises: In response to the acoustic parameter being present in the metadata or being indicated by a control in the metadata, an acoustic parameter specified in the metadata is selected as one or more preset acoustic parameters.
5. The method of claim 1, wherein determining the one or more preset acoustic parameters comprises: In response to whether the acoustic environment of the audio file is indoors or nonexistent, one or more preset acoustic parameters are selected to increase the similarity to the user's current acoustic environment.
6. The method of claim 1, wherein determining the one or more preset acoustic parameters comprises: In response to the association of the audio file with a visual work, one or more preset acoustic parameters are selected to increase the similarity to the acoustic environment of the audio file.
7. The method of claim 1, wherein determining the one or more preset acoustic parameters comprises: In response to the audio file not being associated with a visual work, one or more preset acoustic parameters are selected to increase the similarity to the user's current acoustic environment.
8. The method of claim 1, wherein determining the one or more preset acoustic parameters comprises: The acoustic environment of the audio file is classified, and one or more preset acoustic parameters are selected as the balance between the acoustic environment of the audio file and the acoustic parameters of the user's current acoustic environment.
9. The method of claim 1, wherein determining the acoustic environment of the audio file comprises: Extract content-based acoustic parameters from the audio signal of the audio file or from the metadata.
10. The method of claim 1, wherein the acoustic environment of the audio file is classified according to at least one of the following: room volume, in an open space, or in an enclosed space.
11. A system for audio processing, the system comprising: A microphone that generates microphone signals characterizing the acoustic environment of the system; as well as A non-transitory computer-readable storage device and a processor, the non-transitory computer-readable storage device storing executable instructions, the processor being configured to execute the instructions to cause the system to: Based on the microphone signal, one or more acoustic parameters of the acoustic environment of the system are determined, the one or more acoustic parameters including at least the reverberation time; One or more preset acoustic parameters are determined based on one or more acoustic parameters of the acoustic environment of the system and the acoustic environment of the audio file, wherein the acoustic environment of the audio file is determined based on the audio signal of the audio file or the metadata of the audio file; Rendering the audio signal in space includes: applying one or more preset acoustic parameters to the audio signal to generate a spatialized audio signal; as well as The speaker is driven by the spatialized audio signal.
12. The system of claim 11, wherein the system includes a headset on which the microphone and the speaker are integrated.
13. The system of claim 11, wherein the acoustic environment of the audio file is classified according to spatial type, including: Rooms, libraries, cathedrals, stadiums.
14. The system of claim 11, wherein the acoustic environment for determining the audio file is further based on a video signal associated with the audio file.
15. The system of claim 11, wherein determining the one or more acoustic parameters of the acoustic environment of the system or determining the acoustic parameters based on the audio signal of the audio file is performed using a machine learning model.
16. The system of claim 11, wherein determining the one or more acoustic parameters of the acoustic environment of the system or determining the acoustic parameters based on the audio signal of the audio file is performed using a digital signal processing algorithm including at least one of blind room estimation algorithm, beamforming, or frequency domain adaptive filter (FDAF).
17. The system of claim 11, further comprising: The system stores one or more acoustic parameters of the acoustic environment of the system, and in response to sensing the acoustic environment of the system at a later time, reuses the stored one or more acoustic parameters of the acoustic environment at the later time.
18. The system of claim 11, wherein determining the one or more preset acoustic parameters is performed using a rule-based algorithm, the rule-based algorithm being based on content type, audio scene type, and the one or more acoustic parameters of the acoustic environment of the system.
19. The system of claim 11, wherein determining the one or more preset acoustic parameters is performed by applying a machine learning model to the one or more acoustic parameters of the content type, audio scene type, and / or the acoustic environment of the system.
20. The system of claim 11, wherein the one or more acoustic parameters and the one or more preset acoustic parameters include at least one of reverberation time, direct reverberation ratio (DRR), reflection density, immersion, or speech intelligibility.