Spatial audio capture with depth
By combining depth sensors and audio sensors to capture and encode depth information of the sound field, the problem of difficult to distinguish and render near-field and far-field audio information in the prior art is solved, and higher quality audio signal rendering is achieved.
Patent Information
- Application Number
- CN201980102300.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-10-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2039-10-10
AI Technical Summary
The prior art is difficult to effectively capture and encode the depth quantization of sound field information, resulting in difficulty in distinguishing and rendering of near-field and far-field audio information.
By combining the depth sensor with the audio sensor, the depth information and the audio information are combined to generate a spatial audio signal with depth characteristics. The system may include a microphone array or acoustic field microphone, a depth camera or a depth sensor for capturing and encoding depth information of the sound field.
Accurate depth encoding of sound field information is realized, allowing the decoder to distinguish between near-field and far-field audio information, thereby improving the rendering quality and authenticity of the audio signal.
Smart Images

Figure CN114902330B_ABST
Abstract
Description
Background Art
[0001] Audio and video capture systems, such as those that may include or use a microphone and a camera, respectively, may be co-located in an environment and configured to capture audiovisual information from the environment. The captured audiovisual information may be recorded, transmitted, and played back on demand. In an example, the audiovisual information may be captured in an immersive format, such as using a spatial audio format and a multi-dimensional video or image format.
[0002] In an example, the audio capture system may include a microphone, microphone array, or other sensor including one or more transducers to receive audio information from the environment. The audio capture system may include or use a spatial audio microphone configured to capture a three-dimensional or 360 degree sound field, such as an ambisonic microphone.
[0003] In an example, the video capture system may include a single lens camera or a multi-lens camera system. In an example, the video capture system may be configured to receive 360-degree video information, which is sometimes referred to as immersive video or spherical video. In a 360-degree video, image information from multiple directions may be received and recorded simultaneously. In an example, the video capture system may include or comprise a depth sensor configured to detect depth information of one or more objects in the field of view of the system.
[0004] Various audio recording formats can be used to encode three-dimensional audio cues in a recording. Three-dimensional audio formats include ambisonic and discrete multi-channel audio formats that include height speaker channels. In an example, a downmix can be included in the soundtrack component of a multi-channel digital audio signal. The downmix can be backward compatible and can be decoded by a traditional decoder and reproduced on an existing or traditional playback device. The downmix can include a data stream extension with one or more audio channels that can be ignored by a traditional decoder but can be used by a non-traditional decoder. For example, a non-traditional decoder can recover additional audio channels, subtract their contributions in a backward compatible downmix, and then render them in a target spatial audio format.
[0005] In an example, a target spatial audio format for a soundtrack may be specified at the encoding or production stage. The method allows a multi-channel audio soundtrack to be encoded in the form of a data stream that is compatible with a conventional surround sound decoder and one or more alternative target spatial audio formats also selected at the encoding or production stage. These alternative target formats may include formats suitable for improved reproduction of three-dimensional audio cues. However, one limitation of this approach is that encoding the same soundtrack for another target spatial audio format may require returning to the production facility to record and encode a new version of the soundtrack mixed for the new format.
[0006] Object-based audio scene coding provides a general solution for soundtrack coding that is independent of the target spatial audio format. An example of an object-based audio scene coding system is the MPEG-4 Advanced Audio Binary Format (AABIFS) for scenes. In this method, each source signal is sent separately together with a rendering hint data stream. The data stream carries the time-varying values of the parameters of the spatial audio scene rendering system. The parameter set can be provided in the form of an audio scene description that is independent of the format, so that the soundtrack can be rendered in any target spatial audio format by designing a rendering system according to the format. Each source signal can be combined with its associated rendering hint to define an "audio object". The method enables the renderer to implement accurate spatial audio synthesis technology, thereby rendering each audio object in any target spatial audio format selected at the reproduction end. The object-based audio scene coding system also allows interactive modification of the rendered audio scene in the decoding stage, including remixing, music reinterpretation (e.g., karaoke) or virtual navigation in the scene (e.g., video games). Summary of the invention
[0007] The inventors have recognized that problems to be solved include capturing sound field information into a depth-quantized spatial audio format. For example, the inventors have recognized that a spatial audio signal can include a far-field or omnidirectional component, a near-field component, and information from a mid-field by interpolating or mixing signals from different depths. For example, an auditory event to be simulated in a spatial region between a specified near field and far field can be created by a crossfade between the two depths.
[0008] This problem may include, for example, audio scene information captured using soundfield microphones but without depth information. Such captured audio scene information is often quantized into a general or non-specific "soundfield" and then rendered or encoded as far-field information. A decoder receiving such information may not be configured to distinguish between near-field and far-field sources and may not utilize or use near-field rendering. For example, some information captured using soundfield microphones may include near-field information. However, if depth information is not encoded with the audio scene information, the near-field information may be classified as far-field or other reference soundfield or default depth.
[0009] A solution to the sound field capture or audio capture problem may include using a depth sensor to receive auditory information and visual information about the environment substantially simultaneously with an audio sensor. The depth sensor may include a three-dimensional depth camera, or a two-dimensional image sensor, or multiple sensors with processing capabilities, or the like. The depth sensor may render or provide information about one or more objects in the environment. The audio sensor may include one or more microphone elements that may sense auditory information from the environment. In an example, the solution includes a system or encoder configured to combine information from a depth sensor and an audio sensor to provide a spatial audio signal. The spatial audio signal may include one or more audio objects, and the audio objects may have corresponding depth characteristics.
[0010] This summary is intended to provide an overview of the subject matter of this patent application. It is not intended to provide an exclusive or exhaustive explanation of the invention. Specific embodiments are included to provide further information about this patent application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To facilitate understanding of the discussion of any particular element or act, the most significant digit(s) in a reference number refers to the figure number in which the element is first introduced.
[0012] Figures 1A-1C Schematic diagram generally showing the location of audio sources or objects relative to a listener.
[0013] Figure 2A An example of a system configured to receive audio information and visual information about an environment is generally shown.
[0014] Figure 2B Examples of object recognition and depth analysis for an environment are generally shown.
[0015] Figure 3 Examples showing how information from the environment can be quantized to different depths are generally shown.
[0016] Figure 4 An example of a block diagram of a system for spatial audio capture and encoding is generally shown.
[0017] Figure 5 An example of a first method that may include encoding a spatial audio signal is generally shown.
[0018] Figure 6 An example of a second method that may include encoding a spatial audio signal based on correlation information is generally shown.
[0019] Figure 7An example of a third method is generally shown that may include providing an indication of confidence that audio scene information corresponds to a specified object.
[0020] Figure 8 An example of a fourth method is generally shown which may include determining a correspondence between audio signal characteristics and received information about an audio scene.
[0021] Fig. 9 A schematic diagram is generally shown of a machine in the form of a computer system within which a set of instructions may be executed to cause the machine to perform any one or more of the methodologies discussed herein. DETAILED DESCRIPTION
[0022] In the following description of examples of systems, methods, devices, and apparatus for performing, for example, spatial audio signal processing for coordinating audiovisual program information, reference is made to the accompanying drawings that form a part of the detailed description. By way of illustration, the accompanying drawings show specific embodiments in which the invention disclosed herein may be implemented. These embodiments are generally referred to herein as "examples." Such examples may also include elements in addition to those shown or described. However, the inventors also contemplate examples that provide only those elements shown or described. The inventors contemplate examples using any combination or arrangement of those elements shown or described (or one or more aspects thereof), whether with respect to a particular example (or one or more aspects thereof), or with respect to other examples shown or described herein (or one or more aspects thereof).
[0023] The present subject matter relates to processing audio signals (i.e., signals representing physical sounds). These audio signals are typically represented by digital electronic signals. As used herein, the phrase "audio signal" may include signals representing physical sounds. The audio processing systems and methods described herein may include hardware circuits and / or software configured to process audio signals using audio signals or using various filters. In some examples, the systems and methods may use signals from or corresponding to multiple audio channels. In an example, the audio signal may include a digital signal that includes information corresponding to multiple audio channels. Some examples of the present subject matter may operate in the context of a time sequence of digital bytes or words, where these bytes or words form a discrete approximation of an analog signal or a final physical sound. A discrete digital signal corresponds to a digital representation of a periodically sampled audio waveform.
[0024] The present systems and methods may include an environment capture system. The environment capture system may include optical, visual or auditory sensors, such as one or more cameras, depth sensors, microphones, or other sensors configured to monitor the environment. The systems and methods may be configured to receive audio information from the environment and receive distance or position information about physical objects in the environment. These systems and methods may be configured to identify correlations between audio information or its components and physical objects in the environment. When correlations are identified between audio objects and physical objects, spatial audio signals including audio sources of audio objects may be encoded, such as where the audio sources are located at a virtual distance or position relative to a reference position and correspond to one or more physical objects.
[0025] In an example, the audio information or audio signal received from the microphone may include information from the sound field. The received audio information may be encoded substantially in real time with the depth information. For example, information from a depth sensor (e.g., a three-dimensional depth camera) may be used with the audio information, and the audio information may be encoded in a spatial audio format with depth characteristics, such as including direction or depth magnitude information.
[0026] In an example, a system for performing spatial audio capture with depth may include a microphone array or a sound field microphone configured to capture a sound field or sound scene. The system may include a depth camera or depth sensor configured to determine or estimate the depth of one or more objects in the sensor's field of view, and may optionally be configured to receive depth information from multiple directions (e.g., up / down, left / right, etc.). In an example, the system may utilize the depth or distance information received from the depth sensor to enhance the captured auditory information, and then encode the auditory information and depth information in a spatial audio signal. The spatial audio signal may include components or sources having corresponding depths or distances relative to an origin or reference position.
[0027] In an example, the information from the depth sensor includes information about the direction from a reference position or from a reference direction to one or more physical objects or potential audio sources. The directional information about the physical objects can be related to the audio objects. In an example, the encoded spatial audio information described herein can use head-related transfer functions (HRTFs), for example, which can be synthesized or measured at various distances from a reference head (extending from near field to far field). Additional synthesized or measured transfer functions can be used to extend to the interior of the head, for example for distances closer than the near field. In addition, the relative distance-related gain of each HRTF set can be normalized to the far-field HRTF gain.
[0028] Figures 1A-1C Schematic diagram generally showing near-field and far-field of exemplary audio source or object locations. Figure 1AA first diagram 100A is included that shows the position of the audio object 22 relative to a reference position 101. The reference position 101 may be the position of a listener, the position of a microphone, the position of a camera or depth sensor, or other position used as a reference point in the environment represented by the first diagram 100A. Figure 1A and 1B In the example of , radius R1 may represent the distance from the reference location 101 that coincides with the far field, and radius R2 may represent the distance from the reference location 101 that coincides with the near field or its boundary. The environment may be represented using more than two radii, for example, as discussed below. Figure 1C shown.
[0029] Figure 1B Included is shown Figure 1A A second graph 100B is a spherical expansion (eg, using spherical representation 21) of the first graph 100A. Figure 1B In , an audio object 22 may have an associated height characteristic 23, and an associated projection characteristic 25, for example onto a ground plane, an associated elevation characteristic 27, and an associated azimuth characteristic 29. Figure 1A and Figure 1B In the example of , any suitable number of HRTFs may be sampled over a full 3D sphere of radius Rn, and the samples in each common radius HRTF set need not be identical. Figure 1C A third Figure 10C is included that shows a sound field divided or quantized into an arbitrary number of depths. For example, the object 22 may be located at a far field position, a near field position, somewhere in between, inside the near field, or outside the far field.
[0030] exist Figures 1A-1C In the examples, various HRTFs (Hxy) are shown at locations of radii R1 and R2 centered at a reference position 101, where x represents the ring number or radius and y represents the position on the ring. Such position-related HRTFs may be referred to as a "common radius HRTF set." In these examples, four position weights are shown in the far field set and two position weights are shown in the near field set using the convention Wxy, where x represents the ring number and y represents the position on the ring. Indicators WR1 and WR2 represent radial weights that may be used to decompose an object 22 into a weighted combination of the common radius HRTF set. For example, an object 22 may include a combination of a first source 20 and a second source 24 that provide the object 22 at a desired depth or position when rendered together.
[0031] exist Figure 1A and 1BIn the example of , as the audio object passes through a reference position 101, such as coinciding with the listener's position, the radial distance to the center of the listener's head can be measured. Two measured HRTF data sets that limit the radial distance can be identified. For each set, appropriate HRTF pairs (e.g., ipsilateral and contralateral) can be derived based on the desired azimuth and elevation of the sound source or object position. The final combined HRTF pair can be determined by interpolating the frequency response of each new HRTF pair. The interpolation can be based on the relative distance of the sound source to be rendered and the actual measured distance of each HRTF set. The sound source to be rendered can then be filtered by the derived HRTF pairs, and the gain of the resulting signal can be increased or decreased based on the distance to the listener's head. When the sound source is close to one of the listener's ears, the gain can be limited to avoid saturation.
[0032] Each HRTF set may span a set of measurements or synthesized HRTFs made only in the horizontal plane, or may represent the full range of HRTF measurements around the listener. Additionally, each HRTF set may have a fewer or greater number of samples based on the measured radial distance.
[0033] Various techniques can be used to generate an audio signal with distance or depth information. For example, U.S. Patent No. 9,973,874, entitled "Audiorendering using 6-DOF tracking" (which is incorporated herein by reference in its entirety) discloses a method for generating an audio signal with distance or depth information. Figure 2A -2C includes examples of generating binaural audio with distance cues, and Figure 3 Examples of determining HRTFs and interpolating between pairs of HRTFs are included in A-3C.
[0034] In an example, rendering audio objects in both the near field and the far field can enable rendering not only the depth of the object, but also the depth of any spatial audio mix decoded with active steering / panning, such as using ambisonics, matrix encoding, etc., and thereby achieve full translational head tracking (e.g., user movement) with 6 degrees of freedom (6-DOF) tracking and rendering. Various systems and methods for attaching depth information to an Ambisonic mix created, for example, by capture or by Ambisonic panning are discussed in U.S. Patent No. 9,973,874, entitled "Audio rendering using 6-DOF tracking," which is incorporated herein by reference in its entirety, and certain aspects of which are summarized herein. These techniques generally use first-order Ambisonics as an example, but may also be applied to third-order or other higher-order ambisonics.
[0035] Ambisonic Basics
[0036] While a multichannel mix will capture sounds as contributions from multiple input signals, Ambisonics provides a fixed set of signals that capture or encode the directions of all sounds in the sound field from a single point. In other words, the same ambisonic signal can be used to re-render the sound field on any number of speakers. In the case of multichannel, it can be restricted to reproduce the sources that originate from the combination of channels. For example, if there is no height channel, no height information is transmitted. In surround sound, on the other hand, information about the omnidirectional image can be captured and transmitted, and restrictions are usually only imposed at the point of reproduction.
[0037] Consider a set of 1st order (e.g., B-format) panning equations, which can largely be thought of as virtual microphones at the points of interest:
[0038] W = S*1 / √2, where W = omnidirectional component;
[0039] X=S*cos(θ)*cos(φ), where X= Figure 8 forward;
[0040] Y=S*sin(θ)*cos(φ), where Y= Figure 8 To the right;
[0041] Z = S*sin(φ), where Z = Figure 8 up;
[0042] And S is a signal to pan.
[0043] From these four signals (W, X, Y, and Z), a virtual microphone can be created pointing in any direction. In this way, a decoder receiving the signal can recreate virtual microphones pointing at each speaker used for rendering. This technique is effective for a large part, but in some cases it is only as good as capturing the response with real microphones. As a result, although the decoded signal may have the desired signal for each output channel, each channel may also include a certain amount of leakage or "bleed-through", so there are certain decoder design techniques that best represent the decoder layout, especially if it has non-uniform spacing.
[0044] These kinds of solutions can support head tracking, since decoding is done by combining weights of the WXYZ directional steering signals. To rotate a B-format mix, for example, a rotation matrix can be applied to the WXYZ signal before decoding, and the result will decode to the appropriately adjusted orientation. However, such solutions may not enable panning (e.g., user movement or change in listener position).
[0045] Active decoding extension
[0046] It may be necessary to prevent leakage and improve performance for non-uniform layouts. Active decoding solutions such as Harpex or DirAC do not form a virtual microphone for decoding. Instead, they examine the direction of the sound field, recreate the signal, and render it specifically in the direction they identify for each time-frequency. While this greatly improves the directivity of the decoding, it limits the directivity because hard decisions are used for each time-frequency block. In the case of DirAC, it makes a single directional assumption for each time-frequency. In the case of Harpex, two directional wavefronts can be detected. In either system, the decoder can provide control over how soft or hard the directional decision should be. Such a control is referred to as a "focus" parameter in this article, which can be a useful metadata parameter that allows soft focusing, inner panning, or other methods of softening the directionality assertion.
[0047] Even in the case of active decoders, distance or depth may be the missing function. While direction is directly encoded in the ambisonic panning equation, information about the distance of a sound source cannot be directly encoded, beyond a simple change in level or reverberation ratio based on the distance of the sound source. In ambisonic capture and decoding schemes, spectral compensation can be made for microphone “closeness” or “microphone proximity”, but this may not allow active decoding of one source located at 2 meters, and another at 4 meters, since the signal is limited to carrying only directional information. In fact, the performance of passive decoders relies on the fact that leakage will not be a problem if the listener happens to be located in the optimal position and all channels are equidistant. These conditions maximize the reproduction of the intended sound field.
[0048] Deep Coding
[0049] In an example, depth or distance information about audio objects may be encoded along with other information about the audio source. In an example, the transmission format or panning equation may be modified or enhanced to support the addition of depth indicators during content production. Unlike methods that apply depth cues such as loudness and reverberation changes in the mix, the methods discussed herein may include or allow for measuring or recovering distance or depth information about sources in the mix so that it can be rendered for final playback capabilities rather than production-side capabilities. Various approaches with different trade-offs are discussed in U.S. Patent No. 9,973,874, entitled “Audio rendering using 6-DOF tracking,” which is incorporated herein by reference in its entirety, including depth-based submixing and “D” channel encoding.
[0050] In depth-based sub-mixes, metadata may be associated with each mix. In an example, each mix may be tagged with information about (1) the distance of the mix and (2) the focus of the mix (e.g., an indication of how sharp the mix should be decoded so that, for example, the mix inside the listener's head is not decoded with too much active steering). Other embodiments may use a wet / dry mix parameter to indicate which spatial model to use if there is a choice of HRIR with more or less reflections (or an adjustable reflection engine). Preferably, appropriate assumptions will be made about the layout so that no additional metadata is needed to send it as, for example, an 8-channel mix, making it compatible with existing streams and tools.
[0051] In the 'D' channel encoding, an active decoder that supports depth can use information from the designated steering channel D. The depth channel can be used to encode time-frequency information about the effective depth of the ambisonic mix, which can be used by the decoder to render the distance of the sound source at each frequency. The 'D' channel can be encoded as a normalized distance, which in an example can be restored to the value 0 (head at the origin), 0.25 for just in the near field, and up to 1 for a source fully rendered in the far field. This encoding can be achieved by using an absolute value reference (e.g. 0dBFS) or by the relative amplitude and / or phase of one or more other channels (e.g. the 'W' channel).
[0052] Another method of encoding a distance channel may include using directional analysis or spatial analysis. For example, if only one sound source is detected at a particular frequency, the distance or depth associated with the sound source may be encoded. If more than one sound source is detected at a particular frequency, a combination of distances associated with the sound sources may be encoded, such as a weighted average. Alternatively, a depth or distance channel may be encoded by performing a frequency analysis of each individual sound source at a particular time frame. The distance at each frequency may be encoded, for example, as the distance associated with the most dominant sound source at that frequency, or as a weighted average of the distances associated with the active sound source at that frequency. The above techniques may be extended to additional D channels, for example, to a total of N channels. In the case where a decoder can support multiple sound source directions at each frequency, additional D channels may be included to support extended distances in these multiple directions.
[0053] Depth Rendering and Source Panning
[0054] The distance rendering techniques discussed here can be used to achieve a sense of depth or proximity in binaural rendering. Distance panning can be used to distribute sound sources over two or more reference distances. For example, a weighted balance of far-field and near-field HRTFs can be rendered to achieve a target depth. In the encoding or transmission of depth information, it is also useful to create sub-mixes at different depths using such a distance panner. Typically, the sub-mixes can each include or represent information with the same directionality for scene encoding, and the combination of multiple sub-mixes reveals depth information through their relative energy distribution. Such energy distributions can include direct quantification of depth, such as being evenly distributed or grouped for relevance, such as "near" and "far". In an example, such an energy distribution can include a relative turn relative to a reference distance, or closer or farther away, for example, some signals are understood to be closer than the rest of the far-field mix.
[0055] Audiovisual scene capture and spatial audio signal coding
[0056] Figure 2A An example of a system configured to receive audio information and visual information about an environment is generally shown. Figure 2B Examples of object recognition and depth analysis for the same environment are generally shown.
[0057] Figure 2A An example of includes a first environment 210, which may include various physical objects, and each physical object may emit or generate sound. The physical objects may have corresponding coordinates or positions, which may be defined relative to the origin of the environment, for example. Figure 2A In the example of Figure 2A In the example of FIG. 2 , the reference position 201 coincides with the sensor position.
[0058] Figure 2A Examples include an audio capture device 220 and a depth sensor 230. Information from the audio capture device 220 and / or the depth sensor 230 can be simultaneously received and recorded as an audiovisual program using various recording hardware and software. The audio capture device 220 may include a microphone or microphone array configured to receive audio information from the first environment 210. In an example, the audio capture device 220 includes a sound field microphone or an ambisonic microphone and is configured to capture audio information in a three-dimensional audio signal format.
[0059] The depth sensor 230 may include a camera, for example, having one or more lenses or image receivers. In an example, the depth sensor 230 includes a large field of view camera, such as a 360-degree camera. Information received or recorded from the depth sensor 230 as part of an audiovisual program may be used to provide an immersive or interactive experience to a viewer, such as allowing a viewer to "look around" the first environment 210, such as when the viewer uses a head tracking system or other program navigation tool or device.
[0060] Audio information, such as that received from the audio capture device 220 concurrently with video information received from the depth sensor 230 or camera, may be provided to the viewer. Audio signal processing techniques may be applied to the audio information received from the audio capture device 220 to ensure that the audio information tracks changes in the viewer's position or viewing direction as the viewer navigates the program, such as described in PCT patent application serial number PCT / US2019 / 40837, entitled “Non-coincident Audio-visual Capture System,” which is incorporated herein by reference in its entirety.
[0061] The depth sensor 230 can be implemented in various ways or using various devices. In an example, the depth sensor 230 includes a three-dimensional depth sensor configured to capture a depth image of the field of view of the first environment 210 and provide or determine a depth map from the depth image. The depth map may include information about the distance of one or more surfaces or objects. In an example, the depth sensor 230 includes one or more two-dimensional image sensors configured to receive incident light and capture image information about the first environment 210, and the image information can be processed using a processor circuit to identify objects and associated depth information. The depth sensor 230 may include a device that uses, for example, laser, structured light, time of flight, stereo vision, or other sensor technology to capture depth information about the first environment 210.
[0062] In an example, the depth sensor 230 may include a system with a transmitter and a receiver, and may be configured to determine the depth of an object using an active sampling technique. For example, the transmitter may transmit a signal and use timing information about the bounced signal to establish, for example, a point cloud representation of the environment. The depth sensor 230 may include or use two or more sensors, such as passive sensors, which may receive information from the environment and from different perspectives at the same time. Depth information about various objects in the environment may be determined using parallax in the received data or image. In an example, the depth sensor 230 may be configured to render a data set that can be used for clustering and object recognition. For example, if the data indicates a relatively large continuous plane at a common depth, an object may be identified at the common depth. Other techniques may be used similarly.
[0063] exist Figure 2A In the example of , the first environment 210 includes various objects at respective different depths relative to the reference position 201. The first environment 210 includes some objects that may generate or produce sounds and other objects that may not generate sounds. For example, the first environment 210 includes a first object 211, such as a quacking duck toy, and a second object 212, such as a roaring lion toy. The first environment 210 may include other objects, such as colored panels, boxes, cans, etc.
[0064] Figure 2B A depth map 250 representation of a first environment 210 is generally shown, including in this example a reference location 201, a depth sensor 230, and an audio capture device 220 as context. The depth map 250 displays some physical objects from the first environment 210 in lighter colors to designate these objects as belonging to surfaces that are closer or at smaller depths relative to the reference location 201. The depth map 250 displays other physical objects from the first environment 210 in darker colors to designate these other objects as belonging to surfaces that are farther away or at larger depths relative to the reference location 201. Figure 2B In the example of FIG. 2 , the first object 211 is identified or determined to be closer to the reference location 201 than the second object 212 , as indicated by their relative false color (grayscale) representations.
[0065] In an example, an audio capture device 220 may be used to receive audio or auditory information about the first environment 210. For example, the audio capture device 220 may receive a high-frequency, short-duration "crunch!" sound and a lower-frequency, longer-duration "growl!" sound emanating from the environment. For example, a processor circuit that may be coupled to the audio capture device 220 and the depth sensor 230 may receive audio information from the audio capture device 220 and may receive depth map information from the depth sensor 230. For example, the processor circuit may include a processor circuit that receives audio information from the audio capture device 220 and a processor circuit that receives depth map information from the depth sensor 230. Figure 4The processor circuit of the example processor circuit 410 can identify the correlation between the audio information and the depth information. Based on the identified correlation, the processor circuit can encode a spatial audio signal having information about audio objects at corresponding different depths, for example, using one or more systems or methods discussed herein.
[0066] Figure 3 A quantization example 300 is generally shown, which shows how information from the first environment 210 can be quantized to different depths. Figure 3 In the example of , the reference position 201 corresponds to the origin of the sound field. Figure 3 The viewing direction indicated in may correspond to Figure 2A Or the viewing direction indicated in 2B. In the example shown, the viewing direction is to the right of the reference position 201.
[0067] The quantization example 300 shows a second object 212 mapped to a position coinciding with a far-field depth or a first radius R1 relative to a reference position 201. That is, when the second object 212 is determined to be at a distance R1 from the reference position 201 (e.g., which may be determined using a depth map or other information from a depth sensor 230), sound from the second object 212 (e.g., which may be received using an audio capture device 220) may be encoded as a far-field signal. In an example, the second object 212 may have a position that may be specified by coordinates, such as radial or spherical coordinates, and may include information about a distance and angle (e.g., including an azimuth and / or elevation) from the reference position 201 or from a reference direction (e.g., a viewing direction). Figure 3 In the example of , the second object 212 may have a position defined by a radius R1, an azimuth angle of 0°, and an elevation angle of 0° ( Figure 3 The examples do not illustrate the "elevation" plane).
[0068] The quantization example 300 shows a first object 211 mapped to an intermediate depth or radius R2, which is smaller than the far-field depth or first radius R1 and larger than the near-field depth or first radius R2. N . That is, for example, when the first object 211 is determined to be at a distance R2 from the reference position 201 (e.g., which can be determined using a depth map or other information from the depth sensor 230), the sound from the first object 211 (e.g., which can be received using the audio capture device 220) can be encoded as a signal having a specific or specified depth R2. In an example, the first object 211 can have a position that can be specified by coordinates, such as radial or spherical coordinates, and can include information about the distance and angle (e.g., including azimuth and / or elevation) from the reference position 201 or from a reference direction (e.g., a viewing direction). Figure 3In the example of , the first object 211 may have a position defined by a radius R2, an azimuth angle α°, and an elevation angle 0° ( Figure 3 The examples do not illustrate the "elevation" plane).
[0069] In an example, an audio source or a virtual source may be generated and encoded using information from the audio capture device 220 and the depth sensor 230. For example, if the depth sensor indicates an object located at a distance (or radius) R2 relative to the reference position 210 and at an angle α° relative to the viewing direction, a first spatial audio signal may be provided, and the first spatial audio signal may include audio signal information from the audio capture device 220 (e.g., an audio object or a virtual source) located at R2 and the angle α°. If the depth sensor 230 indicates an object at a distance (or radius) R1 and an azimuth angle of 0°, a second spatial audio signal may be provided, and the second spatial audio signal may include audio signal information from the audio capture device 220 located at the radius R1 and the azimuth angle of 0°.
[0070] In an example, information from the depth sensor 230 may indicate the simultaneous presence of one or more objects in the environment. Various techniques may be used to determine which (if any) of the audio information from the audio capture device 220 corresponds to one or more of the corresponding objects. For example, information about the movement of a physical object over time, such as determined using information from the depth sensor 230, may be associated with changes in the audio information. For example, if a physical object is observed to move from one side of the environment to the other, and at least a portion of the audio information similarly moves from the same side of the environment to the other, a correlation may be found between the physical object and the portion of the audio information, and the audio information may be assigned a depth corresponding to the depth of the moving physical object. In an example, the depth information associated with the audio information may change over time, for example, along with the movement of the physical object. Various threshold conditions or learned parameters may be used to reduce the discovery of false positive correlations.
[0071] In an example, a classifier circuit or a software-implemented classifier module may be used to classify physical objects. For example, the classifier circuit may include a neural network or other recognizer circuit configured to process image information about the environment from the depth sensor 230, or may process image information from an image capture device configured to receive image information about the same environment. In an example, the classifier circuit may be configured to recognize various objects and provide information about corresponding auditory profiles associated with those objects. In an example, the auditory profile may include information about the audio, amplitude, or other characteristics of sounds that are known or believed to be associated with a particular object. Figure 3In an example of , the classifier circuit may be used to identify the first object 211 as a duck or a squeak toy, and in response, provide an indication that an auditory profile of a "squeak" sound (e.g., a sound that includes relatively high frequency information, has a short duration, and is highly transient) may generally be associated with sounds from the first object 211. Similarly, the classifier circuit may be used to identify the second object 212 as a lion, and in response, provide an indication that an auditory profile of a "roar" sound (e.g., a sound that includes relatively low frequency information, has a longer duration, and has large amplitudes and soft transients) may generally be associated with sounds from the second object 212. In an example, the spatial audio encoder circuit may be coupled to or may include the classifier circuit, and may use information about the classified object to identify a correlation between the input audio information and a physical object in the environment.
[0072] Figure 4 An example of a block diagram of an audio encoder system 400 for audio capture and spatial audio signal encoding is generally shown. Figure 4 An example of may include a processor circuit 410, which may include, for example, a spatial audio encoder circuit or module, or an object classifier circuit or module. In an example, a circuit configured according to the block diagram of the audio encoder system 400 may be used to encode or render one or more signals having corresponding directional or depth characteristics. Figure 4 One example of signal flow and processing is shown, and other interconnections or data sharing between or among the functional blocks shown are permitted. Similarly, processing steps can be redistributed between modules to accommodate various processor circuit architectures or optimizations.
[0073] In an example, the audio encoder system 400 may be used to receive an audio signal using an audio capture device 220, receive physical object position or orientation information using a depth sensor 230, and encode a spatial audio signal using the received audio signal and the received physical object information. For example, the circuit may encode a spatial audio signal using information about one or more audio sources or virtual sources in a three-dimensional sound field, such as each source or source group having different corresponding depth characteristics. In an example, the received audio signal may include a sound field or a 3D audio signal including one or more components or audio objects. The received physical object information may include information about classified objects and associated auditory profiles, or may include information about the placement or orientation of one or more physical objects in an environment.
[0074] In an example, spatial audio signal encoding may include using processor circuit 410 or one or more processing modules thereof to receive a first audio signal and determine a position, direction, and / or depth of a component of the audio signal. Reference frame coordinates or origin information for the audio signal components may be received, measured, or otherwise determined. One or more audio objects may be decoded for reproduction via a speaker or headphones, or may be provided to a processor for re-encoding into a new sound field format.
[0075] In an example, the processor circuit 410 may include various modules or circuits or software-implemented processes (such as may be performed using general-purpose or application-specific circuits) for performing audio signal encoding. Figure 4 In an example, the audio signal or data source may include the audio capture device 220. In an example, the audio source provides audio reference frame data or origin information to the processor circuit 410. The audio reference frame data may include information about a fixed or changing origin or reference point of the audio information, such as relative to the environment or relative to the depth sensor 230. The respective origins, reference positions, or orientations of the depth sensor 230 and the audio capture device 220 may change over time and may be taken into account when determining a correlation between a physical object identified in the environment and audio information from the environment.
[0076] In an example, the processor circuit 410 includes an FFT module 404 configured to receive audio signal information from the audio capture device 220 and convert the received signal(s) to the frequency domain. The converted signal may be processed using spatial processing, steering, or panning to change the position, depth, or reference frame of the received audio signal information.
[0077] In an example, the processor circuit 410 may include an object classifier module 402. The object classifier module 402 may be configured to implement one or more aspects of the classifier circuits discussed herein. For example, the object classifier module 402 may be configured to receive image or depth information from the depth sensor 230 and apply artificial intelligence-based tools (e.g., machine learning or neural network-based processing) to identify one or more physical objects present in the environment.
[0078] In an example, the processor circuit 410 includes a spatial analysis module 406, which is configured to receive the frequency domain audio signal from the FFT module 404 and, optionally, receive at least a portion of the audio data associated with the audio signal. The spatial analysis module 406 can be configured to use the frequency domain signal to determine the relative position of one or more signals or their signal components. For example, the spatial analysis module 406 can be configured to determine that a first sound source is located or should be located in front of a listener or a reference video position (e.g., 0° azimuth), and a second sound source is located or should be located to the right of the listener or the reference video position (e.g., 90° azimuth). In an example, the spatial analysis module 406 can be configured to process the received signals and generate a virtual source that is positioned or intended to be rendered at a specified position or depth relative to a reference video or image position, including when the virtual source is based on information from one or more input audio signals and each spatial audio signal corresponds to a corresponding different position, for example, relative to a reference position.
[0079] In an example, the spatial analysis module 406 is configured to determine the audio source position or depth and use a reference frame-based analysis to transform the source to a new position, such as a reference frame corresponding to a video source, as similarly discussed in PCT patent application serial number PCT / US2019 / 40837 entitled “Non-coincident Audio-visual Capture System”, the entire text of which is incorporated herein by reference in its entirety. Spatial analysis and processing of sound field signals including ambisonic signals are discussed in detail in U.S. patent application serial number 16 / 212,387 entitled “Ambisonic Depth Extraction” and U.S. patent application number 9,973,874 entitled “Audio rendering using 6-DOF tracking”, each of which is incorporated herein by reference in its entirety.
[0080] In an example, the processor circuit 410 may include a signal forming module 408. The signal forming module 408 may be configured to use the received frequency domain signal to generate one or more virtual sources, which may be output as sound objects with associated metadata, or may be encoded as spatial audio signals. In an example, the signal forming module 408 may use information from the spatial analysis module 406 to identify or place various sound objects at corresponding specified locations or corresponding depths in the sound field.
[0081] In one example, the signal formation module 408 can be configured to use information from both the spatial analysis module 406 and the object classifier module 402 to identify or place various sound objects identified by the spatial analysis module 406. In an example, the signal formation module 408 can use information about the identified physical object or audio object, such as information about the auditory profile or signature of the identified object, to determine whether the audio data (e.g., received using the audio capture device 220) includes information corresponding to the auditory profile. If there is sufficient correspondence between the auditory profile of a particular object (e.g., different from other objects in the environment) and a particular portion of the audio data (e.g., corresponding to a particular one or more frequency bands, or duration, or other portion of the audio spectrum in time), then the particular object can be associated with the corresponding particular portion of the audio data. In another example, artificial intelligence such as machine learning or neural network-based processing can be used to determine such correspondence.
[0082] In another example, the signal formation module 408 may use the results or products of the spatial analysis module 406 with information from the depth sensor 230 (optionally processed using the object classifier module 402) to determine the audio source location or depth. For example, the signal formation module 408 may use the correlation information or may determine whether there is a correlation between the physical objects or depths identified in the image data and the audio information received from the spatial analysis module 406. In an example, determining the correlation may be performed at least in part by comparing the direction or position of the identified visual object with the direction or position of the identified audio object. Other modules or portions of the processor circuit 410 may be similarly or independently used to determine the correlation between the information in the image data and the information in the audio data.
[0083] In examples with high correspondence or correlation, the signal formation module 408 can use a weighted combination of position information from audio and visual objects. For example, weights can be used to indicate the relative direction of the audio object that best matches the spatial audio distribution and can be used with depth information from the depth sensor visual data or image data. This can provide a final source position encoding that most accurately matches the depth capability of the spatial audio signal output to the auditory environment using the depth sensor and audio capture device.
[0084] In an example, the signal from the signal forming module 408 can be provided to other downstream processing modules that can help generate signals for transmission, reproduction, or other processing. For example, the spatial audio signal output from the signal forming module 408 can include or use virtualization processing, filtering, or other signal processing to shape or modify the audio signal or signal components. The downstream processing module can receive data and / or audio signal input from one or more modules and use signal processing to rotate or translate the received audio signal.
[0085] In an example, multiple downstream modules create multiple vantage points from which to observe the auditory environment. These modules can utilize the methods described in PCT patent application serial number PCT / US2019 / 40837, entitled "Non-coincident Audio-visual Capture System," which is incorporated herein by reference.
[0086] In an alternative example, the audio encoding / rendering portion of the signal forming module 408 can be replicated for each desired vantage point. In an example, the spatial audio signal output may include multiple encodings with corresponding different reference positions or orientations. In an example, the signal forming module 408 may include an inverse FFT module or may provide a signal to the inverse FFT module. The inverse FFT module may generate one or more output audio signal channels with or without metadata. In an example, the audio output from the inverse FFT module may be used as an input to a sound reproduction system or other audio processing system. In an example, the output may include a deeply extended ambisonic signal, such as one that may be decoded by a system or method discussed in U.S. Patent No. 10,231,073 "Ambisonic Audio Rendering with Depth Decoding", which is incorporated herein by reference. In an example, it may be desirable to keep the output format agnostic and support decoding of various layouts or rendering methods, such as including a monophonic backbone with position information, a base / bed mix, or other sound field representations including a surround sound format, for example.
[0087] In an example, multiple depth sensors may be coupled to the processor circuit 410, and the processor circuit 410 may use information from any one or more of the depth sensors to identify depth information about physical objects in the environment. Each depth sensor may have or may be associated with its own reference frame or corresponding reference position in the environment. Therefore, audio objects or sources in the environment may have different relative positions or depths relative to the reference positions of each depth sensor. As the viewer's perspective changes, such as when the video information changes from the perspective of a first camera to the perspective of a different second camera, the listener's perspective may similarly change by updating or adjusting the depth or orientation or rotation of one or more associated audio sources. In an example, the processor circuit 410 may be configured to adjust such perspective changes for the audio information, for example using cross-fading or other signal mixing techniques.
[0088] In an example, multiple audio capture devices (e.g., multiple instances of audio capture device 220) can be coupled to processor circuit 410, and processor circuit 410 can use information from any one or more audio capture devices to receive audio information about the environment. In an example, a specific one or combination of audio capture devices can be selected for use based at least in part on the proximity of a specific audio capture device to a specific physical object identified in the environment. That is, if a first audio capture device in the environment is closer to a first physical object, a deep encoded audio signal for the first physical object can be generated using audio information from the first audio capture device, such as when the first audio capture device captures sound information about the first physical object better than another audio capture device in the environment.
[0089] Figure 5 An example of a first method 500 that may include encoding a spatial audio signal is generally shown. The first method 500 may be performed at least in part using one or more portions of the processor circuit 410. In step 502, the first method 500 may include receiving audio scene information from an audio capture source in an environment. In an example, receiving the audio scene information may include using the audio capture device 220, and the audio scene information may include an audio signal with or without depth information. The audio scene information may optionally have an associated perspective, viewing direction, orientation, or other spatial characteristics.
[0090] In step 504, the first method 500 may include identifying at least one audio component in the received audio scene. Identifying the audio component may include, for example, identifying a signal contribution to a time-frequency representation of the received audio scene information. The audio component may include audio signal information for a specific frequency band or range, such as a duration of an audio program or a discrete portion of a program. In an example, step 504 may include identifying a direction associated with the audio scene information or associated with a portion of the audio scene information.
[0091] In step 506, the first method 500 may include receiving depth characteristic information about one or more objects in the environment from the depth sensor. Step 506 may include or use information from the depth sensor 230. In an example, step 506 may include using circuitry in the depth sensor 230 to receive the image or depth map information to process the information and identify the depth information, or step 506 may include using a different processor circuit coupled to the sensor. In an example, step 506 includes using the processor circuitry to identify one or more physical objects in the environment monitored by the depth sensor 230 in the image or depth map information, for example including identifying boundary information about the objects. In an example, the depth characteristic information may be provided relative to a reference position of the depth sensor 230 or a reference position of the environment.
[0092] In an example, step 506 may include receiving directional information about one or more objects in the environment, such as using information from the depth sensor 230. Step 506 may include identifying corresponding direction or orientation information for any physical objects identified. The direction or orientation information may be provided relative to a reference position or viewing direction. In an example, receiving the directional information in step 506 may include receiving information about an azimuth or altitude angle relative to a reference.
[0093] In step 508, the first method 500 may include encoding a spatial audio signal based on the identified at least one audio component and the depth characteristic information. Step 508 may include encoding the spatial audio signal using the audio scene information received from step 502 and using the depth characteristic received from step 506. That is, the spatial audio signal encoded at step 508 may include information of a virtual source having, for example, audio from the audio scene received at step 502 and depth characteristics from the depth information received at step 506. Encoding the spatial audio signal may be, for example, an ambisonic signal including audio information quantized at different depths. In an example, step 508 may include encoding the spatial audio signal based on the direction information identified in step 504 or received together with the depth characteristic at step 506. Thus, for example, in addition to the depth of the physical object to which the audio corresponds, encoding the spatial audio signal may also include information about the azimuth or elevation angle of the virtual source.
[0094] Figure 6 An example of a second method 600 is generally shown, which may include encoding a spatial audio signal based on correlation information. The second method 600 may be performed at least in part using one or more portions of the processor circuit 410. Figure 6 In an example of the present invention, step 610 may include determining a correlation between audio scene information from an environment and depth characteristics of physical objects identified in the environment. In an example, the audio scene information may be received or determined according to an example of the first method 500. Step 610 may include using the processor circuit 410 to analyze the audio information and determine a correspondence or likelihood of correspondence between the audio information and an object or location of an object in the environment.
[0095] For example, the processor circuit 410 may identify the positions of one or more potential audio sources in the environment that change over time, and the processor circuit 410 may also identify the positions of one or more physical objects in the environment that change over the same time, for example. If the change in position of at least one potential audio source corresponds to a change in position of at least one physical object, the processor circuit 410 may provide a strong correlation or positive indication that the audio source and the physical object are related.
[0096] Various factors or considerations may be used to determine the strength of the correlation or correspondence between the identified audio source and the physical object. For example, information from the object classifier module 402 may be used to provide information about specific audio features that are known or expected to be associated with a particular identified physical object. If an audio source having specific audio characteristics is found near the identified physical object, the audio information and the physical object may be considered to be corresponding or related. The strength or quality of the correspondence may be further identified or calculated to indicate a confidence level that the audio and physical object are related.
[0097] At steps 620 and 630, the strength of the correlation identified at step 610 may be evaluated, for example, using the processor circuit 410. At step 620, the second method 600 includes determining whether there is a strong correlation between the audio scene information and the depth characteristics of the particular physical object. In an example, whether the correlation is strong may be determined based on, for example, a quantized value of the correlation that may be determined at step 610. The quantized value of the correlation may be compared to various threshold levels, which may be specified or programmed, or may be learned over time by a machine learning system, for example. In an example, determining at 620 that the correlation is strong may include determining that the value of the correlation meets or exceeds a specified first threshold.
[0098] exist Figure 6 In the example of , if it is determined at step 620 that the correlation is strong, the second method 600 may proceed to step 622 and encode the spatial audio signal using the received depth characteristic of the specific object. That is, if a strong correlation is determined at step 620, it may be considered that the received or identified audio source information sufficiently corresponds to the specific physical object, so that the audio source may be located at the same depth or position as the specific physical object.
[0099] If, at step 620, the relative strength of the correlation does not meet the criteria from step 620, the second method 600 can proceed to step 630 to further evaluate the correlation. If the value of the correlation meets or exceeds the specified second threshold condition or value, it can be determined that the correlation is weak, and the second method 600 continues at step 632. Step 632 may include encoding the spatial audio signal using a reference depth characteristic of the audio source. In an example, the reference depth characteristic may include a far-field depth or other default depth. For example, if sufficient or minimal correlation is not found between a particular audio source or other audio information from an audio scene and an object identified in the environment, or if a particular or discrete object is not identified or cannot be identified, it can be determined that the audio source belongs to the far field or reference plane.
[0100] If the value of the correlation at step 630 does not satisfy the second threshold condition or value, the second method 600 may continue at step 634. Step 634 may include encoding the spatial audio signal using an intermediate depth characteristic of the audio source. The intermediate depth may be a depth that is closer to the reference position than it is to the far-field depth, and is a depth that is different from the depth of the identified physical object. In an example, if the correlation determined at step 610 indicates an intermediate certainty or confidence that a particular audio signal corresponds to a particular physical object, the particular audio signal may be encoded at a position or depth that is close to the particular physical object but not necessarily at the depth of the particular physical object.
[0101] In an example, the depth information may include an uncertainty measure that may be considered when determining relevance. For example, if a depth map indicates that an object is likely but not certain to be at a particular depth, audio information corresponding to the object may be encoded at a depth different from the particular depth, e.g., closer to the far field than the particular depth. In an example, if a depth map indicates that an object may appear at a range of different depths, audio information corresponding to the object may be encoded at a selected depth within the range, e.g., the farthest depth within the range. Systems and methods for encoding, decoding, and using audio information or a mixture having intermediate depth characteristics are discussed in detail in U.S. patent application serial number 16 / 212,387, entitled “Ambisonic Depth Extraction,” which is incorporated herein by reference in its entirety.
[0102] Figure 7 An example of a third method 700 is generally shown that may include providing an indication of confidence that audio scene information corresponds to a specified physical object. The third method 700 may be performed, at least in part, using one or more portions of the processor circuit 410.
[0103] At step 710, the third method 700 may include receiving physical object depth information using the depth sensor 230. In an example, step 710 may include receiving depth information about multiple objects and determining a combined depth estimate for a single object or for a group of multiple objects. In an example, step 710 may include determining a combined depth estimate that may represent a combination of different object depths of candidate objects in the environment. In an example, the depth information may be based on weighted depths or confidence indications about multiple objects. The confidence indication about the object may indicate the confidence or likelihood that the machine-recognized object corresponds to a particular object of interest or a particular audio object. In an example, the combined depth estimate based on multiple objects may be based on multiple video frames or from depth information that varies over time, for example, to provide a smooth or continuous indication of depth that slowly transitions rather than quickly jumps to different locations.
[0104] At step 720, the third method 700 may include receiving audio scene information from an audio sensor, for example using the audio capture device 220. Figure 7 In an example, step 730 may include parsing the received audio scene information into discrete audio signals or audio components. In an example, the received audio scene information includes information from a directional microphone, or from a microphone array, or from a sound field microphone. In an example, the received audio scene information includes a multi-channel mixture of multiple different audio signals, such as audio information from multiple different reference positions, viewing angles, viewing directions, or may have other similar or different characteristics. Step 730 may include generating discrete signals, such as discrete audio signal channels, time-frequency blocks, or other representations of different parts of the audio scene information.
[0105] Step 740 may include identifying the dominant direction of the audio object in each audio signal. For example, step 740 may include analyzing each discrete signal generated at step 730 to identify the audio objects therein. The audio objects may include, for example, audio information belonging to a specific frequency band, or audio information corresponding to a specific time or duration, or audio information including specified signal characteristics such as transient characteristics. Step 740 may include identifying the direction in which each audio object is detected in the audio scene.
[0106] Step 750 may include comparing the direction identified at step 740 with the object depth information received at step 710. Comparing the directions may include determining whether the direction of the audio object, for example relative to a common reference direction or a viewing direction, corresponds to the direction of a physical object in the environment. If a correspondence is identified, such as when the audio object and the physical object are both identified or determined to be located at an azimuth of 30° relative to a common reference angle, the third method 700 may include providing an indication of confidence that the audio scene (or a particular portion of the audio scene corresponding to the audio object) is associated with the identified physical object in the environment. For example, based on Figure 6 For example, correlation information can be used to encode an audio scene.
[0107] Figure 8 An example of a fourth method 800 that may include determining a correspondence between audio signal characteristics and received information about an audio scene is generally shown. The fourth method 800 may be performed, at least in part, using one or more portions of the processor circuit 410. At step 810, the fourth method 800 may include receiving audio scene information from an audio sensor, for example, using the audio capture device 220. At step 820, the fourth method 800 may include receiving image or video information, for example, from a camera or from a depth sensor 230.
[0108] At step 830, the fourth method 800 may include identifying objects in the image or video information received at step 820. Step 830 may include image-based processing, such as using clustering, artificial intelligence-based analysis, or machine learning, to identify physical objects that are or may be present in the image or field of view of the camera. In an example, step 830 may include determining depth characteristics of any one or more of the different objects identified.
[0109] At step 840, the fourth method 800 may include classifying the object identified at step 830. In an example, step 840 may include receiving image information using a neural network-based classifier or a machine learning classifier, and in response, providing a classification for the identified object. The classifier may be trained on various data, for example, to identify humans, animals, inanimate objects, or other objects that may or may not produce sounds. Step 850 may include determining audio characteristics associated with the classified object. For example, if a human male is identified at step 840, step 850 may include determining an auditory profile corresponding to a human male voice, such as one that may have various frequencies and transient characteristics. If a lion is identified at step 840, step 850 may include determining an auditory profile corresponding to a noise or speech known to be associated with a lion, such as one that may have frequencies and transient characteristics different from those associated with humans. In an example, step 850 may include or use a lookup table to map audio characteristics with various objects or object types.
[0110] At step 860, the fourth method 800 may include determining a correspondence between the audio characteristics determined at step 850 and the audio scene information received at step 810. For example, step 860 may include determining whether the audio scene information includes audio signal content that matches or corresponds to an auditory profile of an object identified in the environment. In an example, information about the correspondence may be used to determine a correlation between the audio scene and the detected physical object, which may be determined, for example, based on Figure 6 to use as an example.
[0111] The various illustrative logical blocks, modules, methods, and algorithmic processes and sequences described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and process actions are described above generally in terms of their functionality. Whether such functionality is implemented in hardware or software depends on the specific application and design constraints imposed on the overall system. The described functionality may be implemented in different ways for each specific application, but such implementation decisions should not be interpreted as causing a departure from the scope of this document. Embodiments of systems and methods for detecting depth information and using correlations between depth and audio information to encode spatial audio signals, as well as other techniques described herein, are described above, for example, in Fig. 9 The discussion described herein is operable in connection with various types of general purpose or special purpose computing system environments or configurations.
[0112] The various illustrative logical blocks and modules described in conjunction with the embodiments disclosed herein may be implemented or executed by a machine, such as a general purpose processor, a processing device, a computing device having one or more processing devices, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. General purpose processors and processing devices may be microprocessors, but alternatively, the processor may be a controller, a microcontroller or a state machine, a combination thereof, etc. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0113] In addition, one or any combination of software, programs, or computer program products implementing some or all of the various examples of virtualization and / or optimal adaptation described herein, or portions thereof, may be stored, received, transmitted, or read from any desired combination of computer or machine readable media or storage devices and communication media in the form of computer executable instructions or other data structures. Although the present subject matter is described in language specific to structural features and methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described herein. Instead, the specific features and actions described above are disclosed as example forms of implementing the claims.
[0114] Various systems and machines may be configured to perform or implement one or more of the signal processing tasks described herein, including but not limited to audio component localization or relocalization, or orientation determination or estimation, for example, using HRTF and / or other audio signal processing. Any one or more of the disclosed circuits or processing tasks may be implemented or performed using a general purpose machine or using a dedicated machine that performs various processing tasks, for example using instructions retrieved from a tangible, non-transitory, processor-readable medium.
[0115] Fig. 9 900 in which instructions 908 (e.g., software, programs, applications, applet, application program, or other executable code) may be executed to cause the machine 900 to implement any one or more of the methodologies discussed herein. For example, the instructions 908 may cause the machine 900 to perform any one or more of the methodologies described herein. The instructions 908 may transform a general-purpose, non-programmed machine 900 into a specific machine 900 that is programmed to perform the functions described and illustrated in the manner described.
[0116] In an example, the machine 900 can operate as a standalone device, or can be coupled (e.g., networked) to other machines or devices or processors. In a networked deployment, the machine 900 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer (or distributed) network environment. The machine 900 may include a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular phone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a network device, a network router, a network switch, a bridge, or any machine capable of sequentially or otherwise executing instructions 908, which specify actions to be taken by the machine 900. In addition, although only a single machine 900 is shown, the term "machine" may be considered to include a collection of machines that execute instructions 908 individually or in conjunction to implement any one or more of the methods discussed herein. In an example, the instructions 908 may include instructions that can be executed using the processor circuit 410 to implement one or more of the methods discussed herein.
[0117] The machine 900 may include various processors and processor circuits (e.g., Fig. 9902), memory 904, and I / O components 942, which may be configured to communicate with each other via a bus 944. In an example, processor 902 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 906 that executes instructions 908 and a processor 910. The term "processor" is intended to include a multi-core processor that may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Fig. 9 Multiple processors are shown, but the machine 900 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with single cores, multiple processors with multiple cores, or any combination thereof, for example, to provide the processor circuit 410.
[0118] The memory 904 may include a main memory 912, a static memory 914, or a storage unit 916, for example, which may be accessed by the processor 902 via a bus 944. The memory 904, the static memory 914, and the storage unit 916 may store instructions 908 that implement any one or more of the methods, functions, or processes described herein. During execution of the instructions 908 by the machine 900, the instructions 908 may also reside, in whole or in part, within the main memory 912, within the static memory 914, within a machine-readable medium 918 within the storage unit 916, within at least one processor (e.g., within a cache memory of a processor), or within any suitable combination thereof.
[0119] I / O components 942 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurements, etc. The specific I / O components 942 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It should be understood that I / O components 942 may include Fig. 9928 and 930. In various example embodiments, the I / O components 942 may include output components 928 and input components 930. The output components 928 may include visual components (e.g., displays such as plasma display panels (PDFs), light emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. The input components 930 may include alphanumeric input components (e.g., keyboards, touch screens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touch pads, trackballs, joysticks, motion sensors, or other pointing instruments), touch input components (e.g., physical buttons, touch screens that provide location and / or force of touch or touch gestures, or other tactile input components), audio input components (e.g., microphones), video input components, etc.
[0120] In an example, the I / O component 942 may include a biometric component 932, a motion component 934, an environmental component 936, or a position component 938, as well as a number of other components. For example, the biometric component 932 includes a component configured to detect the presence or absence of a human, pet, or other individual or object, or a component configured to detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body postures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identify people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or recognition based on brain waves), etc. The motion component 934 may include an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc.
[0121] The environment component 936 may include, for example, an illumination sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor for detecting concentrations of hazardous gases or a gas detection sensor for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. The position component 938 includes a location sensor component (e.g., a GPS receiver component, an RFID tag, etc.), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure from which altitude can be derived), an orientation sensor component (e.g., a magnetometer), etc.
[0122] I / O components 942 may include communication components 940 operable to couple machine 900 to network 920 or device 922 via coupling 924 and coupling 926, respectively. For example, communication components 940 may include a network interface component or another suitable device to interface with network 920. In further examples, communication components 940 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, Components (e.g. Low energy), Components and other communication components that provide communication via other modes. Device 922 can be another machine or any of a variety of peripheral devices (e.g., peripheral devices coupled via USB).
[0123] In addition, the communication component 940 can detect an identifier or include a component operable to detect an identifier. For example, the communication component 940 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional bar codes such as Universal Product Code (UPC) bar codes, multi-dimensional bar codes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar codes, and other optical codes), or an auditory detection component (e.g., a microphone for identifying an audio signal of a tag). In addition, various information can be derived through the communication component 940, such as a location via Internet Protocol (IP), a geographic location via Internet Protocol (IP), a ... Location by signal triangulation, or by detecting NFC beacon signals that can indicate a specific location, etc.
[0124] Various memories (e.g., memory 904, main memory 912, static memory 914, and / or memory of processor 902) and / or storage unit 916 may store one or more instructions or data structures (e.g., software) that implement or are used by any one or more of the methods or functions described herein, which, when executed by a processor or processor circuitry, (e.g., instructions 908) result in various operations to implement the embodiments discussed herein.
[0125] Instructions 908 may be sent or received over network 920 using a transmission medium, via a network interface device (e.g., a network interface component included in communication component 940), and using any of a number of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 908 may be sent or received to device 922 via coupling 926 (e.g., a peer-to-peer coupling) using a transmission medium.
[0126] In this document, as is common in patent documents, the terms "a" or "an" are used to include one or more, independent of any other instance or usage of "at least one" or "one or more." In this document, the term "or" is used to refer to a non-exclusive or, such that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise stated. In this document, the terms "including" and "in which" are used as plain-language equivalents of the respective terms "comprising" and "wherein."
[0127] Conditional language used herein, such as, among others, "can," "might," "may," "could," "e.g.," and the like, unless specifically stated otherwise or otherwise understood in the context of use, is generally intended to convey that certain embodiments include certain features, elements, and / or states, while other embodiments do not include certain features, elements, and / or states. Thus, such conditional language is generally not intended to imply that one or more embodiments require features, elements, and / or states in any way, or that one or more embodiments must include logic for determining, with or without author input or prompting, whether such features, elements, and / or states are included or will be performed in any particular embodiment.
[0128] Although the foregoing detailed description has shown, described and pointed out novel features applicable to various embodiments, it is to be understood that various omissions, substitutions and changes in the form and details of the devices or algorithms shown may be made, as will be appreciated, and that certain embodiments of the invention described herein may be implemented in a form that does not provide all of the features and benefits set forth herein, as some features may be used or implemented separately from other features.
[0129] In addition, although the subject matter has been described in language specific to structural features or methods or acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Instead, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for processing an audio signal, comprising: receiving audio scene information from an audio capture source in the environment; identifying at least one audio component in a received audio scene; receiving, from a depth sensor, depth characteristic information regarding distances from a reference location to a first physical object and a second physical object in the environment; correlating the identified at least one audio component with received depth characteristic information of a selected one of the first and second physical objects; as well as For the selected one of the first and second physical objects, a spatial audio signal is encoded based on the identified at least one audio component and the depth characteristic information.
2. The method of claim 1, wherein: The at least one audio component is determined using information about signal contributions to the received time-frequency representation of the audio scene information. 3 . The method of claim 1 , further comprising determining, for the at least one audio component, a first direction and a reference depth relative to the audio capture source.
4. The method of claim 3, further comprising: determining a confidence that at least a portion of the depth characteristic information from the depth sensor corresponds to the at least one audio component; as well as providing a first depth of the identified at least one audio component using the determined confidence level; Wherein encoding the spatial audio signal comprises using the first depth.
5. The method of claim 4, wherein: Providing the first depth includes: When the confidence level is high, providing the first depth based on information from the depth sensor; When the confidence level is low, providing the first depth as the reference depth; and When the confidence is medium, the first depth is provided as a depth between the reference depth and a depth determined using the depth sensor.
6. The method of claim 4, wherein: Determining the confidence level includes using a computer vision processor to classify objects identified in the environment and determine whether the at least one audio component includes or is likely to include audio from at least one of the objects.
7. The method of claim 4, wherein: Determining the confidence level includes determining a confidence level that the identified at least one audio component corresponds to a particular one of the first and second physical objects.
8. The method of claim 4, wherein: Determining the confidence level includes: identifying one or more data clusters in the depth characteristic information from the depth sensor, and A first direction of the at least one audio component is associated to the identified one or more data clusters.
9. The method of claim 3, further comprising: receiving, from the depth sensor, depth characteristic information about the first and second physical objects having respective depth magnitude and depth direction characteristics; determining, for the first and second physical objects, respective confidence indications that the depth characteristic information corresponds to the at least one audio component; and determining a combined depth characteristic based on the respective confidence indications; Wherein encoding the spatial audio signal comprises using the combined depth characteristic.
10. The method of claim 1, wherein: Encoding the spatial audio signal includes encoding a depth-extended ambisonic signal based on the audio scene and the depth characteristic information.
11. The method of claim 1, wherein: Receiving the audio scene information from an audio capture source includes receiving the audio scene information from one or more of a multi-transducer microphone, a sound field microphone, a microphone array, and an ambisonic microphone.
12. The method of claim 1, wherein: Receiving the depth characteristic information includes receiving time-varying depth characteristic information about the first physical object, the time-varying depth characteristic information indicating movement of the first physical object in the environment, and The encoding of the spatial audio signal includes encoding based on the audio scene and the time-varying depth characteristic information.
13. The method of claim 1, further comprising: determining a classification of the first physical object using an image-based object classifier; as well as The condition for encoding the spatial audio signal is that it is determined based on the classification that the first part of the audio scene information includes or may include audio information from the first physical object.
14. The method of claim 13, further comprising determining whether the first portion of the audio scene information includes or may include audio information from the first physical object based on audio frequency content associated with the classification of the first physical object and audio frequency content of the audio information.
15. The method of claim 1, wherein: Receiving the depth characteristic information includes analyzing information from one or more of a three-dimensional video capture system, a stereo camera, or an active depth detector configured to measure time-of-flight information of a laser or infrared detector signal.
16. A system for processing an audio signal, comprising: an audio capture source configured to capture an audio scene in an environment; a depth sensor configured to provide depth characteristic information about a plurality of objects in the environment relative to a reference position of the depth sensor; as well as The processor circuit is configured to: identifying at least one audio component in the audio scene, the at least one audio component having a first direction and a reference depth relative to the audio capture source selecting a first object of the plurality of objects to be associated with the identified at least one audio component; as well as A spatial audio signal is encoded based on the identified at least one audio component in the audio scene and depth characteristic information associated with the identified at least one audio component.
17. The system of claim 16, wherein: The audio capture source includes one or more of a multi-transducer microphone, a sound field microphone, a microphone array, and an ambisonic microphone.
18. The system of claim 16, wherein: The depth sensor includes one or more of a laser, a modulated light source, a stereo camera, a depth detector, an infrared sensor, and a camera array.
19. The system of claim 16, wherein: The processor circuit is configured to encode the spatial audio signal into a depth extended ambisonic signal based on the audio scene and depth characteristics of the first object.
20. The system of claim 16, wherein: The processor circuit is configured to encode the spatial audio signal using a weighted combination of depth information about the plurality of objects.
21. The system of claim 16, wherein: The processor circuit is configured to determine a confidence that information from the audio scene corresponds to a first object among the plurality of objects in the environment, and wherein the processor circuit is configured to encode the spatial audio signal based on the determined confidence meeting or exceeding a specified confidence threshold.
22. The system of claim 16, wherein: The depth sensor is configured to determine the depth characteristic of the object using information from one or more data clusters identified in the information from the depth sensor.
23. The system of claim 16, further comprising an object classifier circuit configured to determine a classification of the object; in, The processor circuit is configured to determine a correspondence between the classification of the object and the at least one audio component, and wherein the processor circuit is configured to encode the spatial audio signal based on a value of the determined correspondence satisfying a threshold correspondence condition.
24. An audio signal encoder device, comprising: A processor and a non-transitory computer-readable medium operatively coupled to the processor, the non-transitory computer-readable medium including instructions stored in association therewith, the instructions being accessible and executable by the processor, wherein the instructions include: instructions that when executed receive an audio scene from an audio capture source in the environment; instructions that when executed identify a first audio component in the audio scene from a plurality of different audio components in the audio scene; instructions that when executed receive image information regarding the environment, the image information including depth information regarding a distance between a reference position of an object depth sensor and one or more objects in the environment; instructions that when executed identify a classification of a first object of the one or more objects using a neural network-based classifier; Instructions that when executed identify audio characteristics associated with the identified classification of the first object; and Instructions that when executed determine whether the audio characteristic corresponds to the first audio component identified in the audio scene.
25. The audio signal encoder device of claim 24, further comprising instructions that when executed conditionally encode the spatial audio signal, the instructions comprising instructions that when executed perform the following operations: encoding the spatial audio signal based on depth information about the first object in the environment when the audio characteristic corresponds to the first audio component identified in the audio scene; and When the audio characteristic does not correspond to the first audio component identified in the audio scene, the spatial audio signal is encoded based on a reference depth, wherein the reference depth is characteristic of the audio capture source and / or the environment.
26. The audio signal encoder device of claim 24, further comprising instructions that when executed encode a spatial audio signal using the first audio component and using depth information about the first object in the environment.
Citation Information
Patent Citations
Ambisonic audio rendering with depth decoding
US10231073B2
Ambisonic depth extraction
US20190313200A1
Audio rendering using 6-DOF tracking
US9973874B2
Screen-relative rendering of audio and encoding and decoding of audio for such rendering
CN105723740A
Distance panning using near / far-field rendering
CN109891502A