Amplitude-independent window size in audio coding

By using transforms at different frequencies of window size independent of amplitude in audio processing, the problem of inaccurate time response at resonant frequencies is solved, the time and amplitude sensitivity of audio devices is improved, and the spatial decomposition capability in virtual reality and augmented reality is enhanced.

CN113272895BActive Publication Date: 2025-09-05GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980024488.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-16
Publication Date
2025-09-05
Estimated Expiration
2039-12-16

AI Technical Summary

Technical Problem

When existing audio processing technologies deal with resonance phenomena related frequencies, it is difficult to effectively distinguish and process transient sounds, resulting in inaccurate time responses and affect the spatial decomposition capabilities of audio devices in virtual reality and augmented reality.

Method used

Using amplitude-independent window sizes, improve time response by using different window sizes at different frequencies, especially using smaller window sizes at resonance frequencies to improve time resolution, and using larger window sizes at other frequencies to improve amplitude sensitivity.

Benefits of technology

The time response and amplitude sensitivity of audio devices at processing resonance frequency is improved, the spatial decomposition ability of audio devices in virtual reality and augmented reality is enhanced, and the accuracy of sound directional determination is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113272895B_ABST
    Figure CN113272895B_ABST
Patent Text Reader

Abstract

A computer-implemented method may include receiving a first signal corresponding to a first acoustic energy flow; applying a transform to the received first signal using a first amplitude-independent window size at at least a first frequency and a second amplitude-independent window size at a second frequency, the second amplitude-independent window size improving a time response at the second frequency, wherein the second frequency experiences amplitude reduction due to a resonance phenomenon associated with the first frequency; and storing a first encoded signal based on applying the transform to the received first signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document generally relates to amplitude-independent window size in audio coding. Background Art

[0002] Audio processing remains a crucial aspect of today's technological landscape. Digital assistants, used in both personal and professional situations to assist users with various tasks, are trained to recognize speech to detect prompts and instructions. Speech recognition is also used to create digitally accessible records of events people are discussing. In the rapidly growing world of virtual and / or augmented reality, audio processing provides users with a plausible auditory experience to optimally perceive and interact with digital environments. Summary of the Invention

[0003] In an aspect of the present disclosure, a computer-implemented method is provided. The method includes receiving a first signal corresponding to a first acoustic energy flow, applying a transform to the received first signal using a first amplitude-independent window size at at least a first frequency and a second amplitude-independent window size at a second frequency, the second amplitude-independent window size improving a time response at the second frequency, wherein the second frequency experiences amplitude reduction due to a resonance phenomenon associated with the first frequency, and storing a first encoded signal based on applying the transform to the received first signal.

[0004] For example, the first frequency may be approximately 3 kHz, and the second frequency may be approximately 1.5 kHz or approximately 10 kHz. The first amplitude-independent window size may be approximately 18-30 ms (e.g., approximately 24 ms). The second amplitude-independent window size may be approximately 3-9 ms (e.g., approximately 6 ms).

[0005] The method may further include mapping a first amplitude-independent window size to the first frequency based on the first frequency being associated with energy integration in human hearing.

[0006] The method may further include mapping a second amplitude-independent window size to the second frequency based on the second frequency being associated with energy differentiation in human hearing.

[0007] The first amplitude-independent window size may be applied to all frequencies of the received first signal except for a frequency band at the second frequency. The first amplitude-independent window size may be larger than the second amplitude-independent window size. The first amplitude-independent window size may be larger than an integer multiple of the second amplitude-independent window size. The first amplitude-independent window size may be approximately four times larger than the second amplitude-independent window size.

[0008] The method may further include using a third amplitude-independent window size when applying the transform to the first received signal, the third amplitude-independent window size being used at a third frequency not associated with resonance phenomena, the third amplitude-independent window size being different from the first and second amplitude-independent window sizes.

[0009] The third amplitude-independent window size may be smaller than the first amplitude-independent window size. The third amplitude-independent window size may be approximately half the size of the first amplitude-independent window size. The third amplitude-independent window size may be larger than the second amplitude-independent window size. The third amplitude-independent window size may be approximately twice the size of the second amplitude-independent window size. The third amplitude-independent window size may be smaller than the first amplitude-independent window size.

[0010] Applying the transform at a first frequency using a first amplitude-independent window size may generate a first result, wherein applying the transform at a second frequency using a second amplitude-independent window size may generate a second result, the method further comprising storing the second result more frequently than storing the first result.

[0011] The method may further include storing the second result with a lower precision than the first result.

[0012] The method may further include using a third amplitude-independent window size when applying the transform at a third frequency, the third amplitude-independent window size improving the time response at the third frequency, the third frequency experiencing amplitude reduction due to a resonance phenomenon associated with the first frequency.

[0013] The second and third frequencies may be located on opposite sides of the first frequency.

[0014] The third amplitude-independent window size may be approximately equal to the second amplitude-independent window size.

[0015] The second and third amplitude-independent window sizes may be smaller than the first amplitude-independent window size.

[0016] The first audio file may include a first encoded signal, and the method may further include: receiving a second signal corresponding to a second acoustic energy flow; applying a transform to the received second signal using a first amplitude-independent window size at at least a first frequency and a second amplitude-independent window size at a second frequency; storing a second encoded signal, the second encoded signal being based on applying the transform to the received second signal, wherein the second audio file includes the second encoded signal; and determining a difference between the first and second audio files.

[0017] Determining the difference may include playing the first and second audio files into a human hearing model, the model including resonance phenomena.

[0018] In an aspect of the present disclosure, there is provided a computer program product tangibly embodied in a non-transitory storage medium, the computer program product comprising instructions that, when executed by a processor, cause the processor to perform the operations of any of the methods described herein.

[0019] Optional features of one aspect may be combined with any other aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 An example of a system is shown.

[0021] Figure 2 An example of determining the directionality of a sound source is shown.

[0022] Figure 3 An example of an audio signal is shown.

[0023] Figure 4 An example of an audio encoder is shown.

[0024] Figure 5 Examples of window sizes are shown.

[0025] Figure 6 An example of decoding is schematically shown.

[0026] Figure 7 An example of an audio analyzer is shown.

[0027] Figure 8 An example of a method is shown.

[0028] Figure 9 Examples of computer devices and mobile computer devices that can be used to implement the techniques described herein are shown.

[0029] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION

[0030] This document describes examples of audio processing using amplitude-independent window sizes. In some implementations, a relatively large window size can be used when processing signals at frequencies associated with resonance phenomena in the human ear. For example, the window size can be approximately twice the window size used for another frequency. In some implementations, a relatively small window size can be used when processing signals at frequencies that experience amplitude reduction due to resonance phenomena. For example, the window size can be approximately two times smaller than the window size used for another frequency.

[0031] Figure 1An example of a system 100 is shown. The system 100 can be used with one or more other examples described elsewhere herein. The system 100 includes a plurality of sound sensors 102, including but not limited to microphones. For example, one or more omnidirectional microphones and / or microphones with other spatial characteristics can be used. The sound sensor 102 detects audio in a space 104. For example, the space 104 can be characterized by a structure (such as in a recording studio with a specific ambient impulse response), or can be characterized as being substantially free of surrounding structure (such as in a substantially empty space). The output of the sound sensor can be provided to a resonance enhancement encoder 106. The resonance enhancement encoder 106 can perform improved encoding on the audio signal from the sound sensor 102. In some implementations, the resonance enhancement encoder 106 can improve the time response of the sound signal at one or more specific frequencies associated with the resonance phenomenon. The time response can be improved by increasing the time resolution of the encoding process at one or more frequencies. For example, the time resolution can be increased by including relatively less audio content (e.g., a shorter time portion of the signal) when applying the transform. This approach may improve the ability of system 100 (or another component, including but not limited to an audio analyzer) to determine sound directionality; that is, to distinguish two or more sound sources from one another based at least in part on their spatiality.

[0032] Before the resonance enhancement encoder 106 encodes the signal from the sound sensor 102, one or more types of conditioning can be performed on the signal. In some implementations, the signal can be processed to generate a particular representation (e.g., according to a pre-specified format). For example, the representation can be decomposed into individual channels of sound from the sound sensor 102.

[0033] When encoding, the resonance enhancement encoder 106 may apply a transformation to the signal from the sound sensor 102. The transformation may include applying two or more different window sizes to various frequencies (or frequency bands) of the signal from the sound sensor 102. In some implementations, the window size is amplitude-independent, that is, the window size is applied to a specific at least one frequency (frequency band) regardless of the characteristics of the signal in this regard. For example, the resonance enhancement encoder 106 may not consider whether the frequency (frequency band) contains a sustained level of sound energy, and / or whether the frequency (frequency band) contains any transients, such as a relatively short duration region with a higher amplitude than the surrounding portions of the waveform. Using different window sizes can help address listening-related situations, including but not limited to acoustic characteristics, such as resonance phenomena.

[0034] After encoding, the encoded signal may be stored, forwarded, and / or sent to another location.For example, channel 108 represents one or more ways in which the encoded audio signal may be managed, such as by transmission to another system for playback.

[0035] If the audio of the encoded signal should be played, a decoding process can be performed. Such a decoding process can be performed by a resonance enhancement decoder 110. For example, the resonance enhancement decoder 110 can perform operations in a manner substantially opposite to that of the resonance enhancement encoder 106. For example, an inverse transform can be performed in a decoding module that partially or completely recovers the specific representation generated by the resonance enhancement encoder 106. The resulting audio signal can be stored and / or played as appropriate. For example, the system 100 can include two or more audio playback sources 112 (including but not limited to speakers) to which the processed audio signal can be provided for playback.

[0036] A representation of the signal from the sound sensor 102 can be played on headphones, and the system 100 can calculate what should be rendered in the headphones. In some implementations, this can be applied to situations involving virtual reality (VR) and / or augmented reality (AR). In some implementations, the rendering can depend on how the user turns his or her head. For example, a sensor that notifies the system of the head's orientation can be used, and then the system can make the person hear sounds from directions that are independent of the head's orientation. As another example, a representation of the signal from the sound sensor 102 can be played on a set of speakers. That is, first, the system 100 can store or send a description of the sound field surrounding the listener. At the resonance enhancement decoder 110, the content that each speaker should produce can then be calculated to produce a sound field around the listener's head. That is, the method exemplified herein can promote spatial decomposition of sound.

[0037] Figure 2 An example of determining the directionality of a sound source is shown. Examples of spatial outlines are schematically shown here. Physical space 200 may include any spatial extension, including but not limited to a room, an outdoor area, or an atmospheric area. Circle 202 schematically represents a listener in each case. For the purposes of this example, the listener represented by circle 202 may be a device according to the present subject matter (e.g., Figure 1 The device may be a system 100 in the present invention, or a human listener. The listener perceives the sound as a flow of acoustic energy. For example, a device may perceive the sound in order to encode it (e.g., the device may be an encoder according to the present subject matter). As another example, a device may perceive the sound in order to analyze it in order to make a difference determination (e.g., the device may be an audio analyzer according to the present subject matter). As another example, a human listener may perceive the sound in the physical space 200 by being an active or passive listener in or near the space.

[0038] People 204A-C are schematically shown in physical space 200. The human-shaped symbols represent any type of sound source that a listener can hear. Such sounds can be generated by humans (e.g., voices, songs, or other speech patterns), nature (e.g., wind, animals, or other natural phenomena), or technology (e.g., machines, speakers, or other man-made devices). That is, the present subject matter relates to sounds from one or more types of sources, regardless of whether the sounds are caused by humans. The position of people 204A-C around circle 202 indicates that circle 202 can perceive sounds from multiple separate directions. Here, it can be said that each person 204A-C has a corresponding spatial profile 206A-C associated with them. Spatial profiles 206A-C represent the directions from which a listener can perceive sounds arriving. Spatial profiles 206A-C correspond to how sounds from different sources are captured: some sounds arrive directly from the source, while other sounds (generated simultaneously) first bounce off one or more surfaces before being perceived. That is, the sound represented here by person 204A may have spatial profile 206A, the sound represented here by person 204B may have spatial profile 206B, and the sound represented here by person 204C may have spatial profile 206C.

[0039] In the context of a room, the concept of a spatial profile is a generalization of this illustrative example. Here, a spatial profile includes both the direct path and all reflected paths that sound from a source travels along to reach a listener in circle 202. In various circumstances, such as when physical space 200 is relatively structure-free or suppresses echoes and other acoustic reflections, the direct path of acoustic energy may dominate at circle 202. In some implementations, the term "direction" may be considered to have a general meaning and be equivalent to a set of directions representing the direct path and all reflected paths. In some implementations, more or fewer spatial profiles than spatial profiles 206A-C may be present.

[0040] Different listeners represented by circle 202 may have different abilities to spatially resolve arriving sounds having respective spatial profiles 206A-C. For example, a human may be able to identify ten, perhaps fifteen, sound sources in parallel based on their respective spatial profiles 206A-C. On the other hand, a device (e.g., a computer-based system prior to the present subject matter) may be able to distinguish significantly fewer sound sources in parallel than a human listener. For example, existing computers have been able to distinguish fewer than three simultaneous sound sources in parallel (e.g., about two sound sources). This can result in limitations in the ability of audio devices to perform spatial decomposition (e.g., in AR / VR systems). Therefore, using a computer-based system with improved spatial decomposition capabilities can allow the listener of circle 202 to distinguish between more spatial profiles 206A-C.

[0041] Determining the directionality of a sound can depend on a number of factors, including but not limited to time response. In some implementations, time response can represent the system's ability to detect the onset or end of an acoustic phenomenon in time. For example, improved time response corresponds to a system that is better at accurately pinpointing when a sound begins or ends. This applies to any type of sound, including both sustained sound energy levels and transients.

[0042] Figure 3 An example of an audio signal 300 is shown. Audio signal 300 may appear in or be considered in one or more other examples described elsewhere herein. Here, audio signal 300 includes input signals 302A-C, which may be referred to as respective inputs to a system. That is, each input signal 302A-C represents an audio signal (e.g., acoustic energy flow) that may be recorded by a computer system and / or a human listener. Some examples described with reference to signal 300 will be based on human listeners. Input signals 302A-C have different frequencies (or frequency bands). In some implementations, input signal 302A is associated with a frequency of approximately 1.5 kHz. For example, this corresponds to a period of approximately 666 microseconds (μs). In some implementations, input signal 302B is associated with a frequency of approximately 3.0 kHz. For example, this corresponds to a period of approximately 333 μs. In some implementations, input signal 302C is associated with a frequency of approximately 10.0 kHz. For example, this corresponds to a period of approximately 100 μs. Input signals 302A-C may be separate and independent of each other, or they may be part of the same acoustic signal. For example, a bandpass filter array may be used to separate an input signal into multiple components, including but not limited to input signals 302A-C.

[0043] Each input signal 302A-C can include any type of audio signal content. In some implementations, input signal 302A includes waveform 304A. For example, waveform 304A can be a relatively homogeneous set of waves having similar or identical amplitudes and a frequency of approximately 1.5 kHz. In some implementations, input signal 302B includes waveform 304B. For example, waveform 304B can be a relatively homogeneous set of waves having similar or identical amplitudes and a frequency of approximately 3.0 kHz. In some implementations, input signal 302C includes waveform 304C. For example, waveform 304C can be a relatively homogeneous set of waves having similar or identical amplitudes and a frequency of approximately 10.0 kHz.

[0044] One or more acoustic phenomena may affect the perception of input signals 302A-C. In some implementations, resonance may occur. For example, the human ear resonates at approximately 3 kHz, which can be explained by the elastic-viscous properties of the membrane that oscillates in the ear and the interaction of the hair cells on that membrane. This resonance phenomenon is common in all humans and may have some impact on how the human ear receives sound waves.

[0045] Starting with input signal 302B, which is at a resonant frequency of approximately 3.0 kHz, the ear will receive signal 306B affected by the resonance. Resonance can cause amplification of input signal 302B. If input signal 302B has a specific amplitude, signal 306B can have an amplitude that is a larger multiple. For example, the amplitude of signal 306B can be approximately twice the amplitude of input signal 302B (e.g., amplified by approximately +6 dB). Resonance can also cause a temporal localization of transients at frequencies of approximately 3.0 kHz. That is, the accumulation of energy associated with resonance can accumulate signal energy over time. Therefore, the frequency of 3.0 kHz can be associated with energy integration in human hearing. For example, this can blur the temporal characteristics of transients and attenuate them (e.g., by a factor of approximately 2). This blurring can make transients more difficult to detect (e.g., the transients can appear to disappear). This can result in transients being heard for longer than they occurred (e.g., the transients may appear to be smeared in time). For example, signal 306B may include a waveform 308B that is several times longer (eg, three times longer) than waveform 304B.

[0046] Referring now to input signals 302A and 302C, these signals are at approximately two frequencies (1.5 kHz and 10.0 kHz, respectively) that are also affected by resonance in the human ear. Consequently, the ear will receive signals 306A and 306C, respectively, that are also affected by resonance. Specifically, resonance can cause input signals 302A and 302C to be reduced. If input signal 302A has a specific amplitude, signal 306A may have an amplitude that is several times smaller. For example, the amplitude of signal 306A may be approximately half the amplitude of input signal 302A (e.g., reduced by approximately -6 dB). If input signal 302C has a specific amplitude, signal 306C may have an amplitude that is several times smaller. For example, the amplitude of signal 306C may be approximately half the amplitude of input signal 302C (e.g., reduced by approximately -6 dB). Transients at approximately 1.5 and / or 10.0 kHz may become more localized in time (e.g., sharper in time). For example, the resonance at 3.0 kHz can act as a differential filter by eliminating surrounding frequencies, thereby enhancing transients at those frequencies but attenuating energy in sustained waves. This can allow for more quantization, but leaves less room for arranging transients. For example, signal 306A can include waveform 308A that is several times shorter (e.g., three times shorter) than waveform 304A. As another example, signal 306C can include waveform 308C that is several times shorter (e.g., three times shorter) than waveform 304C. Thus, the 1.5 and 10.0 kHz frequencies can each be associated with energy differentiation in human hearing.

[0047] Applying aspects of the present subject matter can facilitate improved audio processing. For example, an audio compressor (e.g., as Figure 1 ) and / or a component for evaluating audio signal similarity (e.g., Figure 7 The audio analyzer 700 in FIG. 1 may achieve higher amplitude sensitivity and / or higher time sensitivity. The present subject matter may be practiced via instructions (e.g., a computer program) stored in a computer program product and executable by at least one processor. In some implementations, executing operations according to the instructions may result in increased amplitude sensitivity at a first frequency (e.g., at approximately 3.0 kHz). For example, the increased amplitude sensitivity may be attributed to a larger amplitude-independent window size (e.g., a window size twice as large) being used at the first frequency compared to another frequency (e.g., a frequency below approximately 1 kHz). In some implementations, executing operations according to the instructions may result in increased time sensitivity at a second frequency (e.g., at approximately 1.5 and / or approximately 10 kHz). For example, the increased time sensitivity may be attributed to a smaller amplitude-independent window size (e.g., a window size twice as small) being used at the second frequency compared to another frequency (e.g., a frequency below approximately 1 kHz).

[0048] Figure 4 An example of an audio encoder 400 is shown. The audio encoder 400 can be used with one or more examples described elsewhere herein. The audio encoder 400 is configured to receive an input 402 (e.g., one or more signals corresponding to an acoustic energy flow), process the input 402 signals, and generate an output 404 (e.g., one or more encoded signals). In some implementations, the audio encoder 400 can be used with high-quality audio (e.g., to provide a high-quality hifi sound system). For example, the audio encoder 400 can support lossless (e.g., the original signal can be perfectly reconstructed using the encoded signal) or near-lossless (e.g., the original signal can be almost perfectly reconstructed using the encoded signal) compression. The audio encoder 400 can be based on a reference signal. Figure 9 The described one or more examples implement the audio encoder 400 .

[0049] The audio encoder 400 may include one or more transforms 406. The transform 406 may convert the audio signal from the time domain to the frequency domain. The transform 406 may be performed over one or more time ranges, sometimes referred to as windows for the transform 406. When a sound evolves slowly, it can be said that the larger the window (e.g., the larger the number of milliseconds (ms) of the transform), the more compressed that portion of the signal can be. For sounds, it is sometimes considered that they evolve relatively slowly in a relevant reference frame. For example, with speech, an audio signal is generated by a column of vibrating air, such that at some given time, the air will vibrate at least substantially as it did 20 milliseconds earlier. In this context, an integral transform may be used to obtain a predictive signature of the vibration. Any frequency-dependent transform may be used, including but not limited to a Fourier transform or a cosine transform. In some implementations, a discrete variant of the transform may be used. For example, a discrete Fourier transform (DFT) may be implemented as a fast Fourier transform (FFT). As another example, a discrete cosine transform (DCT) may be used.

[0050] The audio encoder 400 includes a mapping 408 between a window size and a frequency. The mapping 408 can be based on the resonance phenomenon in the human ear. In some implementations, the mapping 408 can associate a first window size with a frequency that is associated with energy integration in human hearing. For example, the frequency can be approximately 3.0kHz (e.g., with a window size of approximately 18-30ms, such as approximately 24ms). In some implementations, the mapping 408 can associate a second window size with a frequency that is associated with energy differentiation in human hearing. For example, the frequency can be approximately 1.5kHz and / or approximately 10.0kHz (e.g., with a window size of approximately 3-9ms, such as approximately 6ms). In some implementations, the mapping 408 can associate a third window size with a frequency that is not associated with any specific acoustic phenomenon in human hearing (e.g., not associated with any resonance). For example, the frequency can be lower than approximately 1.0kHz and / or higher than approximately 10.0kHz (e.g., with a window size of approximately 6-18ms, such as approximately 12ms). Mapping 408 can implement the association between window size (e.g., in terms of size such as ms) and frequency (e.g., in terms of one or more frequency bands) in any of a number of different ways. For example, mapping 408 can include a lookup table to be used with one or more transforms 406. As another example, mapping 408 can be integrated into one or more transforms 406 so as to be automatically applied to the transform.

[0051] Encoder 400 is an example of an apparatus that can perform methods related to improved encoding. The method can include receiving a first signal corresponding to a first acoustic energy stream (e.g., Figure 3 The method may include performing a transform (e.g., FFT or DCT) on the received first signal. The transform may use at least a first amplitude-independent window size (e.g., approximately 24 ms) at a first frequency (e.g., approximately 3 kHz) and a second amplitude-independent window size (e.g., approximately 6 ms) at a second frequency (e.g., approximately 1.5 kHz and / or approximately 10 kHz). The second amplitude-independent window size may improve the time response at the second frequency (e.g., Figure 3Waveforms 308A and / or 308C in the input signal 302A and 302C may represent transients that are relatively easier to detect. For example, the second amplitude-independent window size may improve the time response by being shorter than the window size used for most of the bandwidth, thereby applying the transform to a shorter span of the audio signal at a time. Due to resonance phenomena associated with the first frequency, the second frequency may experience a reduction in amplitude (e.g., signal 306A or 306C may have a reduced amplitude relative to input signal 302A or 302C, respectively). The method may include storing a first encoded signal (e.g., output 404), the first encoded signal being based on applying the transform to the received first signal.

[0052] Figure 5An example of a window size is shown. The window size is shown relative to an axis 500 representing frequency. For example, the frequencies of axis 500 are the individual frequencies included in the audio signal (e.g., separated by a filter bank). Frequency 502 may be associated with a resonance phenomenon (e.g., in the human ear). For example, resonance may amplify the signal at frequency 502 and attenuate signals at one or more other frequencies. Here, frequency 504 and frequency 506 are indicated. Frequencies 504 and / or 506 may be associated with a resonance phenomenon (e.g., in the human ear). For example, resonance may attenuate signals at frequencies 504 and / or 506. In the transformation, different window sizes may be used for one or more of frequencies 502, 504, or 506, and the window size may be independent of the specific amplitude at any frequency (e.g., not dependent on whether a transient has been detected at that frequency (frequency band)). In some implementations, the window size associated with frequency 502 may be used for all frequencies of the signal except frequencies 504 and / or 506 (e.g., for one or more frequency bands including frequencies 504 and / or 506). Frequencies 504 and 506 may use the same or different window sizes. The window size for frequency 502 may be larger than the window sizes for frequencies 504 and / or 506. Making the window size for frequency 502 larger than the window sizes for frequencies 504 and / or 506 can provide the following advantages: more efficiently processing portions of the audio signal where the time response increases are relatively insignificant (e.g., so that the transform is applied to a larger span of the audio signal each time). For example, a 24 ms window is larger than a 6 ms window. In some implementations, the window size for frequency 502 may be an integer multiple larger than the window size for frequencies 504 and / or 506. For example, a window size of approximately 24 ms is approximately four times the window size of approximately 6 ms. Frequency 508 and frequency 510 are labeled. In some implementations, frequencies 508 and / or 510 are not associated with any acoustic phenomena of the human ear (e.g., frequencies 508 and / or 510 are not amplified or attenuated by resonance at 3 kHz). For example, frequency 508 may be lower than frequency 504 (e.g., at approximately 1 kHz or lower). As another example, frequency 510 may be higher than frequency 506. Frequencies 508 and / or 510 may use a window size that is different from one or more other frequency sizes. In some implementations, the window size for frequencies 508 and / or 510 is smaller than the window size for frequency 502. In some implementations, the window size for frequencies 508 and / or 510 is approximately half the window size for frequency 502. Making the window size for frequencies 508 and / or 510 approximately half the window size for frequency 502 can provide the following advantages: obtaining higher quality encoding in portions of the audio signal where resonance effects do not occur or are relatively insignificant (e.g., so that the transform is applied to a smaller span of the audio signal each time). For example, a window size of approximately 12 ms is smaller than, and approximately half of, a window size of approximately 24 ms.In some implementations, the window size for frequencies 508 and / or 510 can be larger than the window size for frequencies 504 and / or 506. In some implementations, the window size for frequencies 508 and / or 510 can be approximately twice the window size for frequencies 504 and / or 506. Making the window size for frequencies 508 and / or 510 approximately twice the window size for frequencies 504 and / or 506 can provide the following advantages: more efficient encoding in portions of the audio signal where the increase in time response is relatively insignificant (e.g., so that the transform is applied to a larger span of the audio signal at a time). For example, a window size of 12 ms is larger than a window size of approximately 6 ms and is approximately twice as large. Frequencies 504 and 506 can be located on opposite sides of frequency 502. For example, one of frequencies 504 and 506 can be lower than frequency 502, while the other of frequencies 504 and 506 can be lower than frequency 502. That is, the positions here can be defined by frequency. For example, resonance at frequency 502 may cause attenuation at one or more higher frequencies (eg, at frequency 506 ) and at one or more lower frequencies (eg, at frequency 504 ).

[0053] Encoder (e.g. Figure 4The audio encoder 400 in FIG. 4 may be included in the codec. In some implementations, the codec may calculate multiples of the window size. When storing frequencies in different frequency bands, frequencies of approximately 1.5 kHz and approximately 10 kHz may be stored. In some implementations, data may be stored more frequently (e.g., as integer multiples) for these frequencies compared to the resonant frequency (e.g., approximately 3 kHz). Storing data more frequently may provide the following advantages: the window size is shorter, thereby improving the time response, so that the transform is applied to the audio signal of a shorter span each time. For example, data of frequencies of approximately 1.5 kHz and approximately 10 kHz may be stored more frequently because their window sizes are shorter in duration than the window size of the resonant frequency (e.g., approximately 3 kHz), so they have outputs for a given time period. For example, if the 3 kHz window size is four times the 1.5 kHz window size and is approximately the 10 kHz window size, one window may have four outputs, while the other window may have one output, and the output values ​​of each window may be different from each other. In some implementations, relatively lower precision may be used for frequencies of approximately 1.5 kHz and / or approximately 10 kHz. For example, one or two bits may be omitted so that time data is preserved and there is a greater degree of quantization. Quantization may be advantageous in reducing the amount of data stored, thereby requiring fewer system resources. In some implementations, relatively higher precision may be used for frequencies of approximately 3 kHz. For example, one or two bits may be added so that there is more data to capture finer amplitude variations in the region. That is, it can be said that a transformation applied at a resonant frequency (e.g., 3 kHz) produces a first result, and it can be said that a transformation applied at a decaying frequency (e.g., 1.5 and / or 10 kHz) produces a second result. The first result is stored less frequently (e.g., every 24 milliseconds) than the second result (e.g., every 6 milliseconds), including, but not limited to, the second result being stored approximately four times more frequently than the first output.

[0054] Figure 6Examples of decoding are schematically illustrated. These examples of decoding can be used in conjunction with one or more other examples described elsewhere herein. Decoding can be applied to an encoded signal to convert it into another form (e.g., an audio signal). The transforms of different sizes implied by the encoding process can be manipulated and accumulated during decoding. In some implementations, different frequency bands can be represented by different window lengths. That is, when decoding sound, decoding can be performed from each of a variety of transform sizes. In some implementations, to extract a sample, three transforms (e.g., 6ms, 12ms, and 24ms transforms) can be performed. These can be accumulated, and the decoder can transmit a 6-millisecond time period. Transformations 600-1, 600-2, 600-3, and 600-4 are shown here. For example, each of transformations 600-1 to 600-4 corresponds to applying a transform with a specific window size (e.g., 6ms) to one or more frequencies. Transformations 602-1 and 602-2 are shown here. For example, each of transforms 602-1 and 602-2 corresponds to applying a transform with a specific window size (e.g., 12 ms) to one or more frequencies. Transform 604 is shown here. For example, transform 604 corresponds to applying a transform with a specific window size (e.g., 24 ms) to one or more frequencies. Transform 606 schematically represents another application of a transform to an audio signal (e.g., with a smaller or larger window size).

[0055] The following is an example of decoding. Transformations 600-1, 602-1, and 604 may be performed, where transformations 602-1 and 604 may be stored (e.g., via Figure 1 The resonance enhancement decoder 110 in the memory is stored in memory). The transformations 600-1, 602-1, and 604 can then be accumulated and used to output the sound for a portion of time (e.g., 6 ms). Thereafter, transformation 600-2 can be performed. By retrieving the transformations 602-1 and 604 from storage, the transformations 600-2, 602-1, and 604 can be accumulated and used to output the sound for a portion of time (e.g., 6 ms). Then, transformations 600-3 and 602-2 can be performed, where the transformation 602-2 can be stored. Then, the transformations 600-3, 602-2, and 604 can be accumulated and used to output the sound for a portion of time (e.g., 6 ms). Finally, transformation 600-4 can be performed. By retrieving the transformations 602-2 and 604 from storage, the transformations 600-4, 602-2, and 604 can be accumulated and used to output the sound for a portion of time (e.g., 6 ms).

[0056] Figure 7An example of an audio analyzer 700 is shown. The audio analyzer 700 can be used with one or more of the other examples described elsewhere herein. Figure 9 One or more examples described herein implement an audio analyzer 700. In some implementations, the audio analyzer 700 can be used to determine (e.g., model) differences between audio files. Here, audio files 702 and 704 are shown as being input into the audio analyzer 700. Each of the audio files 702 and 704 can be generated according to the present subject matter. For example, the audio encoder 400 ( Figure 4 ) can generate audio files 702 and 704. The audio analyzer 700 includes a difference determination circuit 706. In some implementations, the difference determination circuit 706 can perform an evaluation of the audio files 702 and 704 to determine whether they are the same or different, or what differences there are between them. To name a few examples, the difference determination circuit 706 can perform such an evaluation as part of speech recognition, blind source separation, directionality determination, security control, authentication, music selection, and / or fraud detection. The difference determination circuit 706 can apply each of the audio files 702 and 704 to a human hearing model 708. In some implementations, the model 708 is a software-based representation of how the human ear works (e.g., a psychoacoustic model). For example, the model 708 can specify that sounds at frequencies around 3 kHz are amplified and undergo energy integration (e.g., smearing in time), and sounds at frequencies around 1.5 kHz and around 10 kHz are attenuated and undergo energy differentiation (e.g., transient enhancement). By applying the audio files 702 and 704 to the human hearing model 708 via the difference determination circuit 706, the audio encoder 400 can determine differences, if any, between the audio files 702 and 704. The difference determination circuit 706 can include a user interface 710 to output one or more results of evaluating the audio files 702 and 704. In some implementations, the user interface 710 indicates differences, if any, between the audio files 702 and 704. The user interface 710 can generate an output 712, such as a binary evaluation (e.g., "same" or "not identical") or a quantitative evaluation based on a similarity criterion (e.g., "95% similar"), to name a few examples. The output 712 can be generated for a human user or for another component that relies on the evaluation of the audio analyzer 700.

[0057] Figure 8 An example of method 800 is shown. Method 800 can be used with one or more other examples described elsewhere herein. Method 800 can be performed by Figure 9The method 800 is a computer-implemented method executed by the computing device 900 in FIG. The method 800 may include more or fewer operations than those shown. Unless otherwise indicated, two or more operations of the method 800 may be performed in a different order.

[0058] At 802, a signal may be received. The signal may be an audio signal corresponding to an energy flow. For example, the resonance enhancement encoder 106 may receive a signal ( Figure 1 ).

[0059] At 804, a transform may be applied to the received signal. In some implementations, the transform uses a window size that is independent of amplitude. For example, a DCT or FFT may be applied to any of the input signals 302A-C regardless of the amplitude of the signal. Different window sizes may be applied at different frequencies.

[0060] At 806, the encoded signal may be stored. For example, the resonance enhanced encoder 106 ( Figure 1 ) can store encoded signals.

[0061] Figure 9 Illustrated is an example architecture of a computing device 900 that can be used to implement aspects of the present disclosure, including any of the systems, apparatuses, and / or techniques described herein, or any other systems, apparatuses, and / or techniques that may be described in various possible implementations.

[0062] Figure 9 The computing device illustrated in can be used to execute the operating systems, applications, and / or software modules (including software engines) described herein.

[0063] In some implementations, computing device 900 includes at least one processing device 902 (e.g., a processor), such as a central processing unit (CPU). Various processing devices are available from various manufacturers, such as Intel or Advanced Micro Devices. In this example, computing device 900 also includes system memory 904 and a system bus 906 that couples various system components, including system memory 904, to processing device 902. System bus 906 is one of several types of bus structures that may be used, including, but not limited to, a memory bus or memory controller; a peripheral bus; and a local bus using any of a variety of bus architectures.

[0064] Examples of computing devices that may be implemented using computing device 900 include desktop computers, laptop computers, tablet computers, mobile computing devices (such as smart phones, touchpad mobile digital devices, or other mobile devices), or other devices configured to process digital instructions.

[0065] System memory 904 includes read-only memory 908 and random access memory 910. Basic input / output system 912, containing the basic routines for transferring information within computing device 900, such as during startup, may be stored in read-only memory 908.

[0066] In some implementations, the computing device 900 also includes a secondary storage device 914, such as a hard drive, for storing digital data. The secondary storage device 914 is connected to the system bus 906 via a secondary storage interface 916. The secondary storage device 914 and its associated computer-readable media provide non-volatile and non-transitory storage of computer-readable instructions (including applications and program modules), data structures, and other data for the computing device 900.

[0067] Although the example environment described herein uses a hard drive as a secondary storage device, other types of computer-readable storage media may be used in other implementations. Examples of these other types of computer-readable storage media include magnetic cassettes, flash memory cards, digital video disks, Bernoulli cartridges, compact disc read-only memories, digital versatile disk read-only memories, random access memories, or read-only memories. Some implementations include non-transitory media. For example, a computer program product may be tangibly embodied in a non-transitory storage medium. Additionally, such computer-readable storage media may include local storage or cloud-based storage.

[0068] A number of program modules may be stored in the secondary storage device 914 and / or the system memory 904, including an operating system 918, one or more application programs 920, other program modules 922 (such as the software engines described herein), and program data 924. The computing device 900 may use any suitable operating system, such as Microsoft Windows TM , Google Chrome TM OS, Apple OS, Unix or Linux and their variants and any other operating system suitable for computing devices. Other examples may include Microsoft, Google or Apple operating systems, or any other suitable operating system used in tablet computing devices.

[0069] In some implementations, a user provides input to the computing device 900 through one or more input devices 926. Examples of input devices 926 include a keyboard 928, a mouse 930, a microphone 932 (e.g., for voice and / or other audio input), a touch sensor 934 (such as a touchpad or touch-sensitive display), and a gesture sensor 935 (e.g., for gesture input). In some implementations, the input devices 926 provide detection based on presence, proximity, and / or motion. In some implementations, a user can walk into their home, which can trigger input into the processing device. For example, the input device 926 can then facilitate an automated experience for the user. Other implementations include other input devices 926. The input devices can be connected to the processing device 902 through an input / output interface 936 coupled to the system bus 906. These input devices 926 can be connected through any number of input / output interfaces, such as a parallel port, a serial port, a game port, or a universal serial bus. Wireless communication between the input devices 926 and the input / output interface 936 is also possible, and in some possible implementations includes infrared, Wireless technologies, 802.11a / b / g / n, cellular, ultra-wideband (UWB), ZigBee or other radio frequency communication systems, to name a few.

[0070] In this example implementation, a display device 938, such as a monitor, liquid crystal display device, projector, or touch-sensitive display device, is also connected to the system bus 906 via an interface, such as a video adapter 940. In addition to the display device 938, the computing device 900 may include various other peripheral devices (not shown), such as speakers or a printer.

[0071] The computing device 900 can be connected to one or more networks via a network interface 942. The network interface 942 can provide wired and / or wireless communications. In some implementations, the network interface 942 can include one or more antennas for sending and / or receiving wireless signals. When used in a local area network environment or a wide area network environment (such as the Internet), the network interface 942 can include an Ethernet interface. Other possible implementations use other communication devices. For example, some implementations of the computing device 900 include a modem for communicating over a network.

[0072] The computing device 900 may include at least some form of computer-readable media. Computer-readable media includes any available media that can be accessed by the computing device 900. By way of example, computer-readable media includes computer-readable storage media and computer-readable communication media.

[0073] Computer-readable storage media includes volatile and nonvolatile, removable and non-removable media implemented in any device configured to store information such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media includes, but is not limited to, random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technology, compact disc read-only memory, digital versatile disks or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 900.

[0074] Computer-readable communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" refers to a signal that has one or more of its characteristics set or changed in such a manner as to encode information as a signal. By way of example, computer-readable communication media include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency, infrared, and other wireless media. Any combination of the above is also included within the scope of computer-readable media.

[0075] Figure 9 The computing device illustrated in the figure is also an example of a programmable electronic device, which may include one or more such computing devices, and when multiple computing devices are included, such computing devices may be coupled together with a suitable data communication network to jointly perform various functions, methods, or operations disclosed herein.

[0076] A number of implementations have been described, however, it will be understood that various modifications can be made without departing from the spirit and scope of the invention.

[0077] Additionally, the logic flows depicted in the accompanying drawings do not require the particular order shown or sequential order to achieve the desired results. Furthermore, other steps may be provided or eliminated from the described flows, and other components may be added to or removed from the described systems. Accordingly, other implementations are within the scope of the appended claims.

Claims

1. A computer-implemented method for signal processing, comprising: receiving a first signal corresponding to a first acoustic energy stream; Applying a transformation to the received first signal comprises: using a first window during a sampling duration of the transform, the first window having a first amplitude-independent window size based on a first resonance phenomenon at a first frequency, and the first amplitude-independent window size is not based on acoustic energy of the first signal at the first frequency, and A second window is used during the sampling duration of the transform, the second window having a second amplitude-independent window size based on a second resonance phenomenon at a second frequency, and the second amplitude-independent window size is not based on the acoustic energy of the first signal at the second frequency, the second window improving the time response at the second frequency, wherein the second frequency undergoes a reduction in amplitude due to the first resonance phenomenon associated with the first frequency; and A first coded signal is stored, the first coded signal being based on applying the transform to the received first signal. 2 . The computer-implemented method of claim 1 , further comprising mapping the first amplitude-independent window size to the first frequency based on the first frequency being associated with energy integration in human hearing. 3 . The computer-implemented method of claim 1 , further comprising mapping the second amplitude-independent window size to the second frequency based on the second frequency being associated with energy differentiation in human hearing.

4. The computer-implemented method of claim 1 , wherein: The first amplitude-independent window size is used for all frequencies of the received first signal except for a frequency band at the second frequency.

5. The computer-implemented method of claim 1 , wherein: The first amplitude-independent window size is larger than the second amplitude-independent window size.

6. The computer-implemented method of claim 5, wherein: The first amplitude-independent window size is an integer multiple greater than the second amplitude-independent window size.

7. The computer-implemented method of claim 5, wherein: The first amplitude-independent window size is four times larger than the second amplitude-independent window size.

8. The computer-implemented method of claim 1 , further comprising using a third window when applying the transform to the first received signal, the third window being used at a third frequency not associated with the resonance phenomenon, the third window having a third amplitude-independent window size, the third amplitude-independent window size being different from the first amplitude-independent window size and the second amplitude-independent window size and not based on acoustic energy of the first signal at the third frequency.

9. The computer-implemented method of claim 8, wherein: The third amplitude-independent window size is smaller than the first amplitude-independent window size.

10. The computer-implemented method of claim 8, wherein: The third amplitude-independent window size is half as large as the first amplitude-independent window size.

11. The computer-implemented method of claim 8, wherein: The third amplitude-independent window size is larger than the second amplitude-independent window size.

12. The computer-implemented method of claim 11, wherein: The third amplitude-independent window size is twice as large as the second amplitude-independent window size.

13. The computer-implemented method of claim 11 , wherein: The third amplitude-independent window size is smaller than the first amplitude-independent window size.

14. The computer-implemented method of claim 1 , wherein: Applying the transform using the first window generates a first result, wherein applying the transform using the second window generates a second result, the method further comprising storing the second result more frequently than storing the first result.

15. The computer-implemented method of claim 14, further comprising storing the second result with a lower precision than the first result.

16. The computer-implemented method of claim 1 , further comprising using a third window when applying the transform at a third frequency, the third window improving the time response at the third frequency, the third window having a third amplitude-independent window size that is not based on acoustic energy of the first signal at the third frequency, the third frequency experiencing amplitude reduction due to the resonance phenomenon associated with the first frequency.

17. The computer-implemented method of claim 16, wherein: The second frequency and the third frequency are located on opposite sides of the first frequency.

18. The computer-implemented method of claim 16, wherein: The third amplitude-independent window size is equal to the second amplitude-independent window size.

19. The computer-implemented method of claim 16, wherein: The second amplitude-independent window size and the third amplitude-independent window size are smaller than the first amplitude-independent window size.

20. The computer-implemented method of any one of claims 1-19, wherein: The first audio file includes the first encoded signal, the method further comprising: receiving a second signal corresponding to a second acoustic energy stream; applying a transform to the received second signal using at least the first window at the first frequency and the second window at the second frequency; storing a second encoded signal, the second encoded signal being based on applying the transform to the received second signal, wherein a second audio file comprises the second encoded signal; and Differences between the first audio file and the second audio file are determined.

21. The computer-implemented method of claim 20, wherein: Determining the difference includes playing the first audio file and the second audio file into a human hearing model, the model including the first resonance phenomenon and the second resonance phenomenon.

22. The computer-implemented method of claim 1, wherein: Improving the temporal response comprises increasing the temporal resolution of the first encoded signal by including less audio content when applying the transform.

23. A computer program product tangibly embodied in a non-transitory storage medium, the computer program product comprising instructions that, when executed by a processor, cause the processor to perform operations comprising: receiving a first signal corresponding to a first acoustic energy stream; Applying a transformation to the received first signal comprises: using a first window during a sampling duration of the transform, the first window having a first amplitude-independent window size based on a first resonance phenomenon at a first frequency, and the first amplitude-independent window size is not based on acoustic energy of the first signal at the first frequency, and A second window is used during the sampling duration of the transform, the second window having a second amplitude-independent window size based on a second resonance phenomenon at a second frequency, and the second amplitude-independent window size is not based on the acoustic energy of the first signal at the second frequency, the second window improving the time response at the second frequency, wherein the second frequency undergoes a reduction in amplitude due to the first resonance phenomenon associated with the first frequency; and A first coded signal is stored, the first coded signal being based on applying the transform to the received first signal.

24. The computer program product of claim 23, wherein: Executing the operation according to the instruction increases the amplitude sensitivity at the first frequency.

25. The computer program product of claim 24, wherein: The increase in amplitude sensitivity is due to the first amplitude-independent window size being larger than the second amplitude-independent window size.

26. A computer program product according to claim 23, 24 or 25, wherein: Executing the operation according to the instruction increases the time sensitivity at the second frequency.

27. The computer program product of claim 26, wherein: The improvement in time sensitivity is due to the second amplitude-independent window size being smaller than the first amplitude-independent window size.