Method and apparatus for extracting a pitch-independent timbre attribute from a media signal
By extracting pitch-independent timbre attributes through logarithmic spectrum analysis, the method addresses the challenge of resource-intensive pitch-dependent timbre classification, enabling precise audio source identification and improved media exposure analysis.
Patent Information
- Application Number
- JP2025146214
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-01-03
- Filing Date
- 2025-09-03
- Publication Date
- 2026-01-14
AI Technical Summary
Existing audio processing technologies struggle to extract timbre attributes independently of pitch, requiring large databases of pitch-dependent timbres and consuming significant resources for classification and identification.
The method involves extracting pitch-independent timbre attributes using a logarithmic spectrum analysis, allowing for classification and identification of audio signals based on timbre alone, reducing resource requirements and enhancing accuracy.
This approach enables more precise classification and identification of audio sources, adjusts sound settings for better user experience, and provides detailed media exposure information by analyzing timbre independently of pitch.
Smart Images

Figure 2026004309000001_ABST
Abstract
Description
Field of Disclosure
[0001] This disclosure relates generally to audio processing, and more particularly to methods and apparatus for extracting pitch-independent timbre attributes from media signals.
[0002]
[0002] Timbre (e.g., timbral attribute / timbral attribute) is a quality / characteristic of sound that is unrelated to the pitch or loudness of the sound. Timbre is what makes two different sounds sound different from each other, even if they have the same pitch and loudness. For example, a guitar and a flute playing the same note at the same amplitude will sound different because the guitar and flute have different timbres. Timbre corresponds to the frequency and time envelope of a sound event (e.g., the distribution of energy along time and frequency). The characteristics of sound that correspond to the perception of timbre include the spectrum and the envelope. [Brief explanation of the drawings]
[0003] [Figure 1] FIG. 1 illustrates an exemplary instrument for extracting pitch-independent timbral attributes from a media signal.
[0004] [Figure 2] FIG. 2 is a block diagram of the exemplary speech analyzer and exemplary speech identifyr of FIG.
[0005] [Figure 3]
[0005] Figure 3 is a flowchart representing example machine-readable instructions that may be executed to implement the example speech analyzer of Figures 1 and 2 to extract pitch-independent timbre attributes from a media signal and / or to extract timbre-independent pitch from a media signal.
[0006] [Figure 4]
[0006] Figure 4 is a flowchart representing example machine-readable instructions that may be executed to implement the example audio specifier of Figures 1 and 2 to characterize audio and / or identify media based on the timbre logarithmic spectrum of silent pitches.
[0007] [Figure 5] FIG. 5 illustrates an exemplary speech signal, an exemplary pitch of the speech signal, and an exemplary timbre of the speech signal that may be determined using the exemplary speech analyzer of FIGS.
[0008] [Figure 6] FIG. 6 is a block diagram of a processor platform configured to execute the example machine-readable instructions of FIG. 3 to control the example speech analyzer of FIGS.
[0009] [Figure 7] FIG. 7 is a block diagram of a processor platform configured to execute the example machine-readable instructions of FIG. 4 to control the example voice identifiers of FIGS.
[0010]
[0010] The figures are not to scale. Wherever possible, the same reference numbers are used throughout the drawings and the accompanying written description to refer to the same or like parts. DETAILED DESCRIPTION
[0011]
[0011] An audio meter is a device that captures an audio signal (e.g., directly or indirectly) and processes the audio signal. For example, when a panelist contracts for exposure to media that is monitored by an audience measurement entity, the audience measurement entity may dispatch a technician to the panelist's home to install an instrument (e.g., a media monitor) that can collect media exposure data from media output device(s) (e.g., television receiver, radio, computer, etc.). In another example, the instrument may respond to instructions executed by a processor, for example, on a smartphone, to process the received audio and / or video data to identify characteristics of the media.
[0012]
[0012] Generally, an instrument includes or is otherwise connected to an interface (e.g., a microphone and / or magnetic coupling device for collecting ambient sound) to receive media signals directly or indirectly from a media source. For example, when a media output device is "on," a microphone can receive audio signals transmitted from the media output device. The instrument can process the received audio signals to identify characteristics of the audio that can be used to characterize and / or identify the audio or audio source. When the instrument responds to commands operating within and / or in conjunction with the media output device to receive audio and / or video signals to be output from the media output device, the instrument can process / analyze the incoming audio and / or video signals to directly identify data associated with the signals. For example, an instrument can operate within a set-top box, receiver, mobile phone, etc. to receive and process incoming audio / video data before, during, or after it is output from the media output device.
[0013] In some examples, the audio measurement device / instructions classify and / or identify audio and / or audio sources using various characteristics of the audio. Such characteristics may include the energy of the media signal, the energy of each frequency band of the media signal, the discrete cosine transform (DCT) coefficients of the media signal, etc. In examples disclosed herein, media is classified and / or identified based on the timbre of the audio corresponding to the media signal.
[0014] Timbre (e.g., timbral attribute / timbral attribute) refers to a quality / characteristic of sound that is unrelated to the pitch or loudness of the sound. For example, a guitar and a flute playing the same note at the same amplitude sound different because the guitar and the flute have different timbres. Timbre corresponds to the frequency and time envelope (e.g., the distribution of energy along time and frequency) of a sound event. Traditionally, timbre has been characterized by various features. However, timbre has never been extracted from sound independently of other aspects of the sound (e.g., pitch). Therefore, identifying media based on pitch-dependent timbre measurements requires a large database of reference pitch-dependent timbres that correspond to timbres by category and pitch. In the examples disclosed herein, a pitch-independent timbre logarithmic spectrum is extracted from measured sound that is independent of pitch, thus reducing the resources required to classify and / or identify media based on timbre.
[0015] As described above, the extracted pitch-independent timbre can be used to classify and / or identify media and / or as part of a signature algorithm. For example, the extracted pitch-independent timbre attributes (e.g., logarithmic spectrum) can be used to determine that a measured sound (e.g., a sound sample) corresponds to a violin, regardless of the notes being played by the violin. In some examples, the characteristic sound can be used to adjust sound settings of a media output device to provide a better audio experience to a user. For example, some sound equalizer settings can be better suited to the sounds of a particular instrument and / or genre. Thus, in examples disclosed herein, the sound equalizer settings of the media output device can be adjusted based on the identified instrument / genre corresponding to the extracted timbre. In another example, the extracted pitch-independent timbre can be used to identify media output from a media presentation device (e.g., a television receiver, a computer, a radio, a smartphone, a tablet, etc.) by comparing the extracted pitch-independent timbre attributes to reference timbre attributes in a database. In this manner, the extracted timbre and / or pitch can be used to provide audience measurement entities with more detailed media exposure information than conventional techniques that only consider the pitch of the received audio.
[0016]
[0016] Figure 1 shows an example speech analyzer 100 for extracting pitch-independent timbre attributes from a media signal. Figure 1 includes the example speech analyzer 100, an example media output device 102, example speakers 104a, 104b, an example media signal 106, and an example speech identifyr 108.
[0017] The example speech analyzer 100 of FIG. 1 receives a media signal from a device (e.g., the example media output device 102 and / or the example speakers 104a, 104b) and processes the media signal to identify pitch-independent timbre attributes (e.g., log-spectrum) and timbre-independent pitch attributes. In some examples, the speech analyzer 100 may include or be otherwise connected to a microphone to receive the example media signal 106 by detecting ambient sound. In such examples, the speech analyzer 100 may be implemented in an instrument or other computing device (e.g., a computer, tablet, smartphone, smartwatch, etc.) that utilizes a microphone. In some examples, the speech analyzer 100 includes an interface for receiving the example media signal 106 directly (e.g., via a wired or wireless connection) from the example media output device 102 and / or a media presentation device that presents media to the media output device 102. For example, the audio analyzer 100 may receive the media signal 106 directly from a set-top box, a mobile phone, a gaming device, an audio receiver, a DVD player, a Blu-ray player, a tablet, and / or any other device that provides media to be output from the media output device 102 and / or the exemplary speakers 104 a, 104 b. As further described below in conjunction with FIG. 2, the exemplary audio analyzer 100 extracts pitch-independent timbre attributes and / or timbre-independent pitch attributes from the media signal 106. If the media signal 106 is a video signal that includes an audio component, the exemplary audio analyzer 100 extracts the audio component from the media signal 106 before extracting pitch and / or timbre.
[0018] The exemplary media output device 102 of FIG. 1 is a device that outputs media. While the exemplary media output device 102 of FIG. 1 is illustrated as a television receiver, the exemplary media output device 102 may be a radio, an MP3 player, a video game console, a stereo system, a mobile device, a tablet, a computing device, a tablet, a laptop, a projector, a DVD player, a set-top box, an over-the-top device, and / or any device capable of outputting media (e.g., video and / or audio). The exemplary media output device may include a speaker 104a and / or may be coupled to or otherwise connected to a portable speaker 104b via a wired or wireless connection. The exemplary speakers 104a, 104b output the audio portion of the media output from the exemplary media output device. In the example shown in FIG. 1, the media signal 106 represents the audio output from the exemplary speakers 104a, 104b. Additionally or alternatively, the exemplary media signal 106 may be an audio and / or video signal transmitted to and output from the exemplary media output device 102 and / or the exemplary speakers 104a, 104b. For example, the exemplary media signal 106 may be a signal from a game console transmitted to the exemplary media output device 102 and / or the exemplary speakers 104a, 104b for outputting audio and video from a video game. The exemplary audio analyzer 100 may receive the media signal 106 directly from the media presentation device (e.g., a game console) and / or from ambient audio. In this manner, the audio analyzer 100 may classify and / or identify audio from the media signal even when the speakers 104a, 104b are off, inactive, or have their volume turned down.
[0019] The example voice identifyr 108 of FIG. 1 characterizes speech and / or identifies media based on received pitch-independent timbre attribute measurements from the example voice analyzer 100. For example, the voice identifyr 108 may include a database of reference pitch-independent timbre attributes corresponding to classification and / or identification. In this manner, the example voice identifyr 108 may compare the received pitch-independent timbre attribute(s) to the reference pitch-independent attributes to determine a match. If the example voice identifyr 108 determines a match, the example voice identifyr 108 classifies the speech and / or identifies media for information corresponding to the matched reference timbre attributes. For example, if the received timbre attributes match the reference attributes corresponding to a trumpet, the example voice identifyr 108 classifies the speech corresponding to the received timbre attributes as speech from a trumpet. In such an example, if the audio analyzer 100 is part of a mobile phone, the exemplary audio analyzer 100 can receive audio signals of a trumpet playing a song (e.g., via an interface that receives audio / video signals or via a microphone on the mobile phone that receives audio signals). In this manner, the audio identifyr 108 can identify the instrument corresponding to the received audio as a trumpet and identify it as a trumpet to a user (e.g., using the user interface of the mobile phone). In another example, if the received timbre attributes match reference attributes corresponding to a particular video game, the exemplary audio identifyr 108 can identify the audio corresponding to the received timbre attributes as being from the particular video game. The exemplary audio identifyr 108 can generate a report identifying the audio. In this manner, an audience measurement entity can credit exposure to the video game based on the report. In some examples, the audio identifyr 108 receives the timbre directly from the audio analyzer 100 (e.g., both the audio analyzer 100 and the audio identifyr 108 are located on the same device). In some examples, the voice identifier 108 is located elsewhere and receives the timbre from the example voice analyzer 100 via wireless communication.In some examples, the voice identifyr 108 sends instructions to the exemplary audio media output device 102 and / or the exemplary voice analyzer 100 (e.g., when the exemplary voice analyzer 100 is implemented in the exemplary media output device 102) to adjust voice equalizer settings based on the voice classification. For example, if the voice identifyr 108 classifies the voice being output from the media output device 102 as being from a trumpet, the exemplary voice identifyr 108 can send instructions to adjust the voice equalizer settings to settings corresponding to trumpet voices. The exemplary voice identifyr 108 is further described below in conjunction with FIG. 2.
[0020]
[0020] Figure 2 includes block diagrams of example implementations of the example speech analyzer 100 and example speech identifyr 108 of Figure 1. The example speech analyzer 100 of Figure 2 includes an example media interface 200, an example speech extractor 202, an example speech feature extractor 204, and an example device interface 206. The example speech identifyr 108 of Figure 2 includes an example device interface 210, an example timbre processor 212, an example timbre database 214, and an example speech settings adjuster 216. In some examples, elements of the example speech analyzer 100 may be implemented in the example speech identifyr 108 and / or elements of the example speech identifyr 108 may be implemented in the example speech identifyr 108.
[0021] The example media interface 200 of Figure 2 receives (e.g., samples) the example media signal 106 of Figure 1. In some examples, the media interface 200 can be a microphone used to obtain the media signal 106 as audio by collecting the media signal 106 through detection of ambient sound. In some examples, the media interface 200 can be an interface for directly receiving an audio signal and / or a video signal (e.g., a digital representation of the media signal) to be output from the example media output device 102. In some examples, the media interface 200 can include two interfaces: a microphone for detecting and sampling ambient sound and an interface for directly receiving and / or sampling the audio signal and / or the video signal.
[0022] 2 extracts audio from the received / sampled media signal 106. For example, the audio extractor 202 determines whether the received media signal 106 corresponds to an audio signal or a video signal with an audio component. If the media signal corresponds to a video signal with an audio component, the example audio extractor 202 extracts the audio component to generate an audio signal / sample for further processing.
[0023] The example audio extractor 204 of FIG. 2 processes the audio signal / samples to extract a pitch-independent timbre log-spectrum and / or a timbre-independent pitch log-spectrum. The log-spectrum is the convolution between a pitch-independent (e.g., pitchless) timbre log-spectrum and a timbre-independent (e.g., pitchless) pitch log-spectrum (e.g., X=T*P, where X is the log-spectrum of the audio signal, T is the pitch-independent log-spectrum, and P is the timbre-independent pitch log-spectrum). Thus, in the Fourier domain, the magnitude of the Fourier transform (FT) of the log-spectrum for the audio signal can be matched to an approximation of the FT of the timbre (e.g., F(X)=F(T)×F(P), where F(X) is the Fourier transform, F(T)≈|F(X)|, and F(P)≈e j arg(F(X)) (where, ∑ ∑ a ∑ b ... -1 ) (e.g., T=F -1 (|F(X)|)), (B) The logarithmic pitch spectrum of the toneless tone is calculated based on the inverse transformation of the complex argument of the transform output (e.g., P=F -1 (e j arg(F(X)) )). The logarithmic frequency scale of the audio spectrum of an audio signal allows pitch shifts to be equivalent to vertical translations. Therefore, the example audio feature extractor 204 uses CQT to determine the logarithmic spectrum of the audio signal.
[0024] In some examples, if the example voice feature extractor 204 of FIG. 2 determines that the resulting timbre and / or pitch are not satisfactory, the voice feature extractor 204 filters the results to improve the decomposition. For example, the voice feature extractor 204 can filter the results by emphasizing specific harmonics of timbre or by pushing a single peak / line into pitch and updating other resulting components. The example voice feature extractor 204 can perform a single filtering pass or an iterative algorithm, updating the filter / pitch with each iteration, thereby ensuring that the full convolution of pitch and timbre results in the original logarithmic spectrum of the voice. The voice feature extractor 204 can determine that these results are not satisfactory based on user and / or manufacturer preferences.
[0025] The example device interface 206 of the example speech analyzer 100 of FIG. 2 can interface with the example speech identifyr 108 and / or other devices (e.g., user interfaces, processing devices, etc.). For example, once the speech feature extractor 204 identifies pitch-independent timbre attributes, the example device interface 206 can communicate the attributes to the example speech identifyr 108 to classify the speech and / or identify the media. In response, the device interface 206 can receive (e.g., in the form of a signal or report) a classification result and / or identification information (e.g., an identifier corresponding to the source of the media signal 106) from the example speech identifyr 108. In such an example, the example device interface 206 can communicate the classification result and / or identification information to another device (e.g., a user interface) to display the classification result and / or identification information to a user. For example, when the speech analyzer 100 is used with a smartphone, the device interface 206 can output the classification result and / or identification information to a user of the smartphone via an interface (e.g., a screen) of the smartphone.
[0026] The example device interface 210 of the example speech identifyr 108 of FIG. 2 receives pitch-independent timbre attributes from the example speech analyzer 100. Additionally, the example device interface 210 outputs a signal / report representing the classification and / or identification information identified by the example speech identifyr 108. The report may be a signal corresponding to the received timbre-based classification and / or identification information. In some examples, the device interface 210 communicates the report (e.g., including identification information of media corresponding to the timbre) to a processor (e.g., a processor of an audience measurement entity, etc.) for further processing. For example, the processor of the receiving device may process the report to generate media exposure metrics, audience measurement metrics, etc. In some examples, the device interface 210 communicates the report to the example speech analyzer 100.
[0027] The example timbre processor 212 of FIG. 2 processes the received timbre attributes of the example speech analyzer 100 to characterize the speech and / or identify the speech source. For example, the timbre processor 212 can compare the received timbre attributes with reference attributes in the example timbre database 214. In this manner, if the example timbre processor 212 determines that the received timbre attributes match the reference attributes, it classifies and / or identifies the speech source based on the data corresponding to the matched reference timbre attributes. For example, if the example timbre processor 212 determines that the received timbre attributes match the reference timbre attributes corresponding to a particular commercial, it identifies the speech source as the particular commercial. In some examples, the classification may include genre classification. For example, if the example timbre processor 212 determines several instruments based on their timbres, the example timbre processor 212 can identify the genre of the speech (e.g., classical, rock, hip hop, etc.) based on the identified instruments and / or based on the timbre itself. In some examples, if the timbre processor 212 does not find a match, the example timbre processor 212 stores the received timbre attributes as new reference timbre attributes in the timbre database 214. When the example timbre processor 212 stores the new reference timbre in the example timbre database 214, the example device interface 210 communicates a command to the example speech analyzer 100 to request identification information from the user (e.g., what the classification of the audio is, what the media source is, etc.). In this manner, if the speech analyzer 100 responds with additional information, the timbre database 214 can store the additional information along with the new reference timbre. In some examples, a technician analyzes the new reference timbre to identify the additional information. The example timbre processor 212 generates a report based on the classification results and / or the identification information.
[0028] 2 determines audio equalizer settings based on the classified audio. For example, if the classified audio corresponds to one or more instruments and / or genres, the example audio settings adjuster 216 may determine audio equalizer settings corresponding to the one or more instruments and / or genres. In some examples, if the audio is classified as classical music, the example audio settings adjuster 216 may select classical audio equalizer settings corresponding to classical music (e.g., bass level, vibrance level, etc.). In this manner, the example device interface 210 may communicate the audio equalizer settings to the example media output device 102 and / or the example audio analyzer 100 to adjust the audio equalizer settings of the example media output device 102.
[0029]
[0029] Although an exemplary method for implementing the exemplary speech analyzer 100 and exemplary speech identifyr 108 of Figure 1 is illustrated in Figure 2, one or more of the elements, processes, and / or devices illustrated in Figure 2 may be combined, divided, rearranged, omitted, removed, and / or implemented in any other manner. Furthermore, the exemplary media interface 200, the exemplary speech extractor 202, the exemplary speech feature extractor 204, the exemplary device interface 206, the exemplary speech settings adjuster 216, and / or more generally the exemplary speech analyzer 100 of Figure 2, and / or the exemplary device interface 210, the exemplary timbre processor 212, the exemplary timbre database 214, the exemplary speech settings adjuster 216, and / or more generally the exemplary speech identifyr 108 of Figure 2 may be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, the exemplary media interface 200, the exemplary audio extractor 202, the exemplary audio feature extractor 204, the exemplary device interface 206, and / or more generally the exemplary audio analyzer 100 of FIG. 2, and / or any of the exemplary device interface 210, the exemplary timbre processor 212, the exemplary timbre database 214, the exemplary audio settings adjuster 216, and / or more generally the exemplary audio identifyr 108 of FIG. 2 may be implemented by one or more analog or digital circuit(s), logic circuit(s), programmable processor(s), programmable controller(s), graphic processing unit(s) (GPU(s)), digital signal processor(s) (DPS(s)), application specific integrated circuit(s) (ASIC(s)), programmable logic device(s) (PLD(s)), and / or field programmable logic device(s) (FPLD(s)).When any of the apparatus or system claims of this patent are read as encompassing purely software and / or firmware implementations, at least one of the exemplary media interface 200, the exemplary audio extractor 202, the exemplary audio feature extractor 204, the exemplary device interface 206, and / or more generally the exemplary audio analyzer 100 of FIG. 2, and / or the exemplary device interface 210, the exemplary timbre processor 212, the exemplary timbre database 214, the exemplary audio settings adjuster 216, and / or more generally the exemplary audio identifyr 108 of FIG. 2, is expressly defined herein to include a non-transitory computer-readable storage device or storage disk, such as a memory, digital versatile disc (DVD), compact disc (CD), Blu-ray disc, or the like, that contains the software and / or firmware. Further still, the example speech analyzer 100 and / or example speech identifyr 108 of Figure 1 may include one or more elements, processes, and / or devices in addition to or instead of those shown in Figure 2 and / or may include two or more of any or all of the illustrated elements, processes, and devices. As used herein, the phrase "in communication," including variations thereof, encompasses direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or constant communication, but rather additionally includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.
[0030] A flowchart representing exemplary hardware logic or machine-readable instructions for implementing the speech analyzer 100 of Figure 2 is shown in Figure 3, and a flowchart representing exemplary hardware logic or machine-readable instructions for implementing the speech identifyr 108 of Figure 2 is shown in Figure 4. The machine-readable instructions may be a program or portion of a program for execution by a processor, such as the processor 612, 712 shown in the exemplary processor platforms 600, 700 discussed below in connection with Figures 6 and / or 7. The program may be embodied in software stored on a non-transitory computer-readable storage medium, such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, or memory associated with the processor 612, 712, although the program, or portions thereof, may alternatively be executed by a device other than the processor 612, 712 and / or embodied in firmware or dedicated hardware. Furthermore, although the example programs are described with reference to the flowcharts illustrated in Figures 3-4, many other ways of implementing the example speech analyzer 100 and / or the example speech identifyr 108 may alternatively be used. For example, the order in which the blocks are executed may be changed, and / or some of the illustrated blocks may be modified, eliminated, or combined. Additionally, or alternatively, some or all of these blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to perform their corresponding operations without executing software or firmware.
[0031] As noted above, the exemplary processes of Figures 3-4 can be implemented using executable instructions (e.g., computer- and / or machine-readable instructions) stored on a non-transitory computer- and / or machine-readable medium, such as a hard disk drive, flash memory, read-only memory, compact disk, digital versatile disk, cache, random access memory, and / or any other storage device or storage disk on which information is stored for any duration (e.g., for an extended period of time, permanently, for short instances, for temporary buffering, and / or for caching information). As used herein, the term non-transitory computer-readable medium is expressly defined to include any type of computer-readable storage device and / or storage disk and to exclude propagating signals and to exclude transmission media.
[0032]
[0032] The terms "comprising" and "comprising" (and all their forms and tenses) are used herein as open-ended terms. Thus, whenever any form of "comprising" or "comprises" (e.g., comprises, includes, comprises, includes, has, etc.) is used in a claim preamble or in any type of claim recitation, it is to be understood that additional elements, terms, etc. may be present without departing from the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used, for example, as a transitional term in a claim preamble, this term is open-ended in the same way that the terms "comprising" and "including" are open-ended. The term "and / or," when used in the form of, for example, A, B, and / or C, refers to any combination or subset of A, B, and / or C, such as (1) A alone, (2) B alone, (3) C alone, (4) A with B, (5) A with C, and (6) B with C.
[0033]
[0033] Figure 3 is an exemplary flowchart 300 depicting exemplary machine-readable instructions that may be executed by the exemplary speech analyzer 100 of Figures 1 and 2 to extract pitch-independent timbre attributes from a media signal (e.g., an audio signal of the media signal). Although the instructions of Figure 3 are described in conjunction with the exemplary speech analyzer 100 of Figure 1, the exemplary instructions may be used by a speech analyzer in any environment.
[0034] At block 302, the exemplary media interface 200 receives one or more media signals or samples of media signals (e.g., the exemplary media signal 106). As described above, the exemplary media interface 200 can receive the media signal 106 directly (e.g., as a signal traveling to or from the media output device 102) or indirectly (e.g., as a microphone that detects the media signal by sensing ambient sound). At block 304, the exemplary audio extractor 202 determines whether the media signal corresponds to video or audio. For example, if the media signal is received using a microphone, the audio extractor 202 determines that the media corresponds to audio. However, if the media signal is a received signal, the audio extractor 202 processes the received media signal to determine whether the media signal corresponds to audio or a video signal with an audio component. If the exemplary audio extractor 202 determines that the media signal corresponds to audio (block 304: Audio), the process continues to block 308. If the example audio extractor 202 determines that the media signal corresponds to video (block 306: video), the example audio extractor 202 extracts audio components from the media signal (block 306).
[0035] At block 308, the exemplary speech feature extractor 204 determines the log-spectrum (e.g., X) of the speech signal. For example, the speech feature extractor 204 may determine the log-spectrum of the speech signal by performing a CQT. At block 310, the exemplary speech feature extractor 204 transforms the log-spectrum to the frequency domain. For example, the speech feature extractor 204 performs a FT on the log-spectrum (e.g., F(X)). At block 312, the exemplary speech feature extractor 204 determines the magnitude of the transform update (e.g., |F(X)|). At block 314, the exemplary speech feature extractor 204 determines the pitch-independent timbre log-spectrum of the speech based on the inverse transform (e.g., inverse FT) of the magnitude of the transform output (e.g., T=F -1 |F(X)|). At block 316, the example speech feature extractor 204 identifies the complex argument of the transform output (e.g., e j arg(F(X)) At block 318, the example voice feature extractor 204 determines the timbre-independent log-pitch spectrum of the voice based on the inverse transform (e.g., inverse FT) of the complex argument of the transform output (e.g., P=F -1 (e j arg(F(X) )).
[0036] At block 320, the example audio feature extractor 204 determines whether the result(s) (e.g., the identified pitch and / or the identified timbre) are satisfactory. As discussed above in conjunction with FIG. 2, the example audio feature extractor 204 determines that the result is satisfactory based on the user's and / or manufacturer's result preferences. If the example audio feature extractor 204 determines that the result is satisfactory (block 320: yes), the process continues to block 324. If the example audio feature extractor 204 determines that the result is satisfactory (block 320: no), the example audio feature extractor 204 filters the result (block 322). As discussed above in conjunction with FIG. 2, the example audio feature extractor 204 can filter the result by emphasizing specific harmonics of the timbre or by squeezing a single peak / line into the pitch (e.g., once or repeatedly).
[0037] At block 324, the example device interface 206 communicates the results to the example voice identifyr 108. At block 326, the example voice feature extractor 204 receives the classification results and / or identification data corresponding to the voice signal. Alternatively, if the voice identifyr 108 is unable to match the timbre of the voice signal to the reference, the device interface 206 can send instructions to identify additional data corresponding to the voice signal. In such an example, the device interface 206 communicates a prompt to the user interface to prompt the user to provide the additional data. Thus, the example device interface 206 can provide the additional data to the example voice identifyr 108 to generate new reference timbre attributes. At block 328, the example voice feature extractor 204 communicates the classification results and / or identification information to other connected devices. For example, the voice feature extractor 204 communicates the classification results to the user interface to provide the classification results to the user.
[0038]
[0038] Figure 4 is an example flowchart 400 depicting example machine-readable instructions that may be executed by the example voice identifyr 108 of Figures 1 and 2 to classify audio and / or identify media based on pitch-independent timbre attributes of the audio. Although the instructions of Figure 4 are described in conjunction with the example voice identifyr 108 of Figure 1, the example instructions may be used by a voice identifyr in any environment.
[0039] At block 402, the example device interface 210 receives a measured (e.g., identified or extracted) timbre logarithmic spectrum of a silent pitch from the example speech analyzer 100. At block 404, the example timbre processor 212 compares the measured timbre logarithmic spectrum of the silent pitch with the timbre logarithmic spectrum of a reference silent pitch stored in the example timbre database 214. At block 406, the example timbre processor 212 determines whether a match is found between the timbre attributes of the received silent pitch and the timbre attributes of the reference silent pitch. If the example timbre processor 212 determines that a match is found (block 406: Yes), the example timbre processor 212 uses additional data stored in the example timbre database 214 corresponding to the matched reference timbre attributes to classify the audio (e.g., identify the instrument and / or genre) and / or identify the media corresponding to the audio based on the match (block 408).
[0040] At block 410, the example audio settings adjuster 216 determines whether the audio settings of the media output device 102 can be adjusted. For example, there may be an enabled setting that allows the audio settings of the media output device 102 to be adjusted based on the classification results of the audio being output from the example media output device 102. If the example audio settings adjuster 216 determines that the audio settings of the media output device 102 should not be adjusted (block 410: No), the process proceeds to block 414. If the example audio settings adjuster 216 determines that the audio settings of the media output device 102 should be adjusted (block 410: Yes), the example audio settings adjuster 216 determines media output device setting adjustments based on the classified audio. For example, the example audio settings adjuster 216 may select audio equalizer settings based on one or more identified instruments and / or an identified genre (e.g., by timbre or based on the identified instruments) (block 412). At block 414, the example device interface 210 outputs a report corresponding to the classification results, the identification information, and / or the media output device setting adjustment. In some examples, the device interface 210 outputs the report to another device for further processing / analysis. In some examples, the device interface 210 outputs the report to the example speech analyzer 100 to display the results to a user via a user interface. In some examples, the device interface 210 outputs the report to the example media output device 102 to adjust the audio settings of the media output device 102.
[0041] If the example timbre processor 212 determines that a match is not found (block 406: NO), the example device interface 210 prompts for additional information corresponding to the audio signal (block 416). For example, the device interface 210 can communicate instructions to the example speech analyzer 100 to (A) prompt the user to provide information corresponding to the audio, or (B) prompt the speech analyzer 100 to respond with the complete audio signal. At block 418, the example timbre database 214 stores the measured untoned pitch log spectrum along with any corresponding data that may have been received.
[0042] FIG. 5 shows an exemplary FT of a logarithmic spectrum 500 of an audio signal, an exemplary non-tonal pitch logarithmic spectrum 502 of an audio signal, and an exemplary non-tonal pitch timbre logarithmic spectrum 504 of an audio signal.
[0043] As described in conjunction with FIG. 2, when the exemplary audio analyzer 100 receives the exemplary media signal 106 (e.g., or a sample of the media signal), the exemplary audio analyzer 100 determines an exemplary logarithmic spectrum of the audio signal / sample (e.g., when the media sample corresponds to a video signal and the audio analyzer 100 extracts its audio component). In addition, the exemplary audio analyzer 100 determines a FT of the logarithmic spectrum. The exemplary FT logarithmic spectrum 500 of FIG. 5 corresponds to an exemplary transform output of the logarithmic spectrum of the audio signal / sample. The exemplary tonal pitch logarithmic spectrum 502 corresponds to the inverse FT of the complex argument of the exemplary FT of the logarithmic spectrum 500 (e.g., P=F -1 (e j arg(F(X)) )), the tonal log spectrum 504 of the neutral pitch corresponds to the inverse FT of the magnitude of the exemplary FT of the log spectrum 500 (e.g., T=F -1 5, an example FT of a log spectrum 500 corresponds to the convolution of an example neutral pitch log spectrum 502 with an example neutral pitch timbre log spectrum 504. The peaked convolution of the example pitch log spectrum 502 adds an offset.
[0044]
[0044] Figure 6 is a block diagram of an exemplary processor platform 600 configured to execute the instructions of Figure 3 to implement the speech analyzer 100 of Figure 2. The processor platform 600 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smartphone, a tablet such as an iPad™), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset or other wearable device, or any other type of computing device.
[0045] The processor platform 600 of the illustrated example includes a processor 612. The processor 612 of the illustrated example is hardware. For example, the processor 612 may be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor may be a semiconductor-based (e.g., silicon-based) device. In this example, the processor implements the example media interface 200, the example speech extractor 202, the example speech feature extractor 204, and / or the example device interface of FIG. 2 .
[0046] The processor 612 of the illustrated example includes a local memory 613 (e.g., a cache). The processor 612 of the illustrated example communicates with main memory, including volatile memory 614 and nonvolatile memory 616, via a bus 618. The volatile memory 614 may be implemented with synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The nonvolatile memory 616 may be implemented with flash memory and / or any other desired type of memory device. Access to the main memory 614, 616 is controlled by a memory controller.
[0047] The processor platform 600 of the depicted example also includes an interface circuit 620. The interface circuit 620 may be implemented with any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), a Bluetooth interface, a Near Field Communication (NFC) interface, and / or a PCI Express interface.
[0048] In the depicted example, one or more input devices 622 are connected to the interface circuit 620. The input device(s) 622 allow a user to input data and / or commands into the processor 612. The input device(s) may be implemented, for example, by an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touchscreen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.
[0049]
[0049] One or more output devices 624 are also connected to the interface circuitry 620 of the illustrated example. The output device(s) 624 may be implemented by, for example, a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-place switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Accordingly, the interface circuitry 620 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0050] The interface circuitry 620 of the illustrated example also includes communications devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate exchanging data with external machines (e.g., any type of computing device) over a network 626. Communications may be via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-sight wireless system, a cellular telephone system, etc.
[0051] The processor platform 600 of the depicted example also includes one or more mass storage devices 628 for storing software and / or data. Examples of such mass storage devices 628 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0052]
[0052] The machine-executable instructions 632 of Figure 3 may be stored on the mass storage device 628, the volatile memory 614, the non-volatile memory 616, and / or a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0053]
[0053] Figure 7 is a block diagram of an exemplary processor platform 700 configured to execute the instructions of Figure 4 to implement the speech identifyr 108 of Figure 2. The processor platform 700 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a mobile phone, a smartphone, a tablet such as an iPad™), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a game console, a personal video recorder, a set-top box, a headset or other wearable device, or any other type of computing device.
[0054] The processor platform 700 of the illustrated example includes a processor 712. The processor 712 of the illustrated example is hardware. For example, the processor 712 may be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. A hardware processor may be a semiconductor-based (e.g., silicon-based) device. In this example, the processor implements the example device interface 210, the example timbre processor 212, the example timbre database 214, and / or the example audio settings adjuster 216.
[0055] The processor 712 of the illustrated example includes a local memory 713 (e.g., a cache). The processor 712 of the illustrated example communicates with main memory, including volatile memory 714 and nonvolatile memory 716, via a bus 718. The volatile memory 714 may be implemented with synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS® dynamic random access memory (RDRAM®), and / or any other type of random access memory device. The nonvolatile memory 716 may be implemented with flash memory and / or any other desired type of memory device. Access to the main memory 714, 716 is controlled by a memory controller.
[0056] The processor platform 700 of the illustrated example also includes an interface circuit 720. The interface circuit 720 may be implemented with any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), a Bluetooth interface, a Near Field Communication (NFC) interface, and / or a PCI Express interface.
[0057] In the depicted example, one or more input devices 722 are connected to the interface circuit 720. The input device(s) 722 allow a user to input data and / or commands into the processor 712. The input device(s) may be implemented, for example, by an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touchscreen, a trackpad, a trackball, an isopoint, and / or a voice recognition system.
[0058]
[0058] One or more output devices 724 are also connected to the interface circuitry 720 of the illustrated example. The output device(s) 724 may be implemented by, for example, a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-place switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Accordingly, the interface circuitry 720 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0059] The interface circuitry 720 of the illustrated example also includes communications devices such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces to facilitate exchanging data with external machines (e.g., any type of computing device) over a network 726. Communications may be via, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-sight wireless system, a cellular telephone system, etc.
[0060] The processor platform 700 of the depicted example also includes one or more mass storage devices 728 for storing software and / or data. Examples of such mass storage devices 728 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0061]
[0061] The machine-executable instructions 732 of Figure 4 may be stored on the mass storage device 728, the volatile memory 714, the non-volatile memory 716, and / or a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0062]
[0062] From the foregoing, it can be appreciated that the above-disclosed methods, apparatus, and articles of manufacture extract pitch-independent timbre attributes from a media signal. Examples disclosed herein identify a non-pitched pitch-independent timbre logarithmic spectrum based on audio received directly or indirectly from a media output device. Examples disclosed herein further include classifying audio based on timbre (e.g., identifying musical instruments) and / or identifying the media source of audio based on timbre (e.g., songs, video games, advertisements, etc.). Using the examples disclosed herein, audio can be classified and / or identified using timbre with significantly fewer resources than conventional techniques because the extracted timbre is pitch-independent. Accordingly, audio can be classified and / or identified without requiring multiple reference timbre attributes for multiple pitches. Rather, pitch-independent timbre can be used to classify audio regardless of pitch.
[0063] Although several exemplary methods, apparatus, and articles of manufacture have been described herein, other implementations are possible. The scope of protection of this patent is not limited to these methods, apparatus, and articles of manufacture. Rather, this patent encompasses all methods, apparatus, and articles of manufacture that fall fully within the scope of the claims of this patent. [Item of invention] [Item 1] 1. An apparatus for extracting pitch-independent timbre attributes from a media signal, comprising: an interface for receiving a media signal; determining a spectrum of audio corresponding to the media signal; Identifying pitch-independent timbre attributes of the audio based on an inverse transform of the magnitude of the spectral transform. speech feature extractor for An apparatus comprising: [Item 2] Item 1, wherein the media signal is the audio. [Item 3] Item 10. The apparatus of item 1, wherein the media signal is a video signal including an audio component, and further comprising an audio extractor for extracting the audio from the video signal. [Item 4] Item 10. The apparatus of item 1, wherein the audio feature extractor determines the spectrum of the audio using a constant-Q transform. [Item 5] Item 1. The apparatus of item 1, wherein the audio feature extractor determines the transform of the spectrum using a Fourier transform and the inverse transform using an inverse Fourier transform. [Item 6] Item 10. The apparatus of item 1, wherein the audio feature extractor identifies timbre-independent pitch attributes of the audio based on an inverse transform of a complex argument of the transform of the spectrum. [Item 7] the interface is a first interface, communicating said pitch-independent timbre attributes to a processing device; receiving at least one of a classification result of the audio or an identifier corresponding to the media signal from the processing device in response to communicating the pitch-independent timbre attribute to the processing device; Item 10. The device of item 1, further comprising a second interface for [Item 8] 8. The apparatus of claim 7, wherein the second interface is for communicating at least one of the classification result of the audio or an identifier corresponding to the media signal to a user interface. [Item 9] Item 10. The device of item 1, wherein the interface is a microphone for receiving the media signal via ambient audio. [Item 10] Item 10. The apparatus of item 1, wherein the media signal corresponds to a media signal to be output by a media output device. [Item 11] Item 10. The device of item 1, wherein the interface receives the media signal from a microphone. [Item 12] A non-transitory computer-readable storage medium containing instructions that, when executed, cause a machine to perform at least: accessing media signals; determining an audio spectrum corresponding to said media signal; identifying pitch-independent timbre attributes of the audio based on an inverse transform of the magnitude of the spectral transform; A non-transitory computer-readable storage medium that causes the [Item 13] Item 13. The non-transitory computer-readable storage medium of item 12, wherein the media signal is audio. [Item 14] Item 13. The non-transitory computer-readable storage medium of item 12, wherein the media signal is a video signal including an audio component, and the instructions, when executed, cause the machine to extract the audio from the video signal. [Item 15] Item 13. The non-transitory computer-readable storage medium of item 12, wherein the instructions, when executed, cause the machine to determine the spectrum of the audio using a constant-Q transform. [Item 16] Item 13. The non-transitory computer-readable storage medium of item 12, wherein the instructions, when executed, cause the machine to determine the transform of the spectrum using a Fourier transform and to determine the inverse transform using an inverse Fourier transform. [Item 17] Item 13. The non-transitory computer-readable storage medium of item 12, wherein the instructions, when executed, cause the machine to determine timbre-independent pitch attributes of the audio based on an inverse transform of a complex argument of the transform of the spectrum. [Item 18] The instructions, when executed, cause the machine to: conveying said pitch-independent timbre attributes to a processing device; receiving from the processing device at least one of a classification result of the audio or an identifier corresponding to the media signal in response to communicating the pitch-independent timbre attributes to the processing device; Item 13. The non-transitory computer-readable storage medium of item 12, which causes the computer to execute the steps described above. [Item 19] Item 19. The non-transitory computer-readable storage medium of item 18, wherein the instructions, when executed, cause the machine to communicate at least one of the classification result of the audio or an identifier corresponding to the media signal to a user interface. [Item 20] 1. A method for extracting pitch-independent timbre attributes from a media signal, comprising: determining a spectrum of audio corresponding to a received media signal by executing instructions on a processor; identifying pitch-independent timbre attributes of the voice based on an inverse transform of the magnitude of the spectral transform by executing instructions on the processor; A method comprising:
Claims
1. 1. A method comprising: determining a logarithmic spectrum (X) of the audio signal; transforming the logarithmic spectrum (X) of the audio signal into the frequency domain to generate a transformed output (F(X)); determining the magnitude (|F(X)|) of the transform output; The inverse transformation of the magnitude of the transformed output of the logarithmic spectrum (F -1 determining pitch-independent timbre attributes (T) of the audio signal based on |F(X)|; adjusting at least one audio equalizer setting of a media output device that outputs the audio signal in response to identifying the pitch-independent timbre attribute (T); A method comprising:
2. The method of claim 1 further comprising the step of extracting the audio signal from a video signal.
3. 2. The method of claim 1, further comprising determining a timbre-independent pitch log-spectrum (P) of the audio signal based on an inverse transform of the complex argument of the transform output (F(X)) of the log-spectrum.
4. 4. The method of claim 3, further comprising updating at least one of the timbre-independent pitch logarithmic spectrum (P) or the pitch-independent timbre attribute (T) based on a single peak or line in the timbre-independent pitch logarithmic spectrum (P).
5. 5. The method of claim 4, further comprising the step of iteratively updating at least one of the timbre-independent pitch log-spectrum (P) or the pitch-independent timbre attributes (T) until a convolution of the timbre-independent pitch log-spectrum (P) and the pitch-independent timbre attributes (T) reproduces the original log-spectrum (X).
6. 2. The method of claim 1, wherein the transforming of the logarithmic spectrum (X) into the frequency domain is performed using a Fourier transform, and the inverse transforming is performed using an inverse Fourier transform.
7. The method of claim 1 , wherein the step of determining the logarithmic spectrum (X) comprises using a constant-Q transform.
8. comparing said pitch-independent timbre attributes (T) with one or more reference pitch-independent timbre log-spectra; classifying the audio signal based on the match; The method of claim 1 further comprising:
9. The method of claim 8 , wherein the classification corresponds to at least one of an instrument or a genre.
10. 9. The method of claim 8, further comprising prompting a user to provide additional information corresponding to the audio signal if it is determined that the pitch-independent timbral attributes (T) do not match the one or more reference pitch-independent timbral logarithmic spectra.
11. A computer-readable storage medium containing instructions, The instructions, when executed by a processor, cause the processor to: determining a logarithmic spectrum (X) of the audio signal; transforming the logarithmic spectrum (X) of the audio signal into the frequency domain to generate a transformed output (F(X)); determining the magnitude (|F(X)|) of the transform output; The inverse transformation of the magnitude of the transformed output of the logarithmic spectrum (F -1 determining pitch-independent timbre attributes (T) of the audio signal based on |F(X)|; adjusting at least one audio equalizer setting of a media output device that outputs the audio signal in response to identifying the pitch-independent timbre attribute (T); 10. A computer-readable storage medium for causing a computer to perform operations including:
12. The operation is 12. The computer-readable storage medium of claim 11, further comprising determining a timbre-independent pitch log-spectrum (P) of the audio signal based on an inverse transform of a complex argument of the transform output (F(X)) of the log-spectrum.
13. The operation is 13. The computer-readable storage medium of claim 12, further comprising updating at least one of the timbre-independent pitch logarithmic spectrum (P) or the pitch-independent timbre attributes (T) based on a single peak or line in the timbre-independent pitch logarithmic spectrum (P).
14. The operation is 14. The computer-readable storage medium of claim 13, further comprising iteratively updating at least one of the timbre-independent pitch log-spectrum (P) or the pitch-independent timbre attributes (T) until a convolution of the timbre-independent pitch log-spectrum (P) and the pitch-independent timbre attributes (T) reproduces the original log-spectrum (X).
15. 12. The computer-readable storage medium of claim 11, wherein the transforming of the logarithmic spectrum (X) to the frequency domain is performed using a Fourier transform, and the inverse transforming is performed using an inverse Fourier transform.
16. The operation is comparing said pitch-independent timbre attributes (T) with one or more reference pitch-independent timbre log-spectra; classifying the audio signal based on the match; The computer-readable storage medium of claim 11 , further comprising:
17. 1. An apparatus comprising a processor and a memory for storing instructions, The instructions, when executed by a processor, cause the device to: determining a logarithmic spectrum (X) of the audio signal; transforming the logarithmic spectrum (X) of the audio signal into the frequency domain to generate a transformed output (F(X)); determining the magnitude (|F(X)|) of the transform output; The inverse transformation of the magnitude of the transformed output of the logarithmic spectrum (F -1 determining pitch-independent timbre attributes (T) of the audio signal based on |F(X)|; adjusting at least one audio equalizer setting of a media output device that outputs the audio signal in response to identifying the pitch-independent timbre attribute (T); A device that performs the following.
18. 18. The apparatus of claim 17, wherein the transforming of the logarithmic spectrum (X) to the frequency domain is performed using a Fourier transform and the determining of the inverse transform is performed using an inverse Fourier transform.
19. 18. The apparatus of claim 17, wherein determining the logarithmic spectrum (X) comprises using a constant-Q transform.
20. The instruction:
18. The apparatus of claim 17, further comprising determining a timbre-independent pitch logarithmic spectrum (P) of the audio signal based on an inverse transform of a complex argument of the transform output (F(X)) of the logarithmic spectrum.