Method and device for identifying media
By generating and adjusting the sample media fingerprint, the problem of difficulty in identifying pitch offset, time offset and resampling audio in the prior art is solved, and more efficient and accurate audio fingerprint recognition is achieved.
Patent Information
- Application Number
- CN202080062321.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-27
- Filing Date
- 2020-09-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-09-04
AI Technical Summary
When existing audio fingerprint recognition technology processes audio signals, it is difficult to effectively identify pitch offset, time offset and resampled audio, resulting in a decrease in the accuracy and efficiency of fingerprint recognition.
Generate sample media fingerprints and adjust them to suit pitch offsets, time offsets, and resampled audio. Specific methods include adjusting the bin value of the sample fingerprint to accommodate pitch offsets, adjusting the number of frames to accommodate time offsets, and applying a resampling rate if necessary.
It realizes effective identification of pitch offset, time offset and resampled audio, improves the accuracy and efficiency of audio fingerprint recognition, and reduces the overhead of adjusting and processing audio samples.
Smart Images

Figure CN114341854B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to signatures and, more particularly, to methods and apparatus for identifying media. Background Art
[0002] Media (e.g., sound, speech, music, video, etc.) can be represented as digital data (e.g., electronic, optical, etc.). Media captured (e.g., via a microphone and / or camera) can be digitized, electronically stored, processed, and / or cataloged. One method of cataloging media (e.g., audio information) is to generate a signature (e.g., fingerprint, watermark, audio signature, audio fingerprint, audio watermark, etc.). A signature is a digital summary of the media created by sampling a portion of the media signal. In the past, signatures have been used to identify media and / or verify the authenticity of the media. Summary of the invention
[0003] In one aspect, a non-transitory computer-readable storage medium is disclosed. The non-transitory computer-readable storage medium includes instructions that, when executed, cause one or more processors to at least: generate an adjusted sample media fingerprint by applying an adjustment to the sample media fingerprint in response to a query; compare the adjusted sample media fingerprint to a reference media fingerprint; and, in response to the adjusted sample media fingerprint matching the reference media fingerprint, send information associated with the reference media fingerprint and the adjustment.
[0004] In another aspect, an apparatus is disclosed. The apparatus includes a fingerprint tuner that generates an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint in response to a query, a comparator that compares the adjusted sample media fingerprint with a reference media fingerprint, and a network interface that transmits information associated with the reference media fingerprint and the adjustment in response to the adjusted sample media fingerprint matching the reference media fingerprint.
[0005] In yet another aspect, a method is disclosed. The method includes: in response to a query, generating an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint; comparing the adjusted sample media fingerprint to a reference media fingerprint; and in response to the adjusted sample media fingerprint matching the reference media fingerprint, sending information associated with the reference media fingerprint and the adjustment. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 is a block diagram of an example environment including an example central facility and an example application.
[0007] Figure 2 It is shown Figure 1 A block diagram of an example central facility with further details.
[0008] Figure 3 It is shown Figure 1 A block diagram of an example application with further details.
[0009] Figure 4A and Figure 4B Example of Figure 2 and Figure 3 Example spectrogram generated by the fingerprint generator.
[0010] Figure 4C Example of Figure 2 and Figure 3 Example normalized spectrogram generated by the fingerprint generator.
[0011] Figure 4D Example of Figure 2 and Figure 3 An example sub-fingerprint of an example fingerprint generated by a fingerprint generator based on the normalized spectrogram.
[0012] Figure 5 Example audio samples and corresponding fingerprints are illustrated.
[0013] Fig. 6A and Figure 6B is a flowchart representing a process for running a query associated with pitch-shifted, time-shifted, and / or resampled media, which may be performed using Figure 1 and Figure 2 It is implemented by machine-readable instructions of a central facility.
[0014] Figure 7 is a flow chart showing a process for identifying trends in broadcast and / or presentation media that may be performed to implement Figure 1 and Figure 2 It is implemented by machine-readable instructions of a central facility.
[0015] Figure 8 is a flowchart showing a process for pitch shifting a fingerprint, which may be performed using Figure 1 and Figure 2 It is implemented by machine-readable instructions of a central facility.
[0016] Fig. 9 is a flow chart showing a process for time shifting a fingerprint, which may be performed using Figure 1 and Figure 2 It is implemented by machine-readable instructions of a central facility.
[0017] Fig.10is a flow chart representing a process for transmitting information associated with a pitch shift, a time shift, and / or a resampling rate, which may be implemented using a method that may be executed to implement Figure 1 and Figure 3 The application is implemented by machine-readable instructions.
[0018] Fig.11 is constructed to execute Fig. 6A , Figure 6B , Figure 7 , Figure 8 and Fig. 9 processing to achieve Figure 1 and Figure 2 Block diagram of an example processing platform of a central facility.
[0019] Fig.12 is constructed to execute Fig.10 processing to achieve Figure 1 and Figure 3 Block diagram of an example processing platform for an application of .
[0020] The drawings are not drawn to scale. Generally, the same reference numerals will be used throughout the drawings and the accompanying written description to refer to the same or similar parts. References to connection (e.g., attach, couple, connect, and combine) are to be interpreted broadly and may include intermediate members between sets of elements and relative movement between elements, unless otherwise specified. Thus, references to connection do not necessarily infer that two elements are directly connected and have a fixed relationship to each other.
[0021] When identifying multiple elements or components that can be referred to separately, the descriptors "first", "second", "third", etc. are used herein. Unless otherwise specified or understood from the context of their use, such descriptors are not intended to confer any priority, physical order, or arrangement in a list or time order, but are merely used as labels to refer to multiple elements or components separately to facilitate understanding of the disclosed examples. In some examples, the descriptor "first" may be used to refer to an element in a specific embodiment, while a different descriptor may be used in the claims to refer to the same element, such as "second" or "third". In this case, it should be understood that these descriptors are only used to facilitate reference to multiple elements or components. DETAILED DESCRIPTION
[0022] Fingerprint or signature-based media monitoring techniques typically exploit one or more inherent characteristics of the media being monitored during a monitoring time interval to generate a substantially unique proxy for the media. Such a proxy is called a signature or fingerprint, and can take any form (e.g., a series of digital values, a waveform, etc.) that represents any aspect of a media signal (e.g., an audio signal and / or a video signal that forms a presentation of the media being monitored). A signature can be a series of signatures collected continuously over a time interval. The terms "fingerprint" and "signature" are used interchangeably herein and are defined herein to mean a proxy generated based on one or more inherent characteristics of the media for identifying the media.
[0023] Signature-based media monitoring generally involves determining (e.g., generating and / or collecting) a signature representing a media signal (e.g., an audio signal and / or a video signal) output by a monitored media device, and comparing the monitored signature with one or more reference signatures corresponding to a known (e.g., reference) media source. Various comparison criteria (such as cross-correlation values, Hamming distances, etc.) may be evaluated to determine whether the monitored signature matches a particular reference signature.
[0024] When a match is found between a monitored signature and one of the reference signatures, the monitored media can be identified as corresponding to a particular reference media represented by the reference signature that matches the monitored signature. Since attributes such as media identifier, presentation time, broadcast channel, etc. are collected for the reference signatures, these attributes can then be associated with the monitored media whose monitored signature matches the reference signature. Example systems for identifying media based on codes and / or signatures are already known and were first disclosed in U.S. Pat. No. 5,481,294 to Thomas, the entire contents of which are incorporated herein by reference.
[0025] In the past, audio fingerprinting techniques used the loudest parts of an audio signal (e.g., the parts with the most energy, etc.) to create a fingerprint in a time period. For example, some audio fingerprinting techniques use one and / or more frequency ranges with the greatest energy to create a fingerprint in a time period. However, in some cases, this approach has multiple serious limitations. In some examples, the loudest parts of an audio signal may be associated with noise (e.g., unwanted audio) rather than from audio of interest. For example, if a user attempts to fingerprint a song in a noisy restaurant, the loudest parts of the captured audio signal may be conversations between restaurant patrons rather than the song or media to be identified. In this example, many sampled portions of the audio signal will have background noise rather than music, which reduces the usefulness of the generated fingerprint.
[0026] In addition, media producers (e.g., radio studios, television studios, recording studios, etc.) adjust and / or otherwise manipulate media before and / or at broadcast time. Adjustments may correspond to transformations, pitch shifts, time shifts, resampling, and / or otherwise manipulating media. For example, a radio studio may increase and / or decrease the playback speed of the audio to increase and / or decrease the amount of media that can be played in a certain time period. In some examples, if an artist records a cover of another song, the recording studio may increase and / or decrease the pitch of the recorded audio to allow the artist to record music in a more comfortable range for the artist. In additional or alternative examples, a radio studio may resample and / or otherwise remix the audio to adjust the media. Resampling may refer to an audio adjustment that creates a dependency between the pitch adjustment of the audio and the playback adjustment. For example, when resampling, increasing the playback speed of the recording not only increases the playback speed of the audio, but also increases the pitch of the audio.
[0027] In addition, a disc jockey (DJ) can adjust and / or otherwise manipulate media before and / or during playback. A DJ can adjust and / or otherwise manipulate media for broadcast purposes (e.g., radio broadcasts, television broadcasts, etc.) and / or entertainment (e.g., nightclubs). For example, a DJ can change the pitch and / or playback speed of a song to obtain a dramatic effect, to smooth the transition between songs, and / or create a combination and / or arrangement of one or more audio files. For example, some DJs can adjust the pitch of the audio by up to 12%. In other examples, a DJ can adjust the pitch of the audio by 6%. In addition, a DJ can resample the audio to change the pitch and / or playback speed of the audio. As used herein, a DJ is a person who plays music for a live audience. For example, a DJ can be a professional who often performs in a group of performances (e.g., a DJ in Las Vegas, a DJ in Los Angeles, etc.). Additionally or alternatively, a DJ can be a semi-professional (e.g., a wedding DJ, a DJ hired for dancing, etc.) who performs less frequently than a full-time DJ. In some examples, a DJ can be an individual who mixes audio for himself or for limited purposes.
[0028] Although the above-mentioned adjustment of the pitch and / or playback speed of the audio can make it pleasant to listen to, the adjustment of the pitch and / or playback speed of the audio will change the fingerprint and / or signature associated with the audio, thereby adversely affecting the ability of the media recognition entity to identify the media based on the signature and / or fingerprint. For example, conventional fingerprint and / or signature generation techniques are not robust enough to detect fingerprints based on audio that has been pitch-shifted, time-shifted and / or resampled (e.g., detecting pitch-shifted, time-shifted and / or resampled audio). On the contrary, in order to detect pitch-shifted, time-shifted and / or resampled audio, conventional techniques rely on adjusting audio samples to compensate for the assumed pitch shift, time shift and / or resampling rate, so as to detect fingerprints based on pitch-shifted, time-shifted and / or resampled audio. For example, conventional techniques can reprocess audio samples, generate adjusted sample fingerprints and / or adjusted sample signatures, and compare the adjusted sample fingerprints and / or adjusted sample signatures with one or more reference fingerprints, reference signatures and / or reference audio samples.
[0029] Although conventional techniques for detecting pitch-shifted audio rely on adjusting audio samples and generating fingerprints and / or signatures for the adjusted audio samples, the examples disclosed herein avoid this processing overhead and increase audio detection. Rather than adjusting audio samples to match reference audio samples, the examples disclosed herein generate sample fingerprints and / or sample signatures, and then adjust the sample fingerprints and / or sample signatures to detect pitch-shifted, time-shifted, and / or resampled audio. For example, the examples disclosed herein may adjust sample fingerprints and / or sample signatures to adjust bin values associated with sample fingerprints and / or sample signatures (e.g., of, in, etc., sample fingerprints and / or sample signatures) to accommodate pitch shifts. Additionally or alternatively, the examples disclosed herein may adjust the number of frames in a sample signature and / or sample fingerprint to accommodate time shifts.
[0030] In addition, examples disclosed herein can monitor media over a period of time to determine trends and / or patterns associated with common pitch shifts, time offsets, and / or resampling rates. For example, examples disclosed herein can identify pitch shifts, time offsets, and / or resampling rates based on the musicality of the pitch shifts, time offsets, and / or resampling rates. For example, due to the musicality of one to five percent pitch shifts, one to five percent pitch shifts may be more common than fifty percent pitch shifts. Therefore, examples disclosed herein can identify pitch shifts, time offsets, and / or resampling rates of music.
[0031] In some examples disclosed herein, a user device runs an application that generates a sample fingerprint from the obtained audio. The examples disclosed herein further indicate that an external device (e.g., a server, a central facility, a cloud-based processor, etc.) runs a query based on the sample fingerprint. For example, a query may include one or more pitch offset values, one or more time offset values, and one or more resampling rates corresponding to a conjecture change of an audio signal associated with a sample fingerprint. In addition, the examples disclosed herein send adjustment instructions (e.g., queries) and generated (e.g., sample) fingerprints to an external device. The adjustment instruction identifies one or more pitch offsets, time offsets, and / or resampling rates that an external device should execute in a query to try to find a match with the sample fingerprint. When an external device matches a sample fingerprint with a reference fingerprint (e.g., a reference media fingerprint) stored at an external device, and / or when the adjusted sample fingerprint (e.g., adjusted sample media fingerprint) adjusted according to the adjustment instruction is matched with a reference fingerprint stored at an external device, the external device sends information corresponding to the match (e.g., the author, artist, title, etc. of the audio). In addition, if the adjusted sample fingerprint matches the reference fingerprint, the external device will send to the user device how the audio is pitch-shifted, time-shifted, and / or resampled based on how the sample fingerprint is adjusted to match the reference fingerprint. The examples disclosed herein can report how the matched audio is pitch-shifted, time-shifted, and / or resampled, and / or adjust subsequent adjustment instructions based on how the matched audio is pitch-shifted, time-shifted, and / or resampled.
[0032] In some examples, the external device can adjust the reference fingerprint to match the reference fingerprint with the sample fingerprint. For example, the external device can apply a pitch shift, a time shift, or a resampling rate to the reference fingerprint to add the conjectured pitch shift, time shift, and / or resampling rate so that the external device can compare and / or match the reference fingerprint with the sample fingerprint. In such an example, the external device generates the adjusted fingerprint as (a) one or more pitch-shifted reference fingerprints, (b) one or more time-shifted reference fingerprints, and / or (c) one or more resampled reference fingerprints.
[0033] 1 is a block diagram of an example environment 100. The example environment 100 includes an example client device 102, an example network 104, an example wireless communication system 106, an example end-user device 108, an example media broadcaster 110, and an example central facility 112. Each of the example client device 102 and the example end-user device 108 includes an example application 114.
[0034] exist Figure 1In an example of the present invention, client device 102 is a laptop computer. For example, client device 102 may be a work laptop computer of an employee of a company that owns the rights to a master license for a song or other media (e.g., a publisher, a record company, etc.). In additional or alternative examples, client device 102 may be any number of desktop computers, laptop computers, workstations, mobile phones, tablet computers, servers, any suitable computing devices, or combinations thereof. Client device 102 includes application 114.
[0035] exist Figure 1 In the example shown, the client device 102 is communicatively coupled to the network 104 and the wireless communication system 106. For example, the client device 102 can communicate with the wireless communication system 106 via an example client device communication link 116. The client device 102 is configured to communicate with one or more of the end-user device 108, the media producer 110, the central facility 112, and / or any other device configured to communicate via the network 104 and / or the wireless communication system 106.
[0036] exist Figure 1 In the example shown, the client device 102 can collect and / or otherwise obtain audio signals (e.g., audio samples). For example, the client device 102 can collect audio signals sent from the media producer 110 (e.g., over the radio) and / or audio signals corresponding to music played by the media producer 110 (e.g., music played by DJs at nightclubs, dance parties, and other events). The client device 102 can be configured to execute the application 114 to generate one or more fingerprints and / or signatures and queries including guessed pitch shifts, time shifts, and / or mixes of the audio signal.
[0037] exist Figure 1 In the example shown, the network 104 is the Internet. In other examples, the network 104 can be implemented using any suitable wired and / or wireless network, including, for example, one or more data buses, one or more local area networks (LANs), one or more wireless LANs, one or more cellular networks, one or more private networks, one or more public networks, etc. The network 104 is coupled to the client device 102, the wireless communication system 106, the media producer 110, and the central facility 112. The network 104 is additionally coupled to the client device 102 via the client device communication link 116 and the wireless communication system 106. The network 104 is also coupled to the end-user device 108 via the example end-user device communication link 118 and the wireless communication system 106.
[0038] exist Figure 1In the example of , the example network 104 enables one or more of the client device 102, the end-user device 108, the media producer 110, and the central facility 112 to communicate with one or more of the client device 102, the end-user device 108, the media producer 110, and the central facility 112. As used herein, the phrase "communicating" (including variations thereof) encompasses direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but otherwise includes selective communication at regular or irregular intervals as well as one-time events.
[0039] exist Figure 1 In the example shown, the end-user device 108 is a cellular phone. For example, the end-user device 108 can be a cellular phone for personal and / or professional use. In additional or alternative examples, the end-user device 108 can be any number of desktop computers, laptop computers, workstations, mobile phones, tablet computers, servers, any suitable computing devices, or combinations thereof. The end-user device 108 includes an application 114.
[0040] exist Figure 1 In the example shown, the end-user device 108 is communicatively coupled to the wireless communication system 106. For example, the end-user device 108 can communicate with the wireless communication system 106 via an end-user device communication link 118. The end-user device 108 is configured to communicate with one or more of the client device 102, the media producer 110, the central facility 112, and / or any other device configured to communicate via the network 104 and / or the wireless communication system 106.
[0041] exist Figure 1 In the example shown, the end-user device 108 can collect and / or otherwise obtain audio signals (e.g., audio samples). For example, the end-user device 108 can collect audio signals sent from the media producer 110 (e.g., via radio) and / or audio signals corresponding to music played by the media producer (e.g., music played by DJs in nightclubs, dance parties, and other events). The end-user device 108 can be configured to execute the application 114 to generate one or more fingerprints and / or signatures and queries including guessed pitch shifts, time shifts, and / or mixes of the audio signal. In some examples, the end-user device 108 can be implemented as a client device 102. In additional or alternative examples, the client device 102 can be implemented as an end-user device 108.
[0042] exist Figure 1In the example shown, media producer 110 is an entity that produces one or more media forms (e.g., audio signals, video signals, etc.). For example, media producer 110 may be a radio studio, a television studio, a recording studio, a DJ, and / or any other media production entity. Media producer 110 may adjust and / or otherwise manipulate the media before and / or during playback. For example, media producer 110 may change the pitch and / or playback speed of the media. For example, media producer 110 may adjust the pitch of the audio by up to 12%. In other examples, media producer 110 may adjust the pitch of the audio by 6%. In addition, media producer 110 may resample the audio to change the pitch and / or playback speed of the audio.
[0043] exist Figure 1 In an example of the invention, the central facility 112 is a server that collects and processes media from the client devices 102, the end-user devices 108, and / or the media producers 110 to generate metrics and / or other reports related to audio signals included in the media received from one or more of the client devices 102, the end-user devices 108, and the media producers 110. For example, the central facility 112 may receive one or more fingerprints and / or signatures and / or a query including guessed pitch shifts, time shifts, and / or resampling rates from the client devices 102 and / or the end-user devices 108. The query may indicate a list of pitch shifts, time shifts, and / or resampling rates that a user of the client device 102 and / or the end-user device 108 guesses may correspond to the one or more fingerprints and / or signatures received from the client device 102 and / or the end-user device 108.
[0044] In such examples, the central facility 112 may adjust one or more fingerprints and / or signatures from the client device 102 and / or the end-user device 108, and / or may adjust one or more reference fingerprints and / or reference signatures to identify whether one or more of the conjectured pitch shift, time shift, and / or resampling rate are applied to the audio. For example, the central facility 112 may adjust one or more fingerprints and / or signatures from the client device 102 and / or the end-user device 108 to determine whether one or more fingerprints and / or signatures from the client device 102 and / or the end-user device 108 match one or more reference fingerprints and / or signatures. In some examples, the central facility 112 may adjust one or more reference fingerprints and / or signatures to determine whether one or more reference fingerprints and / or signatures match one or more fingerprints and / or signatures from the client device 102 and / or the end-user device 108. In some examples, the central facility 112 may process queries serially. For example, the central facility 112 may test each guessed pitch offset included in the query until a match is found, then test each guessed time offset included in the query until a match is found, and then test each guessed resampling rate included in the query until a match is found. In additional or alternative examples, the central facility 112 may process the queries in parallel. For example, the central facility 112 may process each guessed pitch offset, each guessed time offset, and each guessed resampling rate included in the query simultaneously or at substantially similar times.
[0045] In such an example, after processing fingerprints and / or signatures and queries received from client device 102 and / or end-user device 108, central facility 112 may generate a report indicating: (a) audio signals that match the audio signal associated with the query, and (b) which, if any, of the guessed pitch shifts, guessed time shifts, and / or guessed resampling rates correspond to any reference fingerprints and / or reference signatures of central facility 112 when applied to one or more fingerprints and / or signatures from client device 102 and / or end-user device 108. In additional or alternative examples, the report may indicate: (a) audio signals that match the audio signal associated with the query, and (b) which, if any, of the guessed pitch shifts, guessed time shifts, and / or guessed resampling rates that, when applied to one or more reference fingerprints and / or reference signatures from the central facility 112, cause the one or more reference fingerprints and / or reference signatures to match one or more fingerprints and / or signatures from the client device 102 and / or end-user device 108 correspond to any of the reference fingerprints and / or reference signatures.
[0046] In an additional or alternative example, the central facility 112 may receive media (e.g., an audio signal) from the media producer 110 to generate one or more fingerprints and / or signatures. In such an example, the central facility 112 may analyze the media (e.g., an audio signal) to generate one or more sample fingerprints and / or sample signatures. In addition, the central facility 112 may adjust one or more sample fingerprints and / or sample signatures to accommodate one or more pitch shifts, one or more time shifts, and / or one or more resampling rates. In some examples, the central facility 112 may adjust one or more reference fingerprints and / or reference signatures to accommodate one or more pitch shifts, one or more time shifts, and / or one or more resampling rates. In addition, for each sample fingerprint and / or sample signature generated, the central facility 112 may compare the adjusted sample fingerprint and / or sample signature generated from the media received from the media producer 110 with one or more reference fingerprints and / or reference signatures. The central facility 112 may identify those adjusted sample fingerprints and / or sample signatures generated from media received from the media producer 110 that match the reference fingerprints and / or reference signatures, as well as the pitch shift, time shift, and / or resampling rate of the adjusted sample fingerprints and / or sample signatures.
[0047] In some examples, the central facility 112 may adjust one or more reference fingerprints and / or reference signatures to accommodate one or more pitch shifts, one or more time shifts, and / or one or more resampling rates. In addition, for each sample fingerprint and / or sample signature generated, the central facility 112 may compare the adjusted reference fingerprint and / or reference signature with one or more sample fingerprints and / or sample signatures generated from the media received from the media producer 110. The central facility 112 may identify those adjusted reference fingerprints and / or reference signatures generated by the central facility 112 that match the sample fingerprints and / or sample signatures, as well as the pitch shifts, time shifts, and / or resampling rates of the adjusted reference fingerprints and / or reference signatures.
[0048] After a threshold time period has elapsed, the central facility 112 may process these matches and / or corresponding pitch offsets, time offsets, and / or resampling rates to generate a report that compares the frequency of occurrence of the various pitch offsets, time offsets, and / or resampling rates. In additional or alternative examples, the central facility 112 may process these matches and / or corresponding pitch offsets, time offsets, and / or resampling rates continuously. In some examples, the central facility 112 may process these matches and / or corresponding pitch offsets, time offsets, and / or resampling rates after a threshold amount of data has been received.
[0049] To process these matches and / or corresponding pitch shifts, time shifts, and / or resampling rates, the central facility 112 may, for example, generate one or more histograms that identify the frequency of occurrence of each pitch shift, the frequency of occurrence of each time shift, and / or the frequency of occurrence of each resampling rate (e.g., one or more frequencies of occurrence). For example, the central facility 112 may generate a report that includes (a) one or more pitch shift values, (b) one or more time shift values, or (c) one or more resampling rates that have a higher frequency of occurrence than (a) one or more other pitch shift values, (b) one or more other time shift values, or (c) one or more other resampling rates, respectively.
[0050] In additional or alternative examples, the central facility 112 may generate one or a suitable graphical analysis tool to identify the frequency of occurrence of each pitch shift, time offset, and / or resampling rate. For example, the central facility 112 may generate a timeline that identifies the frequency of occurrence of each pitch shift, time offset, and / or resampling rate over time. In such an example, the timeline may facilitate the types of pitch shifts, time offsets, and / or resampling rates used over a period of time.
[0051] exist Figure 1 In the example shown, the central facility 112 can receive and / or obtain Internet messages (e.g., Hypertext Transfer Protocol (HTTP) requests) that include media (e.g., audio signals), fingerprints, signatures, and / or queries. Additionally or alternatively, any other method can be used to receive and / or obtain metering information, such as HTTP Secure Protocol (HTTPS), File Transfer Protocol (FTP), Secure File Transfer Protocol (SFTP), etc.
[0052] In some examples, the central facility 112 may adjust the reference fingerprints in order to match the reference fingerprints with the sample fingerprints. For example, the central facility 112 may apply a pitch shift, a time shift, or a resampling rate to the reference fingerprints in order to add the conjectured pitch shift, time shift, and / or resampling rate so that the central facility 112 can compare and / or match the reference fingerprints with the sample fingerprints. In such examples, the central facility 112 generates the adjusted fingerprints as (a) one or more pitch-shifted reference fingerprints, (b) one or more time-shifted reference fingerprints, and / or (c) one or more resampled reference fingerprints.
[0053] In such an example, if the query indicates that the client and / or user guesses that the reference fingerprint corresponds to an audio signal that has been pitch-shifted up, the central facility 112 may increase the bin value associated with (e.g., of, in, etc.) the reference fingerprint based on (e.g., by, etc.) the guessed pitch shift value. If the query indicates that the client and / or user guesses that the reference fingerprint corresponds to an audio signal that has been pitch-shifted down, the central facility 112 may decrease the bin value associated with (e.g., of, in, etc.) the reference fingerprint based on (e.g., by, etc.) the guessed pitch shift value.
[0054] In an additional or alternative example, if the query indicates that the client and / or user guesses that the reference fingerprint corresponds to an audio signal that has been time-shifted up, the central facility 112 may delete frames of the reference fingerprint at positions in the reference fingerprint that correspond to the time offset value. If the query indicates that the client and / or user guesses that the reference fingerprint corresponds to an audio signal that has been time-shifted down, the central facility 112 may duplicate frames of the reference fingerprint at positions in the reference fingerprint that correspond to the time offset value.
[0055] exist Figure 1 In an example of , the application 114 can be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, graphics processing units (GPUs), digital signal processors (DSPs), application specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field programmable logic devices (FPLDs). The application 114 is configured to obtain one or more audio signals (e.g., audio signals from the media producer 110 and / or ambient audio from a microphone), generate one or more sample fingerprints and / or sample signatures from the one or more audio signals, and send one or more sample fingerprints and / or sample signatures and adjustment instructions (e.g., including queries for one or more pitch shifts, time shifts, and / or resampling rates that are conjectured to be implemented in one or more audio signals) in response to one or more instructions from a user of the client device 102 and / or the end-user device 108. In some examples, the query indicates whether the query is to be processed serially or in parallel. The application 114 is also configured to receive responses to the query, and when one or more of the sample fingerprints and / or sample signatures match one or more of the pitch shifts, time shifts, and / or resampling rates, the application 114 can generate a report identifying: (a) the audio signal that matches the audio signal associated with the query, and (b) which of the guessed pitch shifts, guessed time shifts, and / or guessed resampling rates matches the audio signal associated with the query. The example application 114 can display the report to the user (e.g., via a user interface of the end-user device 108), store the report locally, and / or send the report to the example client device 102.
[0056] In some examples, the application 114 can implement the functionality of the central facility 112. In additional or alternative examples, the central facility 112 can implement the functionality of the application 114. In some examples, the functionality of the central facility 112 and the functionality of the application 114 can be distributed between the central facility 112 and the application 114 in a manner suitable for the application. For example, the central facility 112 can send the application 114 from the central facility 112 to the client device 102 and / or the end-user device 108 for installation on the client device 102 and / or the end-user device 108.
[0057] In some examples, client devices 102 and end-user devices 108 may not be able to send information to central facility 112 via network 104. For example, a server upstream of client device 102 and / or end-user device 108 may not provide functional routing capabilities to central facility 112. Figure 1 In the example shown, client device 102 includes the additional capability to send information through wireless communication system 106 (e.g., a cellular communication system) via client device communication link 116. End-user device 108 includes the additional capability to send information through wireless communication system 106 via end-user device communication link 118.
[0058] Figure 1 The client device communication link 116 and the end user device communication link 118 of the illustrated example are cellular communication links. However, any other communication method and / or communication system may be used in addition or alternatively, such as an Ethernet connection, a Bluetooth connection, a Wi-Fi connection, etc. In addition, Figure 1 The client device communication link 116 and the end-user device communication link 118 implement a cellular connection via the Global System for Mobile Communications (GSM). However, any other system and / or protocol for communication may be used, such as Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), Worldwide Interoperability for Microwave Access (WiMAX), Long Term Evolution (LTE), etc.
[0059] Figure 2 It is shown Figure 1 1 is a block diagram of further details of an example central facility 112. The example central facility 112 includes an example network interface 202, an example fingerprint pitch tuner 204, an example fingerprint velocity tuner 206, an example query comparator 208, an example fingerprint generator 210, an example media processor 212, an example report generator 214, and an example database 216.
[0060] exist Figure 2 In the example of Figure 1The network 104 of the present invention obtains information and / or sends information to the network 104. The network interface 202 implements a web server that receives fingerprints, signatures, media (e.g., audio signals), queries and / or other information from one or more of the client device 102, the end-user device 108, or the media producer 110. The fingerprints, signatures, media (e.g., audio signals), queries and / or other information can be formatted as HTTP messages. However, any other message format and / or protocol can be used in addition or alternatively, such as FTP, SMTP, HTTPS protocol, etc. In addition or alternatively, the media (e.g., audio signals) can be received as a radio waveform, an MP3 file, an MP4 file, and / or any other suitable audio format.
[0061] exist Figure 2 In the example shown, the network interface 202 is configured to obtain one or more fingerprints, signatures, queries and / or media from a device (e.g., a client device 102, an end-user device 108, a media producer 110). The network interface 202 may also be configured to identify when to choose to process a query in parallel or in series. In addition, the network interface 202 is also configured to identify whether the query includes a guessed pitch offset, a guessed time offset and / or a guessed resampling rate. When the query includes one or more guessed pitch offsets, one or more guessed time offsets and / or one or more guessed resampling rates, the network interface 202 may select: (a) one of the one or more guessed pitch offsets for processing by the central facility 112, (b) one of the one or more guessed time offsets for processing by the central facility 112, and / or (c) one of the one or more guessed resampling rates for processing by the central facility 112.
[0062] exist Figure 2 In the example of , the network interface 202 can also be configured to identify whether there are additional guessed pitch offsets in the query, whether there are additional guessed time offsets in the query, and / or whether there are additional guessed resampling rates in the query. In addition, the network interface 202 can identify whether the central facility 112 has received and / or otherwise obtained additional queries from other devices in the network 104 (e.g., client devices 102, end-user devices 108, etc.). The network interface 202 can also send reports and / or other information to devices in the network 104 (e.g., client devices 102, end-user devices 108, etc.).
[0063] In some examples, the example network interface 202 implements an example means for interfacing. The interfacing means is implemented by, for example, at least Fig. 6AIn addition or alternatively, the interface means may be implemented by executable instructions such as executable instructions implemented by blocks 602, 604, 606, 608, 610, 616, 620, 622, 628, 632, 634, 642, 646 and 650 of FIG. Figure 6B The interface means may be implemented by executable instructions such as executable instructions implemented by blocks 652, 654, 662, 664, 674, 676, 684, and 688 of FIG. In some examples, the interface means may be implemented by at least Figure 7 The method may be implemented by executable instructions such as the executable instructions implemented by blocks 702, 732, 736, and 744 of FIG. Fig. 6A , Figure 6B and Figure 7 The executable instructions of blocks 602, 604, 606, 608, 610, 616, 620, 622, 628, 632, 634, 642, 646, 650, 652, 654, 662, 664, 674, 676, 684, 688, 702, 732, 736, and 744 may be in a computer such as Fig.11 The interface means is executed on at least one processor of the example processor 1112. In other examples, the interface means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0064] exist Figure 2 In the example of , the fingerprint pitch tuner 204 is a device that can adjust the pitch of one or more sample fingerprints and / or sample signatures and / or one or more reference fingerprints and / or reference signatures. For example, the fingerprint and / or signature can be represented as a spectrogram of the audio signal. The spectrogram may include one or more bins with corresponding bin values. In order to adjust the pitch of the sample fingerprint and / or sample signature, the fingerprint pitch tuner 204 can obtain one or more sample fingerprints and / or sample signatures from the database 216 and / or from the network 104 via the network interface 202. In addition, the fingerprint pitch tuner 204 can identify whether the guessed pitch shift identified by the network interface 202 increases or decreases the pitch of the audio signal corresponding to the sample fingerprint.
[0065] exist Figure 2In the example shown, if the guessed pitch shift increases the pitch of the audio signal, the fingerprint pitch tuner 204 can reduce the bin value associated with the sample fingerprint and / or sample signature (e.g., of the sample fingerprint and / or sample signature, in the sample fingerprint and / or sample signature, etc.) based on (e.g., by, etc.) the guessed pitch shift. For example, if the guessed pitch shift is a 5% pitch increase, the fingerprint pitch tuner 204 can multiply the bin value of the sample fingerprint and / or sample signature by 95% (e.g., 0.95). If the guessed pitch shift decreases the pitch of the audio signal, the fingerprint pitch tuner 204 can increase the bin value associated with the sample fingerprint and / or sample signature (e.g., of the sample fingerprint and / or sample signature, in the sample fingerprint and / or sample signature, etc.) based on (e.g., by, etc.) the guessed pitch shift. For example, if the guessed pitch shift is a 5% pitch reduction, the fingerprint pitch tuner 204 may multiply the bin value of the sample fingerprint and / or sample signature by 105% (e.g., 1.05). Additionally, in response to a query including one or more resampling rates, at least one of the fingerprint pitch tuner 204 or the fingerprint velocity tuner 206 may generate an adjusted sample fingerprint by applying the resampling rate to the sample fingerprint.
[0066] In some examples, when adjusting the bin value associated with the sample fingerprint, the fingerprint pitch tuner 204 can round to the nearest bin value. For example, the sample fingerprint can include data values of 1 in bin values 100, 250, 372, 491, 522, 633, 725, 871, 905, and 910. If the fingerprint pitch tuner 204 applies a hypothetical pitch shift that increases the pitch of the audio signal corresponding to the sample fingerprint by 10% to the sample fingerprint, the adjusted bin values can be 90, 225, 334.8, 441.9, 469.8, 569.7, 652.5, 783.9, 814.5, and 819. In such an example, the fingerprint pitch tuner 204 may round up the bin values such that the pitch shifted sample fingerprint contains data values of 1 at bin values 90, 225, 335, 442, 470, 570, 653, 784, 815, and 819.
[0067] In additional or alternative examples, the sample fingerprint may include data values of 1 at bin values 100, 250, 372, 491, 522, 633, 725, 871, 905, and 910. If the fingerprint pitch tuner 204 applies a guessed pitch shift to the sample fingerprint that lowers the pitch of the audio signal corresponding to the sample fingerprint by 10%, the adjusted bin values may be 110, 275, 409.2, 540.1, 574.2, 696.3, 797.5, 958.1, 995.5, and 1,001. In such an example, the fingerprint pitch tuner 204 may round up the bin values so that the pitch-shifted sample fingerprint includes data values of 1 at bin values 110, 275, 409, 540, 574, 696, 798, 958, 996, and 1,001.
[0068] In some examples, the central facility 112 may adjust the reference fingerprints in order to match the reference fingerprints with the sample fingerprints. For example, the central facility 112 may apply a pitch shift, a time shift, or a resampling rate to the reference fingerprints in order to add the conjectured pitch shift, time shift, and / or resampling rate so that the central facility 112 can compare and / or match the reference fingerprints with the sample fingerprints. In such examples, the central facility 112 generates the adjusted fingerprints as (a) one or more pitch-shifted reference fingerprints, (b) one or more time-shifted reference fingerprints, and / or (c) one or more resampled reference fingerprints.
[0069] For example, if the guessed pitch shift increases the pitch of the audio signal, the fingerprint pitch tuner 204 can increase the bin value associated with the reference fingerprint and / or reference signature (e.g., of the reference fingerprint and / or reference signature, in the reference fingerprint and / or reference signature, etc.) based on (e.g., by, etc.) the guessed pitch shift. For example, if the guessed pitch shift is a 5% pitch increase, the fingerprint pitch tuner 204 can multiply the bin value of the reference fingerprint and / or reference signature by 105% (e.g., 1.05). If the guessed pitch shift decreases the pitch of the audio signal, the fingerprint pitch tuner 204 can decrease the bin value associated with the reference fingerprint and / or reference signature (e.g., of the reference fingerprint and / or reference signature, in the reference fingerprint and / or reference signature, etc.) based on (e.g., by, etc.) the guessed pitch shift. For example, if the guessed pitch shift is a 5% pitch reduction, the fingerprint pitch tuner 204 may multiply the bin values of the reference fingerprint and / or reference signature by 95% (e.g., 0.95). In an additional or alternative example, if the query indicates that the client and / or user guesses that the sample fingerprint corresponds to an audio signal that has been time-shifted, the central facility 112 may delete frames of the reference fingerprint at one or more positions in the reference fingerprint that correspond to the time shift values.
[0070] In some examples, the example fingerprint pitch tuner 204 implements an example means for pitch tuning. The pitch tuning means is implemented by, for example, at least Fig. 6A Additionally or alternatively, the pitch tuning means is implemented by, for example, at least Figure 6B In some examples, the pitch tuning means is implemented by, for example, at least Figure 7 In some examples, the pitch tuning means is implemented by, for example, at least Figure 8 The method may be implemented by executable instructions such as the executable instructions implemented by blocks 802, 804, 806, 808, and 810 of FIG. Fig. 6A , Figure 6B , Figure 7 and Figure 8 The executable instructions of blocks 612, 636, 656, 666, 708, 720, 802, 804, 806, 808, and 810 may be in a computer such as Fig.11 The method is executed on at least one processor of the example processor 1112. In other examples, the pitch tuning means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0071] exist Figure 2 In the example shown, the fingerprint speed tuner 206 is a device that can adjust the playback speed of one or more sample fingerprints and / or sample signatures and / or one or more reference fingerprints and / or reference signatures. For example, the fingerprint and / or signature can be represented as a spectrogram of an audio signal. The spectrogram may include one or more frames that can be called sub-fingerprints. The spectrogram may include a predefined number of sub-fingerprints for a given audio sample length. The examples disclosed herein identify time-shifted media based on the spacing of sub-fingerprints in fingerprints and / or signatures. In order to adjust the playback speed of sample fingerprints and / or sample signatures, the fingerprint speed tuner 206 can obtain one or more sample fingerprints and / or sample signatures from the database 216 and / or from the network 104 via the network interface 202. In addition, the fingerprint speed tuner 206 can identify whether the conjectured time offset identified by the network interface 202 increases or decreases the playback speed of the audio signal corresponding to the sample fingerprint.
[0072] exist Figure 2In the example shown, if the conjectured time offset increases the playback speed of the audio signal, the fingerprint speed tuner 206 can modify the sample fingerprint and / or sample signature based on the increased playback speed. For example, if the conjectured time offset increases the playback speed, the fingerprint speed tuner 206 can add a sub-fingerprint to the sample fingerprint and / or sample signature at one or more positions in the sample fingerprint and / or sample signature corresponding to 1 divided by the percentage of the time offset. For example, if the conjectured time offset corresponds to a 5% increase in playback speed, the fingerprint speed tuner 206 can add (e.g., copy) one or more frames to the sample fingerprint and / or sample signature for the length of the sample fingerprint and / or sample signature at every 20 frames of the sample fingerprint and / or sample signature (e.g., 100% ÷ 5%) (e.g., copy the 20th, 40th, 60th, 80th, 100th, 120th sub-fingerprint of the sample fingerprint and / or sample signature, etc.). In such an example, if the sample fingerprint and / or sample signature with a 5% playback speed increase is assumed to initially include 400 frames, the time-shifted sample fingerprint and / or time-shifted sample signature may include 420 frames (e.g., 400+(400*0.05)).
[0073] In some examples, the fingerprint speed tuner 206 can round to the nearest frame when adapting to a conjectured pitch shift. For example, a sample fingerprint can include frames (e.g., sub-fingerprints) 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10. If the fingerprint speed tuner 206 applies a conjectured time shift to the sample fingerprint corresponding to a 10% increase in the playback time of the audio signal of the sample fingerprint, the fingerprint speed tuner 206 can multiply the number of frames by a percentage based on the conjectured time shift. For example, a 10% increase in playback speed can correspond to a multiplication factor of 1.1. When the fingerprint speed tuner 206 applies such a time shift, the adjusted frame values can be 1.1, 2.2, 3.3, 4.4, 5.5, 6.6, 7.7, 8.8, 9.9, and 11. In such an example, the fingerprint speed tuner 206 may round up the frames so that the time-shifted sample fingerprint includes frames 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and 11. In such an example, frame 5 may be duplicated. In some examples, the fingerprint speed tuner 206 may generate frame 5 by interpolating between frames 4 and 6.
[0074] In an additional or alternative example, a playback speed increase of 10% may correspond to a multiplication factor of 1.1. When the fingerprint speed tuner 206 applies such a time shift, the adjusted frame values may be 1.1, 2.2, 3.3, 4.4, 5.5, 6.6, 7.7, 8.8, 9.9, and 11. In such an example, the fingerprint speed tuner 206 may round the frames up so that the time-shifted sample fingerprint includes frames with data at frames 1, 2, 3, 4, 5, 7, 8, 9, 10, and 11.
[0075] exist Figure 2 In the example of , if the conjectured time offset reduces the playback speed of the audio signal, the fingerprint speed tuner 206 can modify the sample fingerprint and / or sample signature based on the reduced playback speed. For example, if the conjectured time offset reduces the playback speed, the fingerprint speed tuner 206 can remove sub-fingerprints from the sample fingerprint and / or sample signature at one or more positions in the sample fingerprint and / or sample signature corresponding to 1 divided by the percentage of the time offset. For example, if the conjectured time offset corresponds to a 5% reduction in playback speed, the fingerprint speed tuner 206 can remove (e.g., delete) one or more frames from the sample fingerprint and / or sample for the length of the sample fingerprint and / or sample signature (e.g., remove the 20th, 40th, 60th, 80th, 100th, 120th sub-fingerprint from the sample fingerprint and / or sample, etc.) at every 20 frames (e.g., 100% ÷ 5%) of the sample fingerprint and / or sample signature. In such an example, if the sample fingerprint and / or sample signature at a guessed playback speed reduced by 5% initially includes 400 frames, the time-shifted sample fingerprint and / or the time-shifted sample signature may include 380 frames (e.g., 400-(400*0.05)). Additionally, in response to a query including one or more resampling rates, at least one of the fingerprint pitch tuner 204 or the fingerprint speed tuner 206 may generate an adjusted sample fingerprint by applying the resampling rate to the sample fingerprint.
[0076] In some examples, the fingerprint speed tuner 206 may round up to the nearest frame when adapting to the conjectured pitch shift. For example, a sample fingerprint may include frames (e.g., sub-fingerprints) 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10. If the fingerprint speed tuner 206 applies a conjectured time shift corresponding to a 10% reduction in the playback time of the audio signal of the sample fingerprint to the sample fingerprint, the fingerprint speed tuner 206 may multiply the number of frames by a percentage based on the conjectured time shift. For example, a 10% reduction in playback speed may correspond to a multiplication factor of 0.9. When the fingerprint speed tuner 206 applies such a time shift, the adjusted frame values may be 0.9, 1.8, 2.7, 3.6, 4.5, 5.4, 6.3, 7.2, 8.1, and 9.1. In such an example, the fingerprint speed tuner 206 may round up the frames so that the time-shifted sample fingerprint includes frames 1, 2, 3, 4, 5, 6, 7, 8, and 9. In such an example, frame 5 of the original sample fingerprint may be deleted because, upon adjustment, frame 4 of the original sample fingerprint now occupies frame 5 of the time-shifted sample fingerprint.
[0077] In some examples, the central facility 112 may adjust the reference fingerprints in order to match the reference fingerprints with the sample fingerprints. For example, the central facility 112 may apply a pitch shift, a time shift, or a resampling rate to the reference fingerprints in order to add the conjectured pitch shift, time shift, and / or resampling rate so that the central facility 112 can compare and / or match the reference fingerprints with the sample fingerprints (e.g., the sample media fingerprints). In such examples, the central facility 112 generates the adjusted fingerprints as (a) one or more pitch-shifted reference fingerprints, (b) one or more time-shifted reference fingerprints, and / or (c) one or more resampled reference fingerprints.
[0078] exist Figure 2 In the example shown, if the conjectured time offset increases the playback speed of the audio signal, the fingerprint speed tuner 206 can modify the reference fingerprint and / or reference signature based on the increased playback speed. For example, if the conjectured time offset increases the playback speed, the fingerprint speed tuner 206 can remove (e.g., delete) sub-fingerprints from the reference fingerprint and / or reference signature at one or more positions in the reference fingerprint and / or reference signature corresponding to 1 divided by the percentage of the time offset. For example, if the conjectured time offset corresponds to a 5% increase in playback speed, the fingerprint speed tuner 206 can remove (e.g., delete) one or more frames from the reference fingerprint and / or reference signature for the length of the reference fingerprint and / or reference signature (e.g., remove the 20th sub-fingerprint from the reference fingerprint and / or reference signature) at every 20 frames (e.g., 100% ÷ 5%) of the reference fingerprint and / or reference signature.
[0079] exist Figure 2 In the example of , if the conjectured time offset reduces the playback speed of the audio signal, the fingerprint speed tuner 206 can modify the reference fingerprint and / or reference signature based on the reduced playback speed. For example, if the conjectured time offset reduces the playback speed, the fingerprint speed tuner 206 can repeat (e.g., copy) a sub-fingerprint of the reference fingerprint and / or reference signature at one or more positions in the reference fingerprint and / or reference signature corresponding to 1 divided by a percentage of the time offset. For example, if the conjectured time offset corresponds to a 5% reduction in playback speed, the fingerprint speed tuner 206 can repeat (e.g., copy) one or more frames of the reference fingerprint and / or reference signature (e.g., remove the 20th sub-fingerprint from the reference fingerprint and / or reference signature) at every 20 frames of the reference fingerprint and / or reference signature (e.g., 100% ÷ 5%) for the length of the reference fingerprint and / or reference signature.
[0080] In some examples, the example fingerprint speed tuner 206 implements an example means for speed tuning. The speed tuning means includes at least Fig. 6A Additionally or alternatively, the speed tuning means is implemented by, for example, at least Figure 6B In some examples, the speed tuning means is implemented by, for example, at least Figure 7 In some examples, the speed tuning means is implemented by, for example, at least Fig. 9 The method may be implemented by executable instructions such as the executable instructions implemented by blocks 902, 904, 906, 908, and 910 of FIG. Fig. 6A , Figure 6B , Figure 7 and Fig. 9 The executable instructions of blocks 624, 638, 668, 678, 714, 722, 902, 904, 906, 908, and 910 may be in a computer such as Fig.11 The speed tuning means is executed on at least one processor of the example processor 1112. In other examples, the speed tuning means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0081] In additional or alternative examples, the central facility 112 may include any number of fingerprint tuners (eg, one fingerprint tuner, a plurality of fingerprint tuners, etc.) that implement the functionality of the fingerprint pitch tuner 204 and the fingerprint velocity tuner 206 .
[0082] exist Figure 2In the example shown, the query comparator 208 is a device that compares the adjusted sample fingerprint and / or the adjusted sample signature with a reference fingerprint and / or a reference signature and / or compares the adjusted reference fingerprint and / or the adjusted reference signature with the sample fingerprint and / or the sample signature. In addition or alternatively, the query comparator 208 may compare one or more pitch-shifted sample fingerprints with a reference fingerprint. In response to determining that the pitch-shifted sample fingerprint matches the reference fingerprint, the query comparator 208 may indicate a guessed pitch offset that matches the reference fingerprint. In some examples, the query comparator 208 may compare one or more time-shifted sample fingerprints with a reference fingerprint, and in response to determining that the time-shifted sample fingerprint matches the reference fingerprint, the query comparator 208 may indicate a guessed time offset that matches the reference fingerprint.
[0083] In some examples, the query comparator 208 may compare one or more pitch-shifted reference fingerprints to the sample fingerprint. In response to determining that the pitch-shifted reference fingerprint matches the sample fingerprint, the query comparator 208 may indicate a guessed pitch shift that matches the sample fingerprint. In some examples, the query comparator 208 may compare one or more time-shifted sample fingerprints to the sample fingerprint, and in response to determining that the time-shifted reference fingerprint matches the sample fingerprint, the query comparator 208 may indicate a guessed time shift that matches the sample fingerprint.
[0084] In some examples, the fingerprint pitch tuner 204 and the fingerprint speed tuner 206 can be used in combination to adapt to a conjectured resampling rate. For example, the resampling rate can indicate a pitch offset value and / or a time offset value. In such an example, the query comparator 208 can compare one or more resampled sample fingerprints and / or resampled reference fingerprints with a reference fingerprint and / or a sample fingerprint, respectively. In response to determining that a resampled sample fingerprint (e.g., a resampled sample media fingerprint) matches a reference fingerprint and / or a resampled reference fingerprint matches a sample fingerprint, the query comparator 208 can indicate a conjectured resampling rate that matches the reference fingerprint and / or the sample fingerprint, respectively. For example, the query comparator 208 can indicate a pitch offset and a time offset of a resampling rate that matches a reference fingerprint.
[0085] In some examples, the example query comparator 208 implements an example means for comparison. The comparison means includes at least Fig. 6A In addition or alternatively, the comparison means is implemented by, for example, at least Figure 6BThe comparison means may be implemented by executable instructions such as executable instructions implemented by blocks 658, 660, 670, 672, 680 and 682 of FIG. Figure 7 The method may be implemented by executable instructions such as the executable instructions implemented by blocks 710, 712, 716, 718, 724, and 726 of FIG. Fig. 6A , Figure 6B and Figure 7 The executable instructions of blocks 614, 618, 626, 630, 640, 644, 658, 660, 670, 672, 680, 682, 710, 712, 716, 718, 724, and 726 may be in a computer such as Fig.11 The example processor 1112 is executed on at least one processor of the example processor 1112. In other examples, the comparison means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0086] exist Figure 2 In the example of , the fingerprint generator 210 is a device that can generate one or more sample fingerprints and / or one or more sample signatures from a sampled medium (e.g., an audio signal). For example, the fingerprint generator 210 can separate the audio signal (e.g., a digitized audio signal) into time-frequency bins and / or audio signal frequency components. For example, the fingerprint generator 210 can perform a fast Fourier transform (FFT) on the audio signal to transform the audio signal into the frequency domain.
[0087] Additionally, the example fingerprint generator 210 may divide the transformed audio signal into two or more frequency bins (e.g., using a Hamming function, a Hann function, etc.). In this example, each audio signal frequency component is associated with a frequency bin in the two or more frequency bins. Additionally or alternatively, the fingerprint generator 210 may aggregate the audio signal into one or more time periods (e.g., the duration of the audio, a six-second period, a 1-second period, etc.). In other examples, the fingerprint generator 210 may use any suitable technique to transform the audio signal (e.g., discrete Fourier transform, sliding time window Fourier transform, wavelet transform, discrete Hadamard transform, discrete Walsh Hadamard, constant Q transform, discrete cosine transform, etc.). In some examples, the fingerprint generator 210 may include one or more bandpass filters (BPF). In some examples, the processed audio signal may be represented by a spectrogram. The following is in conjunction with FIG. 4A to FIG. 4B The process of the fingerprint generator 210 that generates the spectrogram is discussed.
[0088] exist Figure 2In an example of the present invention, the fingerprint generator 210 may determine an audio characteristic of a portion of the audio signal (e.g., an audio signal frequency component, an audio region around a time-frequency bin, etc.). For example, the fingerprint generator 210 may determine a mean energy (e.g., average power, etc.) of one or more of the audio signal frequency components. Additionally or alternatively, the fingerprint generator 210 may determine other characteristics of a portion of the audio signal (e.g., mode energy, median energy, mode power, median energy, mean energy, mean amplitude, etc.).
[0089] exist Figure 2 In the illustrated example, the fingerprint generator 210 normalizes one or more time-frequency bins by the associated audio characteristics of the surrounding audio area. For example, the fingerprint generator 210 may normalize the time-frequency bins by the mean energy of the surrounding audio area. In other examples, the fingerprint generator 210 normalizes some of the audio signal frequency components by the associated audio characteristics. For example, the fingerprint generator 210 may use the mean energy associated with the audio signal frequency component to normalize each time-frequency bin of the audio signal frequency component. In some examples, the processed bins (e.g., normalized time-frequency bins, normalized audio signal frequency components, etc.) may be represented as a spectrogram. The following is a flowchart of a method for performing a spectrogram analysis of the audio signal frequency components in combination with the method described below. Figure 4C The process of the fingerprint generator 210 that generates a normalized spectrogram is discussed.
[0090] exist Figure 2 In the illustrated example of , the fingerprint generator 210 can select one or more points from the normalized audio signal to be used to generate the fingerprint and / or signature. For example, the fingerprint generator 210 can select multiple energy maxima of the normalized audio signal. In other examples, the fingerprint generator 210 can select any other suitable point of the normalized audio. Figure 4D and Figure 5 The process of the fingerprint generator 210 to generate a fingerprint is discussed.
[0091] Additionally or alternatively, the fingerprint generator 210 may weight the selection of points based on the category of the audio signal. For example, if the category of the audio signal is music, the fingerprint generator 210 may focus the selection of points on common frequency ranges of music (e.g., bass, treble, etc.). In some examples, the fingerprint generator 210 may determine the category of the audio signal (e.g., music, speech, sound effects, advertising, etc.). The example fingerprint generator 210 uses the selected points to generate fingerprints and / or signatures. The example fingerprint generator 210 may generate fingerprints based on the selected points using any suitable method. An example method and apparatus for fingerprinting an audio signal via normalization is disclosed in U.S. patent application Ser. No. 16 / 453,654 to Coover et al., the entire contents of which are incorporated herein by reference.
[0092] In some examples, the example fingerprint generator 210 implements an example means for fingerprint generation. The fingerprint generation means includes at least Figure 7 The method may be implemented by executable instructions such as the executable instructions implemented by blocks 704 and 706 of FIG. Figure 7 The executable instructions of blocks 704 and 706 may be in a computer such as Fig.11 The fingerprint generation means is executed on at least one processor of the example processor 1112. In other examples, the fingerprint generation means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0093] exist Figure 2 In an example of, the media processor 212 is a device that processes one or more fingerprints, one or more signatures, and / or one or more indications of matching pitch offsets, matching time offsets, and / or matching resampling rates. For example, after a threshold time period has passed, the media processor 212 can process these matches and / or corresponding pitch offsets, time offsets, and / or resampling rates. For example, the media processor 212 can identify the frequency of occurrence of each pitch offset, time offset, and / or resampling rate. In additional or alternative examples, the media processor 212 can continuously process these matches and / or corresponding pitch offsets, time offsets, and / or resampling rates. In some examples, the media processor 212 can process these matches and / or corresponding pitch offsets, time offsets, and / or resampling rates after a threshold amount of data has been received.
[0094] To process these matching and / or corresponding pitch offsets, time offsets, and / or resampling rates, the media processor 212 may, for example, generate one or more histograms that identify the frequency of occurrence of each pitch offset, the frequency of occurrence of each time offset, and / or the frequency of occurrence of each resampling rate (e.g., one or more frequencies of occurrence). For example, the report generator 214 may generate a report including the following items: (a) one or more pitch offset values, (b) one or more time offset values, or (c) one or more resampling rates that have a higher frequency of occurrence than (a) one or more other pitch offset values, (b) one or more other time offset values, or (c) one or more other resampling rates, respectively, as identified by the media processor 212.
[0095] In additional or alternative examples, the media processor 212 can generate one or a suitable graphical analysis tool to identify the frequency of occurrence of each pitch shift, time offset, and / or resampling rate. For example, the media processor 212 can generate a timeline that identifies the frequency of occurrence of each pitch shift, time offset, and / or resampling rate over time. In such an example, the timeline can facilitate the types of pitch shifts, time offsets, and / or resampling rates used over a period of time.
[0096] In some examples, the example media processor 212 implements an example means for processing. The processing means includes at least Figure 7 The method may be implemented by executable instructions such as the executable instructions implemented by blocks 728, 730, and 742 of FIG. Figure 7 The executable instructions of blocks 728, 730, and 742 may be in a computer such as Fig.11 The processing means is executed on at least one processor of the example processor 1112. In other examples, the processing means is implemented by hardware logic, hardware-implemented state machines, logic circuits, and / or any other combination of hardware, software, and / or firmware.
[0097] exist Figure 2In the example shown, the report generator 214 is a device configured to generate and / or prepare a report. The report generator 214 prepares a report indicating one or more of commonly used pitch offsets, time offsets and / or resampling rates. In addition, the report generator 214 can suggest one or more pitch offsets, time offsets and / or resampling rates indicated in the query to the client and / or end user. In addition, the report generator 214 can identify one or more audio signals corresponding to one or more reference fingerprints and / or reference signatures that match one or more sample fingerprints. In some examples, the report generator 214 generates a report including one or more graphic analysis tools, which are used to identify the frequency of occurrence of each pitch offset, time offset and / or resampling rate. The report generated by the report generator 214 can include one or more pitch offset values, one or more time offset values and / or one or more time resampling rates with a higher frequency of occurrence than other pitch offset values, time offset values and / or resampling rates tested by the central facility 112.
[0098] exist Figure 2 In an example of , the report generator 214 can prepare a report that associates one or more reference fingerprints with one or more sample fingerprints. For example, the report generator 214 can prepare a report that identifies: (a) the song corresponding to the reference fingerprint, and (b) the pitch shift, time shift, and / or resampling rate used by the central facility 112 to match the sample fingerprint and / or sample signature to the reference fingerprint. For example, the report generator 214 can generate a report that indicates that the audio sample is a sample of "Take it Easy" by the Eagles, and indicates that the sample pitch is shifted down by 5% and the time shift increases the playback speed by 5% (e.g., a report indicating information associated with the reference fingerprint and adjustments (e.g., transformations, pitch shifts, time shifts, etc.)
[0099] In some examples, the example report generator 214 implements an example means for report generation. The report generation means generates reports by, for example, at least Fig. 6A Additionally or alternatively, the report generating means may be implemented by, for example, at least Figure 6B In some examples, the report generation means is implemented by, for example, at least Figure 7 The executable instructions implemented by blocks 734 and 740 may be implemented by executable instructions such as the executable instructions implemented by blocks 734 and 740. Fig. 6A , Figure 6B and Figure 7 The executable instructions of blocks 648, 686, 734 and 740 may be in a computer such as Fig.11The report generation means is executed on at least one processor of the example processor 1112. In other examples, the report generation means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0100] exist Figure 2 In the example shown, the database 216 is configured to record data (e.g., information obtained, messages generated, etc.). For example, the database 216 may store one or more files indicating one or more pitch shifts, time shifts, and / or resampling rates that cause one or more adjusted sample fingerprints to match one or more reference fingerprints. In addition, the database 216 may store one or more reports generated by the report generator 214, one or more sample fingerprints generated by the fingerprint generator 210, and / or one or more sample fingerprints and / or queries received at the network interface 202. In an additional or alternative example, the database 216 may store one or more adjusted reference fingerprints. In such an additional or alternative example, the central facility 112 may compare the sample fingerprint with the adjusted reference fingerprint to determine a match. The adjusted reference fingerprint and / or reference signature may correspond to the highest number (e.g., 40, 100, etc.) on certain radio stations during a specified time period, and / or correspond to certain television programs that are typically pitch shifted, time shifted, and / or resampled. In such examples, the adjusted reference fingerprint and / or reference signature reduces the computational intensity of analyzing the sample fingerprint to obtain the pitch shift, time shift, and / or resampling because the sample fingerprint can be compared to the predetermined adjustment value. For example, the reference fingerprint can be adjusted by a predefined value. The predefined pitch shift value and / or predefined time shift value can correspond to, for example, the pitch shift and time shift applied to the media by the media producer (e.g., iHeart Top 40 pop radio songs were pitch-shifted and / or time-shifted by 2.5%).
[0101] exist Figure 2In the example of , the database 216 can be implemented by a volatile memory (e.g., synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS dynamic random access memory (RDRAM), etc.) and / or a non-volatile memory (e.g., flash memory). The database 216 can be implemented additionally or alternatively by one or more double data rate (DDR) memories (such as DDR, DDR2, DDR3, DDR4, mobile DDR (mDDR), etc.). The database 216 can be implemented additionally or alternatively by one or more mass storage devices (such as hard disk drives, optical disk drives, digital universal disk drives, solid state disk drives, etc.). Although in the example shown, the database 216 is illustrated as a single database, the database 216 can be implemented by any number and / or any type of databases. In addition, the data stored in the database 216 can be in any data format, such as binary data, comma-delimited data, tab-delimited data, structured query language (SQL) structures, etc.
[0102] In the examples disclosed herein, each of the network interface 202, the fingerprint pitch tuner 204, the fingerprint speed tuner 206, the query comparator 208, the fingerprint generator 210, the media processor 212, the report generator 214, and the database 216 communicates with other elements of the central facility 112. For example, the network interface 202, the fingerprint pitch tuner 204, the fingerprint speed tuner 206, the query comparator 208, the fingerprint generator 210, the media processor 212, the report generator 214, and the database 216 communicate via an example communication bus 218. In some examples disclosed herein, the network interface 202, the fingerprint pitch tuner 204, the fingerprint speed tuner 206, the query comparator 208, the fingerprint generator 210, the media processor 212, the report generator 214, and the database 216 can communicate via any suitable wired and / or wireless communication system. Furthermore, in some examples disclosed herein, each of the network interface 202, fingerprint pitch tuner 204, fingerprint speed tuner 206, query comparator 208, fingerprint generator 210, media processor 212, report generator 214, and database 216 can communicate with any component external to the central facility 112 via any suitable wired and / or wireless communication system.
[0103] Figure 3 It is shown Figure 1 1 is a block diagram of further details of an example application 114. The example application 114 includes an example component interface 302, an example fingerprint generator 304, an example network interface 306, an example user interface 308, and an example report generator 310.
[0104] exist Figure 3 In the example shown, component interface 302 is configured to obtain information from and / or send information to a component on a device (e.g., client device 102, end-user device 108, etc.) that includes application 114. For example, component interface 302 can be configured to communicate with a camera and / or microphone to obtain video and / or audio signals.
[0105] In some examples, the example component interface 302 implements an example means for a component interface. The component interface means is implemented by, for example, at least Fig.10 The executable instructions implemented by block 1002 may be implemented by executable instructions such as the executable instructions implemented by block 1002. Fig.10 The executable instructions of block 1002 may be in a Fig.12 The component interface means are executed on at least one processor of the example processor 1212. In other examples, the component interface means are implemented by hardware logic, hardware-implemented state machines, logic circuits, and / or any other combination of hardware, software, and / or firmware.
[0106] exist Figure 3 In the example of , the fingerprint generator 304 is a device that can generate one or more sample fingerprints and / or one or more sample signatures from a sampled medium (e.g., an audio signal). For example, the fingerprint generator 304 can separate the audio signal (e.g., a digitized audio signal) into time-frequency bins and / or audio signal frequency components. For example, the fingerprint generator 304 can perform a fast Fourier transform (FFT) on the audio signal to transform the audio signal into the frequency domain.
[0107] Additionally, the example fingerprint generator 304 may divide the transformed audio signal into two or more frequency bins (e.g., using a Hamming function, a Hann function, etc.). In this example, each audio signal frequency component is associated with a frequency bin in the two or more frequency bins. Additionally or alternatively, the fingerprint generator 304 may aggregate the audio signal into one or more time periods (e.g., the duration of the audio, a six-second period, a 1-second period, etc.). In other examples, the fingerprint generator 304 may use any suitable technique to transform the audio signal (e.g., discrete Fourier transform, sliding time window Fourier transform, wavelet transform, discrete Hadamard transform, discrete Walsh Hadamard, constant Q transform, discrete cosine transform, etc.). In some examples, the fingerprint generator 304 may include one or more bandpass filters (BPF). In some examples, the processed audio signal may be represented by a spectrogram. The following is in conjunction with FIG. 4A to FIG. 4B The process of the fingerprint generator 304 that generates the spectrogram is discussed.
[0108] exist Figure 3In the example of , the fingerprint generator 304 can determine audio characteristics of a portion of the audio signal (e.g., audio signal frequency components, audio regions around time-frequency bins, etc.). For example, the fingerprint generator 304 can determine the mean energy (e.g., average power, etc.) of one or more of the audio signal frequency components. Additionally or alternatively, the fingerprint generator 304 can determine other characteristics of a portion of the audio signal (e.g., mode energy, median energy, mode power, median energy, mean energy, mean amplitude, etc.).
[0109] exist Figure 3 In the example shown, the fingerprint generator 304 can normalize one or more time-frequency bins according to the associated audio characteristics of the surrounding audio area. For example, the fingerprint generator 304 can normalize the time-frequency bins according to the mean energy of the surrounding audio area. In other examples, the fingerprint generator 304 normalizes some of the audio signal frequency components according to the associated audio characteristics. For example, the fingerprint generator 304 can use the mean energy associated with the audio signal frequency component to normalize each time-frequency bin of the audio signal frequency component. In some examples, the processed bins (e.g., normalized time-frequency bins, normalized audio signal frequency components, etc.) can be represented as a spectrogram. The following is combined with Figure 4C The process of the fingerprint generator 304 that generates a normalized spectrogram is discussed.
[0110] exist Figure 3 In the example shown, the fingerprint generator 304 can select one or more points from the normalized audio signal to be used to generate the fingerprint and / or signature. For example, the fingerprint generator 304 can select multiple energy maxima of the normalized audio signal. In other examples, the fingerprint generator 304 can select any other suitable points of the normalized audio. Figure 4D and Figure 5 The process of the fingerprint generator 304 that generates the fingerprint is discussed.
[0111] Additionally or alternatively, the fingerprint generator 304 can weight the selection of points based on the category of the audio signal. For example, if the category of the audio signal is music, the fingerprint generator 304 can focus the selection of points on common frequency ranges for music (e.g., bass, treble, etc.). In some examples, the fingerprint generator 304 can determine the category of the audio signal (e.g., music, speech, sound effects, advertising, etc.). The example fingerprint generator 304 uses the selected points to generate a fingerprint and / or signature. The example fingerprint generator 304 can generate a fingerprint based on the selected points using any suitable method.
[0112] In some examples, the example fingerprint generator 304 implements an example means for fingerprint generation. The fingerprint generation means includes at least Fig.10 The executable instructions implemented by block 1004 may be implemented by executable instructions such as the executable instructions implemented by block 1004. Fig.10 The executable instructions of block 1004 may be in a Fig.12 The fingerprint generation means is executed on at least one processor of the example processor 1212. In other examples, the fingerprint generation means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0113] exist Figure 3 In the example of Figure 1 The network interface 306 receives and / or sends reports, media, and / or other information from the central facility 112 and / or the example client device 102. The reports, media, and / or other information may be formatted as data packets, HTTP messages, text, pdf, etc. However, any other message format and / or protocol may additionally or alternatively be used, such as FTP, SMTP, HTTPS protocol, etc. Additionally or alternatively, media (e.g., audio signals) may be received as radio waveforms, MP3 files, MP4 files, and / or any other suitable audio format.
[0114] exist Figure 3 In the example shown, the network interface 306 is configured to identify whether the client has defined a guessed pitch offset, a guessed time offset, and / or a guessed resampling rate for the query. The network interface 306 may also be configured to send one or more fingerprints and / or one or more signatures and a query including a pitch offset, a time offset, and / or a resampling rate to the central facility 112. In addition, the network interface 306 is also configured to receive a response to the query from the central facility 112. The network interface 306 may also be configured to identify whether the response to the query includes information corresponding to one or more matches to the guessed pitch offset, the guessed time offset, and / or the guessed resampling rate.
[0115] In some examples, the example network interface 306 implements an example means for network interfacing. The network interfacing means is implemented by, for example, at least Fig.10 The method may be implemented by executable instructions such as the executable instructions implemented by blocks 1006, 1012, 1014, 1022 of FIG. Fig.10 The executable instructions of blocks 1006, 1012, 1014, 1022 may be in a file such as Fig.12The network interface means is executed on at least one processor of the example processor 1212. In other examples, the network interface means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0116] exist Figure 3 In the example of , the user interface 308 is configured to obtain information from and / or send information to a user of a device including the application 114. For example, the user interface 308 can implement a graphical user interface (GUI) or any other suitable interface. The user interface 308 can receive and / or send reports, pitch shifts, time offsets, resampling rates, and / or other information from a user of the end-user device 108. For example, the user interface 308 can display a prompt for an identification adjustment instruction to the user, which corresponds to how many and / or what types of pitch shifts, time offsets, and / or resampling rates are to be performed on one or more sample fingerprints in an attempt to identify the audio associated with the sample fingerprint. Additionally or alternatively, the user interface 308 can display the results of the query (e.g., corresponding to the artist, song, identification information, pitch shift information, time offset information, resampling information, etc. of the audio) to the end user. Reports, pitch shifts, time offsets, resampling rates, and / or other information can be formatted as HTTP messages. However, any other message format and / or protocol, such as FTP, SMTP, HTTPS protocols, etc., can be used additionally or alternatively.
[0117] In some examples, the example user interface 308 implements an example means for a user interface. The user interface means includes at least Fig.10 The executable instructions implemented by block 1008 may be implemented by executable instructions such as the executable instructions implemented by block 1008. Fig.10 The executable instructions of block 1008 may be in a file such as Fig.12 The example processor 1212 is executed on at least one processor of the example processor 1212. In other examples, the user interface means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0118] exist Figure 3In the example shown, the report generator 310 is a device configured to generate and / or prepare a report. The report generator 310 prepares reports and / or prompts indicating one or more of commonly used pitch shifts, time offsets and / or resampling rates. In addition, the report generator 310 can suggest one or more pitch shifts, time offsets and / or resampling rates indicated in the query to the client and / or end user based on one or more reports generated by an external device (e.g., central facility 112). In addition, the report generator 310 can identify one or more audio signals corresponding to one or more reference fingerprints and / or reference signatures that match one or more sample fingerprints. In some examples, the report generator 310 generates a report including one or more graphical analysis tools for identifying the frequency of occurrence of each pitch shift, time offset and / or resampling rate.
[0119] exist Figure 3 In an example of , the report generator 310 can prepare a report that associates one or more reference fingerprints with one or more sample fingerprints. For example, the report generator 310 can prepare a report that identifies: (a) the song corresponding to the reference fingerprint, and (b) the pitch shift, time shift, and / or resampling rate used by the central facility 112 to match the sample fingerprint and / or sample signature to the reference fingerprint. For example, the report generator 310 can generate a report that indicates that the audio sample is a sample of "Take it Easy" by the Eagles, and indicates that the sample pitch is shifted down by 5% and the time shift increases the playback speed by 5% (e.g., a report indicating information associated with a reference fingerprint and adjustments (e.g., transformations, pitch shifts, time shifts, etc.)
[0120] In some examples, the report can be data and / or data packets that are stored locally and / or sent to an external device (e.g., the example client device 102). In some examples, the application 114 and / or another device can use the stored report to adjust the adjustment instructions for subsequent queries. For example, if the adjustment instruction identifies a specific pitch offset that does not correspond to a match within X queries (e.g., if a 40 Hz pitch offset is included in the adjustment instruction, and after 100 queries, there has never been a match with the 40 Hz pitch offset), the application 114 and / or another device can adjust the adjustment instruction to remove the specific pitch offset for subsequent queries (e.g., remove the 40 Hz pitch offset from the adjustment instruction). In another example, if the client selects a specific time offset that corresponds to a large number of matches, the application 114 and / or another device can adjust the adjustment instruction so that the central facility 112 performs the specific time offset first during subsequent queries (e.g., subsequent fingerprinting) of multiple different offsets.
[0121] In some examples, the example report generator 310 implements an example means for report generation. The report generation means includes at least Fig.10 The method may be implemented by executable instructions such as the executable instructions implemented by boxes 1016, 1018, and 1020 of FIG. Fig.10 The executable instructions of blocks 1016, 1018, 1020 may be in a file such as Fig.11 The report generation means is executed on at least one processor of the example processor 1112. In other examples, the report generation means is implemented by hardware logic, a hardware-implemented state machine, a logic circuit, and / or any other combination of hardware, software, and / or firmware.
[0122] In the examples disclosed herein, each of the component interface 302, the fingerprint generator 304, the network interface 306, the user interface 308, and the report generator 310 communicates with other elements of the application 114. For example, the component interface 302, the fingerprint generator 304, the network interface 306, the user interface 308, and the report generator 310 communicate via an example communication bus 312. In some examples disclosed herein, the component interface 302, the fingerprint generator 304, the network interface 306, the user interface 308, and the report generator 310 can communicate via any suitable wired and / or wireless communication system. In addition, in some examples disclosed herein, each of the component interface 302, the fingerprint generator 304, the network interface 306, the user interface 308, and the report generator 310 can communicate with any component external to the application 114 via any suitable wired and / or wireless communication system.
[0123] Figure 4A and Figure 4B Example of Figure 2 The fingerprint generator 210 and Figure 3 Example spectrum graph 400 generated by fingerprint generator 304. Figure 4A In the illustrated example of , the example unprocessed spectrogram 400 includes an example first time-frequency bin 406 surrounded by an example first audio region 408. Figure 4B In the illustrated example of , the example unprocessed spectrogram includes an example second time-frequency bin 410 surrounded by an example audio region 412 . Figure 4A and Figure 4B An example of an unprocessed spectrogram 400, Figure 4C The normalized spectrogram 402 of Figure 4D The sub-fingerprints 404 each include an example vertical axis 414 representing frequency bins and an example horizontal axis 416 representing time bins. Figure 4A and Figure 4BExample audio regions 408 and 412 are illustrated from which the fingerprint generator 210 and / or the fingerprint generator 304 obtain normalized audio characteristics, and the fingerprint generator 210 and / or the fingerprint generator 304 use the normalized audio characteristics to normalize the first time-frequency bin 406 and the second time-frequency bin 410, respectively. In the illustrated example, each time-frequency bin of the unprocessed spectrogram 400 is normalized to generate the normalized spectrogram 402. In other examples, any suitable number of time-frequency bins of the unprocessed spectrogram 400 may be normalized to generate the normalized spectrogram 402. Figure 4C A normalized spectrum diagram 402 of .
[0124] The example vertical axis 414 has frequency bin units generated by a fast Fourier transform (FFT) and has a length of 1024 FFT bins. In other examples, the example vertical axis 414 can be measured by any other suitable technique for measuring frequency (e.g., Hertz, another transform algorithm, etc.). In some examples, the vertical axis 414 covers the entire frequency range of the audio signal. In other examples, the vertical axis 414 can cover a portion of the audio signal.
[0125] In the example shown, the example horizontal axis 416 represents a time period of 11.5 seconds for the total length of the unprocessed spectrogram 400. In the example shown, the horizontal axis 416 has sixty-four millisecond (ms) intervals as units. In other examples, the horizontal axis 416 can be measured in any other suitable unit (e.g., 1 second, etc.). For example, the horizontal axis 416 covers the complete duration of the audio. In other examples, the horizontal axis 416 can cover a portion of the duration of the audio signal. In the example shown, the size of each time-frequency bin of the spectrogram 400, 402 is 64ms×1FFT bin.
[0126] exist Figure 4AIn the illustrated example of , the first time-frequency bin 406 is associated with the intersection of the frequency bin and the time bin of the unprocessed spectrogram 400 and a portion of the audio signal associated with the intersection. The example first audio region 408 includes time-frequency bins within a predefined distance from the example first time-frequency bin 406. For example, the fingerprint generator 210 and / or the fingerprint generator 304 can determine the vertical length of the first audio region 408 (e.g., the length of the first audio region 408 along the longitudinal axis 414, etc.) based on a set number of FFT bins (e.g., 5 bins, 11 bins, etc.). Similarly, the fingerprint generator 210 and / or the fingerprint generator 304 can determine the horizontal length of the first audio region 408 (e.g., the length of the first audio region 408 along the horizontal axis 416, etc.). In the illustrated example, the first audio region 408 is square. Alternatively, the first audio region 408 can be any suitable size and shape and can contain any suitable combination of time-frequency bins within the unprocessed spectrogram 400 (e.g., any suitable group of time-frequency bins, etc.). The example fingerprint generator 210 and / or fingerprint generator 304 may then determine audio characteristics (e.g., mean energy, etc.) of the time-frequency bins contained within the first audio region 408. Using the determined audio characteristics, the fingerprint generator 210 and / or fingerprint generator 304 may normalize the associated values of the first time-frequency bin 406 (e.g., the energy of the first time-frequency bin 406 may be normalized by the mean energy of the respective time-frequency bins within the first audio region 408).
[0127] exist Figure 4BIn the illustrated example of , the second time-frequency bin 410 is associated with the intersection of the frequency bin and the time bin of the unprocessed spectrogram 400 and a portion of the audio signal associated with the intersection. The example second audio region 412 includes time-frequency bins within a predefined distance from the example second time-frequency bin 410. Similarly, the fingerprint generator 210 and / or the fingerprint generator 304 can determine the horizontal length of the second audio region 412 (e.g., the length of the second audio region 412 along the horizontal axis 416, etc.). In the illustrated example, the second audio region 412 is square. Alternatively, the second audio region 412 can be any suitable size and shape, and can include any suitable combination of time-frequency bins within the unprocessed spectrogram 400 (e.g., any suitable group of time-frequency bins, etc.). In some examples, the second audio region 412 can overlap with the first audio region 408 (e.g., include some of the same time-frequency bins, be shifted on the horizontal axis 416, be shifted on the vertical axis 414, etc.). In some examples, the second audio region 412 can have the same size and shape as the first audio region 408. In other examples, the second audio region 412 can have a different size and shape than the first audio region 408. The fingerprint generator 210 and / or the fingerprint generator 304 can then determine audio characteristics (e.g., mean energy, etc.) of the time-frequency bins contained in the second audio region 412. Using the determined audio characteristics, the fingerprint generator 210 and / or the fingerprint generator 304 can normalize the associated values of the second time-frequency bins 410 (e.g., the energy of the second time-frequency bins 410 can be normalized by the mean energy of the bins located within the second audio region 412).
[0128] Figure 4C Example of Figure 2 The fingerprint generator 210 and / or Figure 3 An example of a normalized spectrogram 402 generated by the fingerprint generator 304 of FIG. The fingerprint generator 210 and / or the fingerprint generator 304 may be configured to generate a normalized spectrogram 402 of FIG. FIG. 4A to FIG. 4B The normalized spectrogram 402 is generated by normalizing a plurality of time-frequency bins of the unprocessed spectrogram 400. For example, some or all of the time-frequency bins of the unprocessed spectrogram 400 may be normalized in a manner similar to the manner in which the time-frequency bins 404A and 404B are normalized. The region 402 has now been normalized by the local mean energy within the local region surrounding the region. Figure 4C The resulting frequency bins are normalized. As a result, the darker regions are the regions with the maximum energy in their respective local regions. This enables the fingerprint to contain relevant audio features in regions with even lower energy than the typically louder bass frequency regions.
[0129] Figure 4DThe normalized spectrum diagram 402 is illustrated by Figure 2 The fingerprint generator 210 and / or Figure 3 4. The sub-fingerprint 404 of the example fingerprint generated by the fingerprint generator 304 of the embodiment of the present invention. The sub-fingerprint 404 is a representation of a portion of the audio signal. The fingerprint generator 210 and / or the fingerprint generator 304 can generate the sub-fingerprint 404 by filtering out low energy bin values from the normalized spectrogram 402. For example, the fingerprint generator 210 and / or the fingerprint generator 304 can generate the sub-fingerprint 404 for a 64 ms frame of the normalized spectrogram 402. In other examples, the fingerprint generator 210 and / or the fingerprint generator 304 can generate the sub-fingerprint 404 by any suitable fingerprint generation technique. In the examples disclosed herein, the fingerprint pitch tuner 204 can adjust the fingerprint (e.g., one or more sub-fingerprints (e.g., sub-fingerprint 404) of the fingerprint) to accommodate the pitch shift by multiplying the bin values (e.g., the associated energy values of the bins) of the fingerprint (e.g., one or more sub-fingerprints (e.g., sub-fingerprint 404) of the fingerprint) based on the conjectured pitch shift. Figure 4D The example sub-fingerprint 404 of includes 20 bin values representing extreme values during a portion of the audio signal. In additional or alternative examples, the number of bin values in the sub-fingerprint 404 can be variable. For example, the number of bin values in the sub-fingerprint 404 can be based on a threshold number of bin values. In some examples, the number of bin values in the sub-fingerprint 404 can be based on a trained neural network to determine which extreme values are most likely to contribute to a match.
[0130] exist Figure 4D In the example of , sub-fingerprint 404 is illustrated in a graphical setting. In some examples, the fingerprint can be represented by an array including data values (e.g., 0, 1, logical high value, logical low value, etc.). For example, the array can be 1024 rows×94 columns. For example, each column can correspond to a sub-fingerprint and / or a frame (e.g., sub-fingerprint 404), and each row can correspond to a bin value. In such an example, each bin value can include an associated data value (e.g., 0, 1, logical high value, logical low value, etc.), and each sub-fingerprint can include 20 data values 1 and / or logical high values.
[0131] Figure 5 An example audio sample 502 and an example fingerprint 504 are illustrated. Figure 5 In the example of , fingerprint 504 corresponds to audio sample 502. Figure 5 In the example of , audio sample 502 is an audio recording. For example, audio sample 502 may be a short segment of a song, speech, concert, etc. Audio sample 502 may be represented by a plurality of fingerprints including fingerprint 504 (e.g., indicated by a dashed rectangle on the audio sample).
[0132] exist Figure 5 In the example shown, the example fingerprint 504 is a representation of features of the audio sample 502. For example, the fingerprint 504 includes data representing characteristics of the audio sample 502 for a time frame of the fingerprint 504. For example, the fingerprint 504 can be a compressed digital summary of the audio sample 502, including a number of highest output values, a number of lowest output values, a frequency value, other characteristics, etc. of the audio sample 502.
[0133] exist Figure 5 In the example shown, the example sub-fingerprints 506 are segmented portions of the fingerprint 504. For example, the fingerprint 504 can be divided into a specified number of sub-fingerprints 506 (e.g., ten, fifty, one hundred, etc.), which can then be processed separately. For example, the sub-fingerprints 506 can correspond to frames of the fingerprint 504. In the examples disclosed herein, the fingerprint speed tuner 206 can adjust the fingerprint (e.g., fingerprint 504) to accommodate the time shift by copying and / or deleting frames (e.g., sub-fingerprints 506) from the fingerprint (e.g., sub-fingerprints 506).
[0134] Although Figure 2 Demonstrates the implementation Figure 1 An example method of central facility 112 and Figure 3 Demonstrates the implementation Figure 1 The example method of application 114, however Figure 2 and / or Figure 3 One or more of the illustrated elements, processes and / or devices may be combined, divided, rearranged, omitted, eliminated and / or implemented in any other manner. Figure 2 example network interface 202, example fingerprint pitch tuner 204, example fingerprint speed tuner 206, example query comparator 208, example fingerprint generator 210, example media processor 212, example report generator 214, example database 216, and / or (more generally) example central facility 112, and / or Figure 3 The example component interface 302, the example fingerprint generator 304, the example network interface 306, the example user interface 308, the example report generator 310, and / or the example application 114 may be implemented in hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, Figure 2 example network interface 202, example fingerprint pitch tuner 204, example fingerprint speed tuner 206, example query comparator 208, example fingerprint generator 210, example media processor 212, example report generator 214, example database 216, and / or (more generally) example central facility 112, and / or Figure 3Any of the example component interface 302, the example fingerprint generator 304, the example network interface 306, the example user interface 308, the example report generator 310, and / or (more generally) the example application 114 can be implemented by one or more analog or digital circuits, logic circuits, programmable processors, programmable controllers, graphics processing units (GPUs), digital signal processors (DSPs), application specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field programmable logic devices (FPLDs). When any of the apparatus or system claims of this patent is understood to cover pure software and / or firmware implementations, Figure 2 example network interface 202, example fingerprint pitch tuner 204, example fingerprint speed tuner 206, example query comparator 208, example fingerprint generator 210, example media processor 212, example report generator 214, example database 216, and / or (more generally) example central facility 112, and / or Figure 3 At least one of the example component interface 302, the example fingerprint generator 304, the example network interface 306, the example user interface 308, the example report generator 310, and / or (more generally) the example application 114 is thereby explicitly defined as comprising a non-transitory computer-readable storage device or storage disk (such as a memory, a digital versatile disk (DVD), a compact disk (CD), a Blu-ray disk, etc.) having software and / or firmware. Further, Figure 2 An example of a central facility 112 and / or Figure 3 Example applications 114 may include, in addition to Figure 2 and / or Figure 3 Elements, processes and / or devices other than or in place of the illustrated elements, processes and / or devices Figure 2 and / or Figure 3 One or more of the illustrated elements, processes and / or devices, and / or may include more than one of any or all of the illustrated elements, processes and devices. As used herein, the phrase "communicating" (including variations thereof) encompasses direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but additionally includes selective communication at regular intervals, scheduled intervals, non-periodic intervals and / or one-time events.
[0135] Fig. 6A , Figure 6B , Figure 7 , Figure 8 and / or Fig. 9 A flowchart representing example hardware logic, machine-readable instructions, hardware-implemented state machines, and / or any combination thereof for implementing the central facility 112 is shown. Fig.10114. A flowchart representing example hardware logic, machine readable instructions, hardware implemented state machines, and / or any combination thereof for implementing application 114 is shown. The machine readable instructions may be instructions for a computer processor (such as the following in conjunction with Fig.11 and / or Fig.12 The example processor platforms 1100, 1200 discussed herein may be one or more executable programs or one or more portions of executable programs executed by the processors 1112, 1212 shown in the example processor platforms 1100, 1200. The programs may be embodied in software stored on a non-transitory computer-readable storage medium such as a CD-ROM, floppy disk, hard drive, DVD, Blu-ray disk, or memory associated with the processors 1112, 1212, but all of the programs and / or portions thereof may alternatively be executed by a device other than the processors 1112, 1212, and / or embodied in firmware or dedicated hardware. Furthermore, although reference is made to Fig. 6A , Figure 6B , Figure 7 , Figure 8 , Fig. 9 and / or Fig.10 The illustrated flowcharts describe example programs, but many other methods of implementing the example central facility 112 and / or the example application 114 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the blocks may be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) that are configured to perform corresponding operations without executing software or firmware.
[0136] The machine-readable instructions described herein may be stored in one or more of a compressed format, an encrypted format, a segmented format, a compiled format, an executable format, a packaged format, etc. The machine-readable instructions as described herein may be stored as data (e.g., portions of instructions, codes, representations of codes, etc.) that can be used to create, manufacture, and / or generate machine-executable instructions. For example, the machine-readable instructions may be segmented and stored on one or more storage devices and / or computing devices (e.g., servers). The machine-readable instructions may require one or more of installation, modification, adaptation, update, combination, supplementation, configuration, decryption, decompression, unpacking, distribution, redistribution, compilation, etc., to make them directly readable, interpretable, and / or executable by a computing device and / or other machine. For example, the machine-readable instructions may be stored in multiple parts that are individually compressed, encrypted, and stored on separate computing devices, where the parts, when decrypted, decompressed, and combined, form a set of executable instructions that implement the program described herein.
[0137] In another example, the machine-readable instructions may be stored in a state where they can be read by a computer, but require the addition of a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), an application programming interface (API), etc., in order to execute the instructions on a particular computing device or other device. In another example, the machine-readable instructions and / or corresponding programs may need to be configured before they can be executed in whole or in part (e.g., stored settings, data inputs, recorded network addresses, etc.). Thus, the disclosed machine-readable instructions and / or corresponding programs are intended to encompass such machine-readable instructions and / or programs regardless of the particular format or state in which the machine-readable instructions and / or programs are stored or otherwise stationary or transported.
[0138] The machine-readable instructions described herein may be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, the machine-readable instructions may be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, Hypertext Markup Language (HTML), Structured Query Language (SQL), Swift, etc.
[0139] As described above, the present invention may be implemented using executable instructions (e.g., computer and / or machine readable instructions) stored on a non-transitory computer and / or machine readable medium (such as a hard drive, flash memory, read-only memory, compact disk, digital versatile disk, cache, random access memory, and / or any other storage device or storage disk in which information is stored for any duration (e.g., for an extended period of time, permanently, for a simple instance, for temporary buffering, and / or for caching information)). Fig. 6A , Figure 6B , Figure 7 , Figure 8 , Fig. 9 and / or Fig.10 As used herein, the term non-transitory computer readable medium is expressly defined to include any type of computer readable storage devices and / or storage disks and to exclude propagating signals and to exclude transmission media.
[0140] "Include" and "comprising" (and all forms and tenses thereof) are used herein as open-ended terms. Thus, whenever a claim employs any form of "includes" or "comprising" (e.g., comprises, includes, comprising, including, having, etc.) as a preamble or within any type of claim recitation, it will be understood that additional elements, terms, etc. may be present without falling outside the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used as a transitional term in, for example, a preamble of a claim, it is open-ended in the same manner as the terms "includes" and "comprising" are open-ended. When used, for example, in a form such as A, B, and / or C, the term "and / or" refers to any combination or subset of A, B, C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, and (7) A and B and C. As used herein, in the context of describing structures, components, items, objects, and / or things, the phrase “at least one of A and B” is intended to mean implementations that include any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein, in the context of describing structures, components, items, objects, and / or things, the phrase “at least one of A or B” is intended to mean implementations that include any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. As used herein, in the context of describing the execution of processes, instructions, actions, activities, and / or steps, the phrase “at least one of A and B” is intended to mean implementations that include any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B. Similarly, as used herein, in the context of describing the execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" is intended to mean an implementation that includes any of the following: (1) at least one A, (2) at least one B, and (3) at least one A and at least one B.
[0141] As used herein, singular references (e.g., "one," "first," "second," etc.) do not exclude the plural. As used herein, the term "one" entity refers to one or more of that entity. The terms "one," "one or more," and "at least one" are used interchangeably herein. In addition, although listed separately, multiple devices, elements, or method actions may be implemented by, for example, a single unit or processor. In addition, although individual features may be included in different examples or claims, these features may be combined, and inclusion in different examples or claims does not mean that a combination of features is not feasible and / or advantageous.
[0142] Fig. 6A and Figure 6B is a flow chart representing a process 600 for running a query associated with pitch-shifted, time-shifted, and / or resampled media, which may be performed using a method that may be executed to implement Figure 1 and Figure 2 The process 600 begins at block 602, where the network interface 202 obtains a sample fingerprint. For example, the network interface 202 may receive and / or otherwise obtain the sample fingerprint from the client device 102 and / or the end-user device 108 via the network 104. At block 604, the network interface 202 obtains a query from an external device. For example, the network interface 202 may obtain a query including a guessed pitch shift, a guessed time shift, and / or a guessed resampling rate from the client device 102 and / or the end-user device 108.
[0143] exist Fig. 6A In the example shown, at block 606, network interface 202 determines whether the query includes an indication to be run serially. If network interface 202 determines that the query includes an indication to be run serially (block 606: yes), process 600 proceeds to block 608. If network interface 202 determines that the query does not include an indication to be run serially (block 606: no), process 600 proceeds to block 609. Figure 6B At block 608, network interface 202 determines whether the query includes a guessed pitch shift.
[0144] exist Fig. 6A In the example of , if the network interface 202 determines that the query includes a guessed pitch shift (block 608: yes), the process 600 proceeds to block 610. If the network interface 202 determines that the query does not include a guessed pitch shift (block 608: no), the process 600 proceeds to block 620. At block 610, the network interface 202 selects a guessed pitch shift value for analysis by the central facility 112. At block 612, the fingerprint pitch tuner 204 pitch shifts one or more sample fingerprints based on the selected pitch shift. At block 614, the query comparator 208 determines whether the one or more pitch-shifted sample fingerprints match one or more reference fingerprints.
[0145] exist Fig. 6AIn the example shown, if the query comparator 208 determines that one or more pitch-shifted sample fingerprints match one or more reference fingerprints (block 614: yes), the process 600 proceeds to block 618. If the query comparator 208 determines that one or more pitch-shifted sample fingerprints do not match one or more reference fingerprints (block 614: no), the process 600 proceeds to block 616. At block 616, the network interface 202 determines whether there are additional guessed pitch shifts in the query to be analyzed by the central facility 112.
[0146] exist Fig. 6A In the example of , if the network interface 202 determines that there are additional guessed pitch shifts in the query (block: yes), the process 600 proceeds to block 610. If the network interface 202 determines that there are no additional guessed pitch shifts in the query (block: 616: no), the process 600 proceeds to block 620. At block 618, the query comparator 208 indicates such a guessed pitch shift: when the guessed pitch shift is applied to the sample fingerprint, it causes the sample fingerprint to match one or more reference fingerprints. For example, the query comparator 208 can store the matched pitch shift value and / or other information about the matched reference fingerprint in the database 216.
[0147] exist Fig. 6A In the example shown, at block 620, the network interface 202 determines whether the query includes a guessed time offset. If the network interface 202 determines that the query includes a guessed time offset (block 620: yes), the process 600 proceeds to block 622. If the network interface 202 determines that the query does not include a guessed time offset (block 620: no), the process 600 proceeds to block 632. At block 622, the network interface 202 selects a guessed time offset value for analysis by the central facility 112. At block 624, the fingerprint velocity tuner 206 time-shifts one or more sample fingerprints based on the selected time offset. At block 626, the query comparator 208 determines whether the one or more time-shifted sample fingerprints match one or more reference fingerprints.
[0148] exist Fig. 6A In the example shown, if the query comparator 208 determines that the one or more time-shifted sample fingerprints match the one or more reference fingerprints (block 626: yes), the process 600 proceeds to block 630. If the query comparator 208 determines that the one or more time-shifted sample fingerprints do not match the one or more reference fingerprints (block 626: no), the process 600 proceeds to block 628. At block 628, the network interface 202 determines whether there are additional guessed time offsets in the query to be analyzed by the central facility 112.
[0149] exist Fig. 6AIn the example of FIG. 6 , if the network interface 202 determines that there are additional guessed time offsets in the query (block: 628: yes), the process 600 proceeds to block 622. If the network interface 202 determines that there are no additional guessed time offsets in the query (block: 628: no), the process 600 proceeds to block 632. At block 630, the query comparator 208 indicates such a guessed time offset: when the guessed time offset is applied to the sample fingerprint, it causes the sample fingerprint to match one or more reference fingerprints. For example, the query comparator 208 can store the matched time offset value and / or other information about the matched reference fingerprint in the database 216.
[0150] exist Fig. 6A In the example shown, at block 632, the network interface 202 determines whether the query includes a guessed resampling rate. If the network interface 202 determines that the query includes a guessed resampling rate (block 632: yes), the process 600 proceeds to block 634. If the network interface 202 determines that the query does not include a guessed resampling rate (block 632: no), the process 600 proceeds to block 646. At block 634, the network interface 202 selects a guessed resampling rate for analysis by the central facility 112. At block 636, the fingerprint pitch tuner 204 pitch-shifts one or more sample fingerprints based on the selected resampling rate. At block 638, the fingerprint velocity tuner 206 time-shifts the one or more sample fingerprints that were pitch-shifted by the fingerprint pitch tuner 204 at block 636 based on the selected resampling rate. At block 640, the query comparator 208 determines whether the one or more resampled sample fingerprints match one or more reference fingerprints.
[0151] exist Fig. 6A In the example shown, if the query comparator 208 determines that one or more resampled sample fingerprints match one or more reference fingerprints (block 640: yes), the process 600 proceeds to block 644. If the query comparator 208 determines that one or more resampled sample fingerprints do not match one or more reference fingerprints (block 640: no), the process 600 proceeds to block 642. At block 642, the network interface 202 determines whether there are additional guessed resampling rates in the query to be analyzed by the central facility 112.
[0152] exist Fig. 6AIn the example of , if the network interface 202 determines that there are additional guessed resampling rates in the query (block: 642: yes), the process 600 proceeds to block 634. If the network interface 202 determines that there are no additional guessed resampling rates in the query (block: 642: no), the process 600 proceeds to block 646. At block 644, the query comparator 208 indicates such a guessed resampling rate: when the guessed resampling rate is applied to the sample fingerprint, it causes the sample fingerprint to match one or more reference fingerprints. For example, the query comparator 208 can store the matching pitch offset value, time offset value, and / or other information about the matching reference fingerprint in the database 216.
[0153] exist Fig. 6A In the example shown, at block 646, the network interface 202 determines whether the central facility 112 has received additional queries. If the network interface 202 determines that the central facility 112 has received additional queries (block 646: yes), the process 600 proceeds to block 606. If the network interface 202 determines that the central facility 112 has not received additional queries (block 646: no), the process 600 proceeds to block 648. At block 648, the report generator 214 generates a report based on one or more of the matched pitch offsets, the matched time offsets, and / or the matched resampling rates. At block 650, the network interface 202 sends the report to one or more external devices (e.g., client device 102, end-user device 108, etc.). After block 650, the process 600 terminates.
[0154] Figure 6B Describes the parallel processing of queries. Figure 6B The process 600 of begins at blocks 652, 662, and 674. At block 652, the network interface 202 determines whether the query includes a guessed pitch shift. If the network interface 202 determines that the query includes a guessed pitch shift (block 652: yes), the process 600 proceeds to block 654. If the network interface 202 determines that the query does not include a guessed pitch shift (block 652: no), the process 600 proceeds to block 684. At block 654, the network interface 202 selects each guessed pitch shift value included in the query. At block 656, the fingerprint pitch tuner 204 pitch shifts one or more sample fingerprints based on the selected pitch shift. At block 658, the query comparator 208 determines whether one or more pitch-shifted sample fingerprints match one or more reference fingerprints.
[0155] exist Figure 6BIn the example shown, if the query comparator 208 determines that one or more pitch-shifted sample fingerprints match one or more reference fingerprints (block 658: yes), the process 600 proceeds to block 660. If the query comparator 208 determines that one or more pitch-shifted sample fingerprints do not match one or more reference fingerprints (block 658: no), the process 600 proceeds to block 684. At block 660, the query comparator 208 indicates a conjectured pitch shift that, when applied to the sample fingerprint, causes the sample fingerprint to match one or more reference fingerprints. For example, the query comparator 208 may store the matched pitch shift value and / or other information about the matched reference fingerprint in the database 216.
[0156] exist Figure 6B In the example shown, at block 662, the network interface 202 determines whether the query includes a guessed resampling rate. If the network interface 202 determines that the query includes a guessed resampling rate (block 662: yes), the process 600 proceeds to block 664. If the network interface 202 determines that the query does not include a guessed resampling rate (block 662: no), the process 600 proceeds to block 684. At block 664, the network interface 202 selects each guessed resampling rate in the query to analyze. At block 666, the fingerprint pitch tuner 204 pitch-shifts one or more sample fingerprints based on the selected resampling rate. At block 668, the fingerprint velocity tuner 206 time-shifts the one or more sample fingerprints that were pitch-shifted by the fingerprint pitch tuner 204 at block 666 based on the selected resampling rate. At block 670, the query comparator 208 determines whether the one or more resampled sample fingerprints match one or more reference fingerprints.
[0157] exist Figure 6B In the example shown, if the query comparator 208 determines that one or more resampled sample fingerprints match one or more reference fingerprints (block 670: yes), the process 600 proceeds to block 672. If the query comparator 208 determines that one or more resampled sample fingerprints do not match one or more reference fingerprints (block 670: no), the process 600 proceeds to block 684. At block 672, the query comparator 208 indicates a guessed resampling rate that, when applied to the sample fingerprint, causes the sample fingerprint to match one or more reference fingerprints. For example, the query comparator 208 may store the matched pitch offset value, time offset value, and / or other information about the matched reference fingerprint in the database 216.
[0158] exist Figure 6BAs shown, at block 674, the network interface 202 determines whether the query includes a guessed time offset. If the network interface 202 determines that the query includes a guessed time offset (block 674: yes), the process 600 proceeds to block 676. If the network interface 202 determines that the query does not include a guessed time offset (block 674: no), the process 600 proceeds to block 684. At block 676, the network interface 202 selects each guessed time offset value included in the query for analysis. At block 678, the fingerprint velocity tuner 206 time-shifts one or more sample fingerprints based on the selected time offset. At block 680, the query comparator 208 determines whether the one or more time-shifted sample fingerprints match one or more reference fingerprints.
[0159] exist Figure 6B In the example shown, if the query comparator 208 determines that one or more time-shifted sample fingerprints match one or more reference fingerprints (block 680: yes), the process 600 proceeds to block 682. If the query comparator 208 determines that one or more time-shifted sample fingerprints do not match one or more reference fingerprints (block 680: no), the process 600 proceeds to block 684. At block 682, the query comparator 208 indicates a guessed time offset that, when applied to the sample fingerprint, causes the sample fingerprint to match the one or more reference fingerprints. For example, the query comparator 208 may store the matched time offset value and / or other information about the matched reference fingerprint in the database 216.
[0160] exist Figure 6B In the example shown, at block 684, the network interface 202 determines whether the central facility 112 has received additional queries. If the network interface 202 determines that the central facility 112 has received additional queries (block 684: yes), the process 600 proceeds to block 606. If the network interface 202 determines that the central facility 112 has not received additional queries (block 684: no), the process 600 proceeds to block 686. At block 686, the report generator 214 generates a report based on one or more of the matched pitch offsets, the matched time offsets, and / or the matched resampling rates. At block 688, the network interface 202 sends the report to one or more external devices (e.g., client device 102, end-user device 108, etc.). After block 688, the process 600 terminates.
[0161] Figure 7 is a flow chart showing a process 700 for identifying trends in broadcast and / or presentation media that may be performed using Figure 1 and Figure 2The process 700 begins at box 702, where the network interface 202 collects broadcast media. For example, the network interface 202 may collect media broadcast by the media producer 110. In box 704, the fingerprint generator 210 generates one or more normalized spectrograms based on the collected broadcast media. In box 706, the fingerprint generator 210 generates one or more sample fingerprints from the normalized spectrogram. In box 708, the fingerprint pitch tuner 204 pitch-shifts one or more sample fingerprints based on a predefined set of pitch shift values. For example, the predefined pitch shift value may correspond to a range of pitch shifts (e.g., -10% to 10%). In box 710, the query comparator 208 determines whether one or more pitch-shifted sample fingerprints match one or more reference fingerprints.
[0162] exist Figure 7 In the example shown, if the query comparator 208 determines that one or more pitch-shifted sample fingerprints match one or more reference fingerprints (block 710: yes), the process 700 proceeds to block 712. If the query comparator 208 determines that one or more pitch-shifted sample fingerprints do not match one or more reference fingerprints (block 710: no), the process 700 proceeds to block 714. At block 712, the query comparator 208 indicates a conjectured pitch shift that, when applied to the sample fingerprint, causes the sample fingerprint to match one or more reference fingerprints. For example, the query comparator 208 may store the matched pitch shift value and / or other information about the matched reference fingerprint in the database 216.
[0163] exist Figure 7 In the example shown, at block 714, the fingerprint speed tuner 206 time-shifts one or more sample fingerprints based on a predefined set of time offset values. For example, the predefined time offset values may correspond to a range of time offsets (e.g., -10% to 10%). At block 716, the query comparator 208 determines whether the one or more time-shifted sample fingerprints match one or more reference fingerprints.
[0164] exist Figure 7In the example shown, if the query comparator 208 determines that one or more time-shifted sample fingerprints match one or more reference fingerprints (block 716: yes), the process 700 proceeds to block 718. If the query comparator 208 determines that one or more time-shifted sample fingerprints do not match one or more reference fingerprints (block 716: no), the process 700 proceeds to block 720. At block 718, the query comparator 208 indicates a guessed time offset that, when applied to the sample fingerprint, causes the sample fingerprint to match the one or more reference fingerprints. For example, the query comparator 208 may store the matched time offset value and / or other information about the matched reference fingerprint in the database 216.
[0165] exist Figure 7 In the example of , at block 720, the fingerprint pitch tuner 204 pitch shifts one or more sample fingerprints based on a predefined set of pitch shift values. For example, the predefined pitch shift values may correspond to a range of pitch shifts (e.g., -10% to 10%). At block 722, the fingerprint velocity tuner 206 time shifts the one or more sample fingerprints that were pitch shifted by the fingerprint pitch tuner 204 at block 720 based on a corresponding predefined set of time shift values. For example, the predefined time shift values may correspond to a range of time shifts (e.g., -10% to 10%). At block 724, the query comparator 208 determines whether the one or more resampled sample fingerprints match one or more reference fingerprints.
[0166] exist Figure 7 In the example shown, if the query comparator 208 determines that one or more resampled sample fingerprints match one or more reference fingerprints (block 724: yes), the process 700 proceeds to block 726. If the query comparator 208 determines that one or more resampled sample fingerprints do not match one or more reference fingerprints (block 724: no), the process 700 proceeds to block 728. At block 726, the query comparator 208 indicates a guessed resampling rate that, when applied to the sample fingerprint, causes the sample fingerprint to match the one or more reference fingerprints. For example, the query comparator 208 may store the matched pitch offset value, time offset value, and / or other information about the matched reference fingerprint in the database 216.
[0167] exist Figure 7In the example of FIG. 7 , at block 728, the media processor 212 determines whether the central facility 112 has monitored the broadcast media for a threshold period of time. If the media processor 212 determines that the central facility 112 has monitored the broadcast media for a threshold period of time (block 728: yes), the process 700 proceeds to block 730. If the media processor 212 determines that the central facility 112 has not monitored the broadcast media for a threshold period of time (block 728: no), the process 700 proceeds to block 702. At block 730, the media processor 212 processes these matching and / or corresponding pitch offsets, time offsets, and / or resampling rates to identify the frequency of occurrence of each pitch offset, time offset, and / or resampling rate.
[0168] exist Figure 7 In the example shown, at block 732, network interface 202 determines whether central facility 112 has received a request for a recommendation including one or more pitch shifts, time shifts, and / or resampling rates that may be included in the query. If network interface 202 determines that central facility 112 has received the request (block 732: yes), process 700 proceeds to block 734. If network interface 202 determines that central facility 112 has not received the request (block 732: no), process 700 proceeds to block 732. At block 734, report generator 214 generates a report based on the processed matches.
[0169] exist Figure 7 In the example of FIG. 7 , at block 736, network interface 202 determines whether central facility 112 has received the query. If network interface 202 determines that central facility 112 has received the query (block 736: yes), process 700 proceeds to block 738. If network interface 202 determines that central facility 112 has not received the query (block 736: no), process 700 proceeds to block 736. At block 738, central facility 112 processes the query. For example, central facility 112 may process the query based on Fig. 6A and Figure 6B Process 600 is used to process the query.
[0170] exist Figure 7 In the example shown, at block 740, report generator 214 generates a report based on the query. At block 742, media processor 212 determines whether there is additional media to be analyzed. For example, media processor 212 may cause central facility 112 to process broadcast media again at a predefined frequency (e.g., every 2 months). If media processor 212 determines that there is additional broadcast media to be analyzed (block 742: yes), process 700 proceeds to block 702. If media processor 212 determines that there is no additional broadcast media to be analyzed (block 742: no), process 700 proceeds to block 744.
[0171] exist Figure 7In the example of FIG. 74 , the network interface 202 determines whether the central facility 112 is to continue operating at block 744. If the network interface 202 determines that the central facility 112 is to continue operating (block 744: yes), the process 700 proceeds to block 732. If the network interface 202 determines that the central facility 112 is not to continue operating (block 744: no), the process 700 terminates.
[0172] Figure 8 is a flow chart showing a process for adjusting a fingerprint, which may be performed to implement Figure 1 and Figure 2 The central facility 112 may be implemented by machine-readable instructions. For example, Figure 8 The processing can be achieved Fig. 6A , Figure 6B and / or Figure 7 . Boxes 612, 636, 656, 666, 708 and / or 720.
[0173] exist Figure 8 In the example of , the process begins at block 802, where the fingerprint pitch tuner 204 obtains a sample fingerprint. For example, the fingerprint pitch tuner 204 may access one or more fingerprints from the database 216 and / or from the network 104 via the network interface 202. At block 804, the fingerprint pitch tuner 204 determines whether the pitch shift increases the pitch of the audio associated with the sample fingerprint. If the fingerprint pitch tuner 204 determines that the pitch shift increases the pitch of the audio associated with the sample fingerprint (block 804: yes), then Figure 8 Processing proceeds to block 806. If the fingerprint pitch tuner 204 determines that the pitch shift does not increase the pitch of the audio associated with the sample fingerprint (block 804: No), then Figure 8 Processing proceeds to box 808.
[0174] exist Figure 8 In the example shown, at block 806, the fingerprint pitch tuner 204 decreases a bin value associated with the sample fingerprint (e.g., of the sample fingerprint, in the sample fingerprint, etc.) based on the indicated pitch shift (e.g., by, etc.). At block 808, the fingerprint pitch tuner 204 increases a bin value associated with the sample fingerprint (e.g., of the sample fingerprint, in the sample fingerprint, etc.) based on the indicated pitch shift (e.g., by, etc.). At block 810, the fingerprint pitch tuner 204 generates one or more adjusted fingerprints based on the adjusted bin values. After block 810, Figure 8 Processing returns to blocks 614 , 638 , 656 , 668 , 710 , and / or 720 .
[0175] Fig. 9 is a flow chart showing a process for adjusting a fingerprint, which may be performed to implement Figure 1 and Figure 2 The central facility 112 may be implemented by machine-readable instructions. For example, Fig. 9 The processing can be achieved Fig. 6A , Figure 6B and / or Figure 7 .box 624, 638, 668, 678, 714 and / or 722.
[0176] exist Fig. 9 In the example of FIG. 1 , the process begins at block 902, where the fingerprint speed tuner 206 obtains a sample fingerprint. For example, the fingerprint speed tuner 206 may access one or more sample fingerprints from the database 216 and / or from the network 104 via the network interface 202. At block 904, the fingerprint speed tuner 206 determines whether the time offset increases the playback speed of the audio associated with the sample fingerprint. If the fingerprint speed tuner 206 determines that the time offset increases the playback speed of the audio associated with the sample fingerprint (block 904: yes), then Fig. 9 Processing proceeds to block 906. If the fingerprint speed tuner 206 determines that the time shift does not increase the playback speed of the audio associated with the sample fingerprint (block 904: No), then Fig. 9 Processing proceeds to box 908.
[0177] exist Fig. 9 In the example shown, at block 906, the fingerprint speed tuner 206 modifies the sample fingerprint based on the increased playback speed. At block 908, the fingerprint speed tuner 206 modifies the sample fingerprint based on the increased playback speed. At block 910, the fingerprint speed tuner 206 generates one or more adjusted fingerprints based on these modifications. After block 910, Fig. 9 Processing returns to blocks 626 , 640 , 670 , 680 , 716 , and / or 722 .
[0178] Fig.10 is a flow chart representing a process 1000 for transmitting information associated with a pitch shift, a time shift, and / or a resampling rate, which process 1000 may be implemented using a method that may be executed to implement Figure 1 and Figure 3 The application is implemented by machine-readable instructions.
[0179] At block 1002, the example component interface 302 obtains an audio signal. For example, the example component interface 302 may obtain an audio signal from a microphone embedded in and / or attached to the example end-user device 108. Additionally or alternatively, the example component interface 302 may interface with a processing component and / or another application to obtain an audio signal output by the example end-user device 108. For example, the component interface 302 may obtain an audio signal from a different application (e.g., an Internet radio application, a browser, a podcast, a streaming music service, etc.) running on the example end-user device 108 that is outputting audio.
[0180] At block 1004, the example fingerprint generator 304 generates one or more fingerprints from the audio signal, as described above in conjunction with Figure 3 At block 1006, the example network interface 306 determines whether an adjustment instruction (eg, a query) corresponding to a pitch shift, a time shift, and / or a resampling rate has been defined by the client (eg, via Figure 1 102 ). For example, when the application 114 is owned by a client, the application may be accompanied by such adjustment instructions: the adjustment instructions correspond to how much and / or what type of pitch shift, time shift, and / or resampling rate to perform on the fingerprint in an attempt to recognize the audio. If the example network interface 306 determines that the adjustment instructions corresponding to the pitch shift, time shift, and / or resampling rate have been defined by the client (box 1006: yes), control continues to box 1012, as further described below. If the example network interface 306 determines that the adjustment instructions do not correspond to the pitch shift, time shift, and / or resampling rate that have been defined by the client (box 1006: no), the example user interface 308 displays an adjustment instruction prompt (box 1008) to the end user of the example end-user device 108. The prompt prompts the user to select how much and / or what type of pitch shift, time shift, and / or resampling rate to perform on the fingerprint when attempting to recognize the audio at the central facility 112.
[0181] At block 1010, the example component interface 302 receives an adjustment instruction from a user via the example user interface 308. At block 1012, the example network interface 306 sends the fingerprint to the example central facility 112 along with the adjustment instruction / query (e.g., pitch shift, time shift, and / or resampling rate query instruction) from the client and / or end user of the example end-user device 108. As described above, the example central facility 112 adjusts the fingerprint based on the adjustment instruction (e.g., pitch shift, time shift, and / or resampling rate query instruction) and compares the received fingerprint and / or the adjusted fingerprint to a fingerprint database with identification information to identify a match. The central facility 112 returns the query results to the example application 114 via the example network interface 306. For example, if the query results in the identification of audio that has been pitch shifted, the central facility 112 returns information corresponding to the identified audio, as well as an indication of the audio being pitch shifted and how much the audio is pitch shifted.
[0182] At block 1014, the example network interface 306 receives a response to the fingerprint query (e.g., whether the fingerprint corresponds to a match, information corresponding to the matching audio, whether the audio is pitch-shifted, time-shifted, and / or resampled, and / or how the audio is pitch-shifted, time-shifted, and / or resampled). At block 1016, the example report generator 310 determines whether the response corresponds to matching audio that is pitch-shifted, time-shifted, and / or resampled. If the example report generator 310 determines that the response corresponds to matching audio that is pitch-shifted, time-shifted, and / or resampled (block 1016: yes), the example report generator 310 generates a report that identifies information corresponding to the matched audio (e.g., song name, album name, artist, publisher, etc.) and pitch-shifted, time-shifted, and / or resampled information (e.g., how the audio is pitch-shifted, time-shifted, and / or resampled based on the fingerprint query) (block 1020). For example, if the fingerprint of the audio matches a sped-up version of "Cover Me Up" by Jason Isbell (e.g., corresponds to a time and pitch shift), the report generator 310 generates a report including identification information (e.g., identifying that the obtained audio corresponds to a song, artist, record label, etc.), and identifies how much the song has been pitch and time shifted and the pitch and time shifts increased. The report can be a report in a displayable form (e.g., a pdf, word document, photo, text, picture, etc.) and / or one or more data packets including identification information that can be stored and / or sent to another component or device (e.g., locally or externally).
[0183] If the example report generator 310 determines that the response does not correspond to matching audio that was pitch-shifted, time-shifted, and / or resampled (block 1016: No), the example report generator 310 generates a report identifying information corresponding to the matched audio or generates a report identifying that the fingerprint does not match (block 1018). At block 1022, the example component interface 302 and / or the example network interface 306 sends the report. In some examples, the network interface 306 can send the report to a local storage unit (e.g., Fig.12 The example local memory 1213, the example volatile memory 1214 and / or the example non-volatile memory 1216 of the example application 114 can be used for storage. In such an example, the example application 114 and / or another component can use the stored information to adjust subsequent adjustment instructions. For example, if a specific pitch offset, time offset and / or resampling rate triggers a large number of matches or few or no matches, the specific pitch offset, time offset and / or resampling rate can be prioritized, reduced and / or eliminated in subsequent queries. Additionally or alternatively, the example component interface 302 can send a report to the user interface 308 to display the report and / or the information in the report to the end user. Additionally or alternatively, the example network interface 306 can send a report as a data packet to the example client device 102 and / or another device (e.g., a server) of the client to provide information about how the audio is adjusted when it is obtained to the client. In this way, the client has information that can be used to modify the adjustment instructions for subsequent queries.
[0184] Fig.11 is constructed to execute Fig. 6A , Figure 6B , Figure 7 , Figure 8 and / or Fig. 9 Instructions to achieve Figure 2 1100 of an example processor platform 1100 of a central facility 112 of FIG. 1100. For example, the processor platform 1100 may be a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cell phone, a smart phone, an iPad, etc.), or a processor platform 1100 of an example processor platform 1100 of a central facility 112 of FIG. TM tablet computer), personal digital assistant (PDA), Internet appliance, DVD player, CD player, digital video recorder, Blu-ray player, game console, set-top box, head-mounted device or other wearable device, or any other type of computing device.
[0185] The processor platform 1100 of the illustrated example includes a processor 1112. The processor 1112 of the illustrated example is hardware. For example, the processor 1112 may be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor may be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 1112 implements Figure 2 An example network interface 202, an example fingerprint pitch tuner 204, an example fingerprint speed tuner 206, an example query comparator 208, an example fingerprint generator 210, an example media processor 212, and an example report generator 214 are provided.
[0186] The processor 1112 of the illustrated example includes a local memory 1113 (e.g., cache). The processor 1112 of the illustrated example communicates with a main memory including a volatile memory 1114 and a non-volatile memory 1116 via a bus 1118. The volatile memory 1114 may be comprised of a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), Dynamic Random Access Memory The non-volatile memory 1116 may be implemented by a flash memory and / or any other desired type of memory device. Access to the main memory 1114, 1116 is controlled by a memory controller. In this example, the local memory 1113 implements Figure 2 Additionally or alternatively, the example volatile memory 1114 and the example non-volatile memory 1116 may implement Figure 2 An example database 216 is provided.
[0187] The processor platform 1100 of the illustrated example also includes an interface circuit 1120. The interface circuit 1120 may be implemented by any type of interface standard (such as an Ethernet interface, a universal serial bus (USB), interface, near field communication (NFC) interface and / or PCI express interface).
[0188] In the example shown, one or more input devices 1122 are connected to the interface circuit 1120. The input devices 1122 allow a user to input data and / or commands into the processor 1112. For example, the input devices may be implemented by an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touch screen, a trackpad, a trackball, isopoint, and / or a voice recognition system.
[0189] One or more output devices 1124 are also connected to the interface circuit 1120 of the illustrated example. The output device 1124 can be implemented, for example, by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-plane switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Therefore, the interface circuit 1120 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0190] The interface circuitry 1120 of the illustrated example also includes communication devices (such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces) to facilitate exchanging data with external machines (e.g., any kind of computing device) via the network 1126. For example, the communication may be via an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a direct-line wireless system, a cellular telephone system, etc.
[0191] The processor platform 1100 of the illustrated example also includes one or more mass storage devices 1128 for storing software and / or data. Examples of such mass storage devices 1128 include floppy disk drives, hard disk drives, optical disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0192] Fig. 6A , Figure 6B , Figure 7 , Figure 8 and / or Fig. 9 The machine-executable instructions 1132 may be stored in the mass storage device 1128, the volatile memory 1114, the non-volatile memory 1116, and / or on a removable, non-transitory computer-readable storage medium such as a CD or DVD.
[0193] Fig.12 is constructed to execute Fig.10 Instructions to achieve Figure 3 114. For example, the processor platform 1200 can be a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cell phone, a smart phone, an iPad, etc.). TM tablet computer), personal digital assistant (PDA), Internet appliance, DVD player, CD player, digital video recorder, Blu-ray player, game console, set-top box, or any other type of computing device.
[0194] The processor platform 1200 of the illustrated example includes a processor 1212. The processor 1212 of the illustrated example is hardware. For example, the processor 1212 may be implemented by one or more integrated circuits, logic circuits, microprocessors, GPUs, DSPs, or controllers from any desired family or manufacturer. The hardware processor may be a semiconductor-based (e.g., silicon-based) device. In this example, the processor implements Figure 3 An example component interface 302 , an example fingerprint generator 304 , an example network interface 306 , an example user interface 308 , and an example report generator 310 .
[0195] The processor 1212 of the illustrated example includes a local memory 1213 (e.g., cache). The processor 1212 of the illustrated example communicates with a main memory including a volatile memory 1214 and a non-volatile memory 1216 via a bus 1218. The volatile memory 1214 may be comprised of a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), Dynamic Random Access Memory The non-volatile memory 1216 may be implemented by a flash memory and / or any other desired type of memory device. Access to the main memory 1214, 1216 is controlled by a memory controller.
[0196] The processor platform 1200 of the illustrated example also includes an interface circuit 1220. The interface circuit 1220 may be implemented by any type of interface standard (such as an Ethernet interface, a universal serial bus (USB), interface, near field communication (NFC) interface and / or PCI express interface).
[0197] In the example shown, one or more input devices 1222 are connected to the interface circuit 1220. The input devices 1222 allow a user to input data and / or commands into the processor 1212. For example, the input devices may be implemented by an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touch screen, a trackpad, a trackball, isopoint, and / or a voice recognition system.
[0198] One or more output devices 1224 are also connected to the interface circuit 1220 of the illustrated example. The output device 1224 can be implemented, for example, by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube display (CRT), an in-plane switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Therefore, the interface circuit 1220 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics driver processor.
[0199] The interface circuitry 1220 of the illustrated example also includes communication devices (such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces) to facilitate exchanging data with external machines (e.g., any kind of computing device) via the network 1226. For example, the communication may be via an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a direct-line wireless system, a cellular telephone system, etc.
[0200] The processor platform 1200 of the illustrated example also includes one or more mass storage devices 1228 for storing software and / or data. Examples of such mass storage devices 1228 include floppy disk drives, hard disk drives, optical disk drives, Blu-ray disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0201] Fig.10 The machine-executable instructions 1232 may be stored in the mass storage device 1228, the volatile memory 1214, the non-volatile memory 1216, and / or on a removable, non-transitory computer-readable storage medium such as a CD or DVD.
[0202] Based on the foregoing, it will be understood that example methods, devices, and products for identifying media have been disclosed. Example methods, devices, and products reduce the number of computing cycles used to process media and identify information associated with the media. Some technologies used to identify media change and / or otherwise adjust audio signals, and the examples disclosed herein adjust sample fingerprints and / or reference fingerprints to identify media. Because the examples disclosed herein adjust sample fingerprints and / or reference fingerprints, the examples disclosed herein are robust in terms of pitch shift, time shift, and / or resampling of media. The disclosed methods, devices, and products improve the efficiency of using computing devices by reducing the computational burden of identifying altered media. The disclosed methods, devices, and products improve the efficiency of using computing devices by reducing the power consumption of computing devices when identifying media. In addition, the disclosed methods, devices, and products improve the efficiency of using computing devices by reducing the processing overhead used to identify media. The disclosed methods, devices, and products are therefore directed to one or more improvements in computer functions.
[0203] Example methods, apparatus, systems, and articles of manufacture for identifying media are disclosed herein. Further examples and combinations thereof include the following: Example 1 includes a non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium including instructions that, when executed, cause one or more processors to at least: generate an adjusted sample media fingerprint by applying an adjustment to the sample media fingerprint in response to a query; compare the adjusted sample media fingerprint to a reference media fingerprint; and in response to the adjusted sample media fingerprint matching the reference media fingerprint, send information associated with the reference media fingerprint and the adjustment.
[0204] Example 2 includes the non-transitory computer-readable storage medium of Example 1, wherein the query includes at least one of the following items: (1) a pitch shift value, (2) a time shift value, or (3) a resampling rate, the pitch shift value, the time shift value, and the resampling rate corresponding to a hypothesized change to the audio signal associated with the sample media fingerprint.
[0205] Example 3 includes the non-transitory computer-readable storage medium of Example 2, wherein the instructions, when executed, cause the one or more processors to: determine whether the pitch offset value increases the pitch of the audio signal associated with the sample media fingerprint; in response to the pitch offset value increasing the pitch of the audio signal associated with the sample media fingerprint, reduce one or more bin values associated with the sample media fingerprint based on the pitch offset value; in response to the pitch offset value decreasing the pitch of the audio signal associated with the sample media fingerprint, increase the one or more bin values associated with the sample media fingerprint based on the pitch offset value; and generate the adjusted sample media fingerprint.
[0206] Example 4 includes the non-transitory computer-readable storage medium of Example 2, wherein the instructions, when executed, cause the one or more processors to: determine whether the time offset value increases a playback speed of the audio signal associated with the sample media fingerprint; in response to the time offset value increasing the playback speed of the audio signal associated with the sample media fingerprint, add a first frame to the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; in response to the time offset value decreasing the playback speed of the audio signal associated with the sample media fingerprint, delete a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; and generate the adjusted sample media fingerprint.
[0207] Example 5 includes the non-transitory computer-readable storage medium of Example 2, wherein the instructions, when executed, cause the one or more processors to: generate a first adjusted sample media fingerprint by applying the pitch offset value to the sample media fingerprint; generate a second adjusted sample media fingerprint by applying the time offset value to the sample media fingerprint; generate a third adjusted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and compare the first adjusted sample media fingerprint, the second adjusted sample media fingerprint, and the third adjusted sample media fingerprint to the reference media fingerprint.
[0208] Example 6 includes the non-transitory computer-readable storage medium of Example 2, wherein the instructions, when executed, cause the one or more processors to: select the resampling rate; generate a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and generate the adjusted sample media fingerprint by applying the resampling rate to the pitch-shifted sample media fingerprint.
[0209] Example 7 includes the non-transitory computer-readable storage medium of Example 1, wherein the information associated with the reference media fingerprint identifies at least one of: a title of a song associated with the reference media fingerprint, or an artist of a song associated with the reference media fingerprint.
[0210] Example 8 includes an apparatus: a fingerprint tuner that generates an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint in response to a query; a comparator that compares the adjusted sample media fingerprint with a reference media fingerprint; and a network interface that sends information associated with the reference media fingerprint and the adjustment in response to the adjusted sample media fingerprint matching the reference media fingerprint.
[0211] Example 9 includes the apparatus of Example 8, wherein the query comprises at least one of: (1) a pitch offset value, (2) a time offset value, or (3) a resampling rate, the pitch offset value, the time offset value, and the resampling rate corresponding to a hypothesized change to an audio signal associated with the sample media fingerprint.
[0212] Example 10 includes the apparatus of Example 9, further comprising: a fingerprint pitch tuner, which: determines whether the pitch offset value increases the pitch of the audio signal associated with the sample media fingerprint; in response to the pitch offset value increasing the pitch of the audio signal associated with the sample media fingerprint, reduces one or more bin values associated with the sample media fingerprint based on the pitch offset value; in response to the pitch offset value decreasing the pitch of the audio signal associated with the sample media fingerprint, increases the one or more bin values associated with the sample media fingerprint based on the pitch offset value; and generates the adjusted sample media fingerprint.
[0213] Example 11 includes the apparatus of Example 9, further comprising: a fingerprint speed tuner, the fingerprint speed tuner: determining whether the time offset value increases the playback speed of the audio signal associated with the sample media fingerprint; in response to the time offset value increasing the playback speed of the audio signal associated with the sample media fingerprint, adding a first frame to the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; in response to the time offset value decreasing the playback speed of the audio signal associated with the sample media fingerprint, deleting a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; and generating the adjusted sample media fingerprint.
[0214] Example 12 includes the apparatus of Example 9, further comprising: a fingerprint pitch tuner that generates a first adjusted sample media fingerprint by applying the pitch offset value to the sample media fingerprint; a fingerprint speed tuner that generates a second adjusted sample media fingerprint by applying the time offset value to the sample media fingerprint; at least one of the fingerprint pitch tuner or the fingerprint speed tuner that generates a third adjusted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and the comparator compares the first adjusted sample media fingerprint, the second adjusted sample media fingerprint, and the third adjusted sample media fingerprint with the reference media fingerprint.
[0215] Example 13 includes the apparatus of Example 9, further comprising: a network interface that selects the resampling rate; a fingerprint pitch tuner that generates a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and a fingerprint speed tuner that generates the adjusted sample media fingerprint by applying the resampling rate to the pitch-shifted sample media fingerprint.
[0216] Example 14 includes the apparatus of Example 8, wherein the information associated with the reference media fingerprint identifies at least one of: a title of a song associated with the reference media fingerprint, or an artist of the song associated with the reference media fingerprint.
[0217] Example 15 includes a method that: in response to a query, generates an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint; compares the adjusted sample media fingerprint with a reference media fingerprint; and in response to the adjusted sample media fingerprint matching the reference media fingerprint, sends information associated with the reference media fingerprint and the adjustment.
[0218] Example 16 includes the method of Example 15, wherein the query includes at least one of the following items: (1) a pitch offset value, (2) a time offset value, or (3) a resampling rate, the pitch offset value, the time offset value, and the resampling rate corresponding to a hypothesized change to the audio signal associated with the sample media fingerprint.
[0219] Example 17 includes the method of Example 16, further comprising: determining whether the pitch offset value increases the pitch of the audio signal associated with the sample media fingerprint; in response to the pitch offset value increasing the pitch of the audio signal associated with the sample media fingerprint, reducing one or more bin values associated with the sample media fingerprint based on the pitch offset value; in response to the pitch offset value decreasing the pitch of the audio signal associated with the sample media fingerprint, increasing the one or more bin values associated with the sample media fingerprint based on the pitch offset value; and generating the adjusted sample media fingerprint.
[0220] Example 18 includes the method of Example 16, further comprising: determining whether the time offset value increases the playback speed of the audio signal associated with the sample media fingerprint; in response to the time offset value increasing the playback speed of the audio signal associated with the sample media fingerprint, adding a first frame to the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; in response to the time offset value decreasing the playback speed of the audio signal associated with the sample media fingerprint, deleting a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; and generating the adjusted sample media fingerprint.
[0221] Example 19 includes the method of Example 16, further comprising: generating a first adjusted sample media fingerprint by applying the pitch offset value to the sample media fingerprint; generating a second adjusted sample media fingerprint by applying the time offset value to the sample media fingerprint; generating a third adjusted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and comparing the first adjusted sample media fingerprint, the second adjusted sample media fingerprint, and the third adjusted sample media fingerprint with the reference media fingerprint.
[0222] Example 20 includes the method of Example 16, further comprising: selecting the resampling rate; generating a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and generating the adjusted sample media fingerprint by applying the resampling rate to the pitch-shifted sample media fingerprint.
[0223] Example 21 includes a non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to at least: generate a fingerprint from an audio signal; send the fingerprint and adjustment instructions to a central facility to facilitate a query, the adjustment instructions identifying at least one of a pitch shift, a time shift, or a resampling rate; receive a response including an identifier of the audio signal and information corresponding to how the audio signal was adjusted; and store information indicating the identifier and the information in a database.
[0224] Example 22 includes the non-transitory computer-readable storage medium of Example 21, wherein the instructions cause one or more processors to send a prompt to a user identifying an adjustment instruction.
[0225] Example 23 includes the non-transitory computer-readable storage medium of Example 21, wherein the adjustment instruction is defined by a client.
[0226] Example 24 includes the non-transitory computer-readable storage medium of Example 21, wherein the information is included in a data packet sent to the client.
[0227] Example 25 includes the non-transitory computer-readable storage medium of Example 21, wherein the information can be displayed on a user interface.
[0228] Example 26 includes the non-transitory computer-readable storage medium of Example 21, wherein the instructions cause one or more processors to modify the adjustment instructions based on the information for subsequent fingerprint recognition.
[0229] Example 27 includes the non-transitory computer-readable storage medium of Example 21, wherein the instructions cause the one or more processors to send a data packet including the information to a client.
[0230] Example 28 includes an apparatus comprising a fingerprint generator for generating a fingerprint from an audio signal, and a network interface that: sends the fingerprint and adjustment instructions to a central facility to facilitate a query, the adjustment instructions identifying at least one of a pitch shift or a time shift; receives a response including an identifier of the audio signal and information corresponding to how the audio signal was adjusted; and stores information indicating the identifier and the information in a database.
[0231] Example 29 includes the apparatus of Example 28, further comprising a user interface that displays a prompt to a user to identify the adjustment instruction.
[0232] Example 30 includes the apparatus of Example 28, wherein the adjustment instruction is defined by a client.
[0233] Example 31 includes the apparatus of Example 28, wherein the network interface is used to send a data packet including the information to a client.
[0234] Example 32 includes the apparatus of Example 28, further comprising a user interface for displaying the information to a user.
[0235] Example 33 includes the apparatus of Example 28, wherein the information is used to modify the adjustment instruction for subsequent fingerprint recognition.
[0236] Example 34 includes the apparatus of Example 28, wherein the network interface is to send a data packet including the information to a client.
[0237] Example 35 includes a method comprising: generating a fingerprint from an audio signal by executing instructions using a processor; sending the fingerprint and adjustment instructions to a central facility to facilitate a query, the adjustment instructions identifying at least one of a pitch shift or a time shift; receiving a response including an identifier of the audio signal and information corresponding to how the audio signal was adjusted; and storing information indicating the identifier and the information in a database.
[0238] Example 36 includes the method of Example 35, further comprising sending a prompt to a user to identify the adjustment instruction.
[0239] Example 37 includes the method of Example 35, wherein the adjustment instructions are defined by a client.
[0240] Example 38 includes the method of Example 35, further comprising sending the information as a data packet to a client.
[0241] Example 39 includes the method of Example 35, further comprising displaying the information to a user.
[0242] Example 40 includes the method of Example 35, wherein the information is used to modify the adjustment instruction for subsequent fingerprint recognition.
[0243] Example 41 includes a non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to at least: compare at least one of (a) a pitch-shifted sample media fingerprint, (b) a time-shifted sample media fingerprint, or (c) a resampled sample media fingerprint to a reference media fingerprint; in response to a match between any of the (a) pitch-shifted sample media fingerprint, (b) a time-shifted sample media fingerprint, or (c) a resampled sample media fingerprint and the reference media fingerprint, generate one or more indications of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate that caused the match; in response to collecting broadcast media for a threshold time period, process the one or more indications; and in response to a request for a recommendation of information associated with a query, send a recommendation including one or more frequencies of occurrence of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate.
[0244] Example 42 includes the non-transitory computer-readable storage medium of Example 41, wherein the instructions, when executed, cause the one or more processors to: collect broadcast media; generate a sample media fingerprint based on the broadcast media; and generate at least one of (a) a pitch-shifted sample media fingerprint, (b) a time-shifted sample media fingerprint, or (c) a resampled sample media fingerprint by applying at least one of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate to the sample media fingerprint.
[0245] Example 43 includes the non-transitory computer-readable storage medium of Example 42, wherein the instructions, when executed, cause one or more processors to: determine whether the pitch offset value increases the pitch of an audio signal associated with a sample media fingerprint; in response to the pitch offset value increasing the pitch of the audio signal associated with the sample media fingerprint, reduce one or more bin values associated with the sample media fingerprint based on the pitch offset value; in response to the pitch offset value decreasing the pitch of the audio signal associated with the sample media fingerprint, increase one or more bin values associated with the sample media fingerprint based on the pitch offset value; and generate a pitch-shifted sample media fingerprint.
[0246] Example 44 includes the non-transitory computer-readable storage medium of Example 42, wherein the instructions, when executed, cause the one or more processors to: determine whether the time offset value increases a playback speed of an audio signal associated with the sample media fingerprint; in response to the time offset value increasing the playback speed of the audio signal associated with the sample media fingerprint, add a first frame to the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; in response to the time offset value decreasing the playback speed of the audio signal associated with the sample media fingerprint, delete a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; and generate a time-shifted sample media fingerprint.
[0247] Example 45 includes the non-transitory computer-readable storage medium of Example 42, wherein the instructions, when executed, cause one or more processors to: select a resampling rate; generate a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and generate a resampled sample media fingerprint by applying the resampling rate to the pitch-shifted sample media fingerprint.
[0248] Example 46 includes the non-transitory computer-readable storage medium of Example 41, wherein the recommendation includes (a) one or more pitch offset values, (b) one or more time offset values, or (c) one or more resampling rates having a higher frequency of occurrence than (a) one or more additional pitch offset values, (b) one or more additional time offset values, or (c) one or more additional resampling rates, respectively.
[0249] Example 47 includes the non-transitory computer-readable storage medium of Example 46, wherein the instructions, when executed, cause one or more processors to: in response to a query including a recommendation, compare the sample media fingerprint with (a) one or more pitch offset values, (b) one or more time offset values, or (c) one or more resampling rates.
[0250] Example 48 includes a device comprising a comparator that compares at least one of (a) a pitch-shifted sample media fingerprint, (b) a time-shifted sample media fingerprint, or (c) a resampled sample media fingerprint to a reference media fingerprint; in response to a match between any of the (a) pitch-shifted sample media fingerprint, (b) time-shifted sample media fingerprint, or (c) resampled sample media fingerprint and the reference media fingerprint, generates one or more indications of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate that caused the match; a processor that processes the one or more indications in response to collecting broadcast media for a threshold time period; and a network interface that sends a recommendation including one or more frequencies of occurrence of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate in response to a request for a recommendation of information associated with a query.
[0251] Example 49 includes the apparatus of Example 48, wherein the network interface collects broadcast media; the apparatus further comprises a fingerprint generator for generating a sample media fingerprint based on the broadcast media, and a fingerprint tuner, wherein the fingerprint tuner generates at least one of (a) a pitch-shifted sample media fingerprint, (b) a time-shifted sample media fingerprint, or (c) a resampled sample media fingerprint by applying at least one of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate to the sample media fingerprint.
[0252] Example 50 includes the apparatus of Example 49, further comprising a fingerprint pitch tuner that: determines whether a pitch offset value increases the pitch of an audio signal associated with a sample media fingerprint; in response to the pitch offset value increasing the pitch of the audio signal associated with the sample media fingerprint, decreases one or more bin values associated with the sample media fingerprint based on the pitch offset value; in response to the pitch offset value decreasing the pitch of the audio signal associated with the sample media fingerprint, increases one or more bin values associated with the sample media fingerprint based on the pitch offset value; and generates a pitch-shifted sample media fingerprint.
[0253] Example 51 includes the apparatus of Example 49, further comprising a fingerprint speed tuner that: determines whether the time offset value increases the playback speed of the audio signal associated with the sample media fingerprint; in response to the time offset value increasing the playback speed of the audio signal associated with the sample media fingerprint, adds a first frame to the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; in response to the time offset value decreasing the playback speed of the audio signal associated with the sample media fingerprint, deletes a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; and generates a time-shifted sample media fingerprint.
[0254] Example 52 includes the apparatus of Example 49, further comprising: the network interface selecting a resampling rate; the fingerprint pitch tuner generating a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and at least one of the fingerprint velocity tuner or the fingerprint pitch tuner generating a resampled sample media fingerprint by applying the resampling rate to the pitch-shifted sample media fingerprint.
[0255] Example 53 includes the apparatus of Example 49, wherein the recommendation includes (a) one or more pitch offset values, (b) one or more time offset values, or (c) one or more resampling rates having a higher frequency of occurrence than (a) one or more additional pitch offset values, (b) one or more additional time offset values, or (c) one or more additional resampling rates, respectively.
[0256] Example 54 includes the apparatus of Example 53, wherein the comparator is configured to: in response to a query comprising a recommendation, compare the sample media fingerprint with (a) one or more pitch offset values, (b) one or more time offset values, or (c) one or more resampling rates.
[0257] Example 55 includes a method comprising: comparing at least one of (a) a pitch-shifted sample media fingerprint, (b) a time-shifted sample media fingerprint, or (c) a resampled sample media fingerprint to a reference media fingerprint; in response to a match between any of the (a) pitch-shifted sample media fingerprint, (b) time-shifted sample media fingerprint, or (c) resampled sample media fingerprint and the reference media fingerprint, generating one or more indications of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate that caused the match; in response to collecting broadcast media for a threshold time period, processing one or more indications; and in response to a request for a recommendation of information associated with a query, sending a recommendation including one or more frequencies of occurrence of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate.
[0258] Example 56 includes the method of Example 55, which further includes: collecting broadcast media; generating a sample media fingerprint based on the broadcast media; and generating at least one of (a) a pitch-shifted sample media fingerprint, (b) a time-shifted sample media fingerprint, or (c) a resampled sample media fingerprint by applying at least one of (a) a pitch shift value, (b) a time shift value, or (c) a resampling rate to the sample media fingerprint.
[0259] Example 57 includes the method of Example 56, which further includes: determining whether the pitch offset value increases the pitch of an audio signal associated with the sample media fingerprint; in response to the pitch offset value increasing the pitch of the audio signal associated with the sample media fingerprint, reducing one or more bin values associated with the sample media fingerprint based on the pitch offset value; in response to the pitch offset value decreasing the pitch of the audio signal associated with the sample media fingerprint, increasing one or more bin values associated with the sample media fingerprint based on the pitch offset value; and generating a pitch-shifted sample media fingerprint.
[0260] Example 58 includes the method of Example 56, further comprising: determining whether the time offset value increases the playback speed of the audio signal associated with the sample media fingerprint; in response to the time offset value increasing the playback speed of the audio signal associated with the sample media fingerprint, adding a first frame to the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; in response to the time offset value decreasing the playback speed of the audio signal associated with the sample media fingerprint, deleting a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; and generating a time-shifted sample media fingerprint.
[0261] Example 59 includes the method of Example 56, further comprising: selecting a resampling rate; generating a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and generating a resampled sample media fingerprint by applying the resampling rate to the pitch-shifted sample media fingerprint.
[0262] Example 60 includes the method of Example 59, wherein the recommendation includes (a) one or more pitch offset values, (b) one or more time offset values, or (c) one or more resampling rates having a higher frequency of occurrence than (a) one or more additional pitch offset values, (b) one or more additional time offset values, or (c) one or more additional resampling rates, respectively.
[0263] Although certain example methods, apparatus, and articles of manufacture have been disclosed herein, the scope of coverage of this patent is not limited thereto. On the contrary, this patent covers all methods, apparatus, and articles of manufacture that fully fall within the scope of the claims of this patent.
[0264] The following claims are incorporated into this Detailed Description by reference, with each claim standing on its own as a separate embodiment of the present disclosure.
Claims
1. A non-transitory computer-readable storage medium comprising instructions that, when executed, cause one or more processors to at least: In response to the query, generating an adjusted sample media fingerprint by applying the adjustment to the sample media fingerprint, the sample media fingerprint being associated with the audio signal, in, The query includes at least one of: (1) a pitch shift value, (2) a time shift value, or (3) a resampling rate, wherein the sample media fingerprint includes a plurality of time-frequency bins representing the audio signal, and wherein generating an adjusted sample media fingerprint includes: determining whether the time offset value increases a playback speed of the audio signal associated with the sample media fingerprint; in response to the time offset value increasing a playback speed of the audio signal associated with the sample media fingerprint, adding a first frame to the sample media fingerprint at a location in the sample media fingerprint corresponding to the time offset value; In response to the time offset value reducing a playback speed of the audio signal associated with the sample media fingerprint, deleting a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; comparing the adjusted sample media fingerprint with a reference media fingerprint; and In response to the adjusted sample media fingerprint matching the reference media fingerprint, information associated with the reference media fingerprint and the adjustment is sent.
2. The non-transitory computer-readable storage medium according to claim 1, in, The pitch shift value, the time shift value, and the resampling rate correspond to a hypothesized alteration of the audio signal associated with the sample media fingerprint.
3. The non-transitory computer-readable storage medium according to claim 1, in, The instructions, when executed, cause the one or more processors to: determining whether the pitch shift value increases the pitch of the audio signal associated with the sample media fingerprint; in response to the pitch shift value increasing the pitch of the audio signal associated with the sample media fingerprint, decreasing bin values of one or more time-frequency bins associated with the sample media fingerprint based on the pitch shift value; In response to the pitch shift value lowering the pitch of the audio signal associated with the sample media fingerprint, increasing bin values of one or more time-frequency bins associated with the sample media fingerprint based on the pitch shift value; as well as The adjusted sample media fingerprint is generated.
4. The non-transitory computer-readable storage medium according to claim 1, in, The instructions, when executed, cause the one or more processors to: generating a first adjusted sample media fingerprint by applying the pitch shift value to the sample media fingerprint; generating a second adjusted sample media fingerprint by applying the time offset value to the sample media fingerprint; generating a third adjusted sample media fingerprint by applying the resampling rate to the sample media fingerprint; as well as The first adjusted sample media fingerprint, the second adjusted sample media fingerprint, and the third adjusted sample media fingerprint are compared to the reference media fingerprint.
5. The non-transitory computer-readable storage medium according to claim 1, in, The instructions, when executed, cause the one or more processors to: selecting the resampling rate; generating a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; as well as The adjusted sample media fingerprint is generated by applying the resampling rate to the pitch-shifted sample media fingerprint.
6. The non-transitory computer-readable storage medium according to claim 1, in, The information associated with the reference media fingerprint identifies at least one of: a title of a song associated with the reference media fingerprint, or an artist of the song associated with the reference media fingerprint.
7. A device for identifying media, the device include: A fingerprint tuner, the fingerprint tuner generating an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint in response to a query, the sample media fingerprint being associated with an audio signal, wherein the query comprises at least one of: (1) a pitch shift value, (2) a time shift value, or (3) a resampling rate, wherein the sample media fingerprint comprises a plurality of time-frequency bins representing the audio signal, wherein generating the adjusted sample media fingerprint comprises: determining whether the time offset value increases a playback speed of the audio signal associated with the sample media fingerprint; in response to the time offset value increasing a playback speed of the audio signal associated with the sample media fingerprint, adding a first frame to the sample media fingerprint at a location in the sample media fingerprint corresponding to the time offset value; In response to the time offset value reducing a playback speed of the audio signal associated with the sample media fingerprint, deleting a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; a comparator that compares the adjusted sample media fingerprint with a reference media fingerprint; and A network interface is provided to transmit information associated with the reference media fingerprint and the adjustment in response to the adjusted sample media fingerprint matching the reference media fingerprint.
8. The device according to claim 7, in, The pitch shift value, the time shift value, and the resampling rate correspond to a hypothesized alteration of the audio signal associated with the sample media fingerprint.
9. The device according to claim 7, further comprising: include: Fingerprint Pitch Tuner, the Fingerprint Pitch Tuner: determining whether the pitch shift value increases the pitch of the audio signal associated with the sample media fingerprint; in response to the pitch shift value increasing the pitch of the audio signal associated with the sample media fingerprint, decreasing bin values of one or more time-frequency bins associated with the sample media fingerprint based on the pitch shift value; In response to the pitch shift value lowering the pitch of the audio signal associated with the sample media fingerprint, increasing bin values of one or more time-frequency bins associated with the sample media fingerprint based on the pitch shift value; as well as The adjusted sample media fingerprint is generated.
10. The device according to claim 7, further comprising: include: a fingerprint pitch tuner that generates a first adjusted sample media fingerprint by applying the pitch offset value to the sample media fingerprint; a fingerprint speed tuner that generates a second adjusted sample media fingerprint by applying the time offset value to the sample media fingerprint; At least one of the fingerprint pitch tuner or the fingerprint velocity tuner generates a third adjusted sample media fingerprint by applying the resampling rate to the sample media fingerprint; and The comparator compares the first adjusted sample media fingerprint, the second adjusted sample media fingerprint, and the third adjusted sample media fingerprint with the reference media fingerprint.
11. The device according to claim 7, further comprising: include: a network interface, the network interface selecting the resampling rate; a fingerprint pitch tuner that generates a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; as well as A fingerprint tempo tuner is provided to generate the adjusted sample media fingerprint by applying the resampling rate to the pitch-shifted sample media fingerprint.
12. The device according to claim 7, in, The information associated with the reference media fingerprint identifies at least one of: a title of a song associated with the reference media fingerprint, or an artist of the song associated with the reference media fingerprint.
13. A method for identifying media, the method comprising: include: In response to a query, generating an adjusted sample media fingerprint by applying an adjustment to a sample media fingerprint, the sample media fingerprint being associated with an audio signal, wherein the query comprises at least one of: (1) a pitch shift value, (2) a time shift value, or (3) a resampling rate, wherein the sample media fingerprint comprises a plurality of time-frequency bins representing the audio signal, wherein generating the adjusted sample media fingerprint comprises: determining whether the time offset value increases a playback speed of the audio signal associated with the sample media fingerprint; in response to the time offset value increasing a playback speed of the audio signal associated with the sample media fingerprint, adding a first frame to the sample media fingerprint at a location in the sample media fingerprint corresponding to the time offset value; In response to the time offset value reducing a playback speed of the audio signal associated with the sample media fingerprint, deleting a second frame of the sample media fingerprint at a position in the sample media fingerprint corresponding to the time offset value; comparing the adjusted sample media fingerprint with a reference media fingerprint; and In response to the adjusted sample media fingerprint matching the reference media fingerprint, information associated with the reference media fingerprint and the adjustment is sent.
14. The method according to claim 13, in, The pitch shift value, the time shift value, and the resampling rate correspond to a hypothesized alteration of the audio signal associated with the sample media fingerprint.
15. The method according to claim 13, further comprising: include: determining whether the pitch shift value increases the pitch of the audio signal associated with the sample media fingerprint; in response to the pitch shift value increasing the pitch of the audio signal associated with the sample media fingerprint, decreasing bin values of one or more time-frequency bins associated with the sample media fingerprint based on the pitch shift value; In response to the pitch shift value lowering the pitch of the audio signal associated with the sample media fingerprint, increasing bin values of one or more time-frequency bins associated with the sample media fingerprint based on the pitch shift value; as well as The adjusted sample media fingerprint is generated.
16. The method according to claim 13, further comprising: include: generating a first adjusted sample media fingerprint by applying the pitch shift value to the sample media fingerprint; generating a second adjusted sample media fingerprint by applying the time offset value to the sample media fingerprint; generating a third adjusted sample media fingerprint by applying the resampling rate to the sample media fingerprint; as well as The first adjusted sample media fingerprint, the second adjusted sample media fingerprint, and the third adjusted sample media fingerprint are compared to the reference media fingerprint.
17. The method according to claim 13, further comprising: include: selecting the resampling rate; generating a pitch-shifted sample media fingerprint by applying the resampling rate to the sample media fingerprint; as well as The adjusted sample media fingerprint is generated by applying the resampling rate to the pitch-shifted sample media fingerprint.
Citation Information
Patent Citations
Methods and apparatus to fingerprint an audio signal via normalization
US20200082835A1
Audience measurement system utilizing ancillary codes and passive signatures
US5481294A
Detecting distorted audio signals based on audio fingerprinting
US20150199974A1
Pitch shift resistant audio matching
US9052986B1