Database for efficient fuzzy matching

a database and fuzzy matching technology, applied in the field of efficient fuzzy matching, can solve the problems of difficult problem solving, inability to search each of the possible samples in the database to find a match, and inability to match with a segment in the repository, so as to reduce the number of false negatives

Inactive Publication Date: 2007-12-20
ETSY INC
View PDF14 Cites 6 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Benefits of technology

[0009] In one embodiment, a database includes a repository of data segments to be searched, called standard streams. Rather than searching all possible segments of each standard stream, the database includes a set of index files that reference a number of different segments in the repository. Each index file provides information about whether various data segments in the repository are likely to match a given test stream, although in the presence of noise there may be multiple possible matches. By consulting a number of the index files, a searching algorithm identifies a set of candidate data segments to test and thus reduces the number of streams that must be tested, thus saving computing resources that would otherwise be devoted to testing each stream for a match.
[0012] It can be appreciated that there are no false positives within the given error tolerance, as the final test preferably returns only those streams from the repository that actually matching the test stream within the error tolerance. Beneficially, using multiple indexes may reduce the number of false negatives, even in the presence of noise up to a 30% bit-error rate. For many practical applications, a matching algorithm need not give a perfect answer in all cases, but only in most of the cases. The error tolerance, number of indexes used, and other variables can be adjusted according to the needs of a particular application.

Problems solved by technology

An important class of problems involves searching through a data repository for a match to particular item of test data, where the data repository contains a large number of data segments.
Because of measurement noise and other real-world problems, the acquired test segment is not expected to match exactly with a segment in the repository.
This problem is made more difficult where the streams in the repository are longer than the test segment.
Although such a brute-force method would likely give a correct answer, it can also be quite inefficient.
In many applications, the repository could contain millions of streams, making searching each of the possible samples in the database to find a match impractical for real world applications.
But applying those solutions to this problem quickly becomes unmanageable for high dimensions, corresponding to a wide feature vector, as described in “Approximate Closest-Point Queries in High Dimensions,” by M. Bern, Information Processing Letters (1993).
This solution, however, does not function well in the presence of noise levels of 20% or more.
Searching time-sequenced data has also been studied, for example, in “Efficient Similarity Search in Sequence Databases,” by Agrawal, Faloutsos, and Swami, but the combination of multi-dimensional feature vectors plus time-sequencing is a difficult problem.

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • Database for efficient fuzzy matching
  • Database for efficient fuzzy matching
  • Database for efficient fuzzy matching

Examples

Experimental program
Comparison scheme
Effect test

Embodiment Construction

[0020] A database and method of matching to items in the database allow for efficient fuzzy matching of test data while avoiding the impracticalities of searching prohibitively large data repositories. FIG. 1 illustrates one example of an application for which the fuzzy matching algorithm can be used. An event or item 10 is sampled at various locations in a sequence to yield a number of frames 20 of data representative of the event or item 10 at a number of instances of the event or item. Preferably, the sampling rate is constant and is consistent across all the data in the repository and the data to be tested. The event or item 10 may be any number of things from which representative data can be obtained. For example, the event 10 may be an audio or video signal, a data signal representing a measurement over time, or any number of time-sequenced events. It may be obtained from a transmission broadcast, decoded from a digital file, or obtained in any other known way. Alternatively, ...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

PUM

No PUM Login to View More

Abstract

A database includes a repository of data segments to be searched, called standard streams. But rather than searching all possible segments of each standard stream, the database includes a set of index files that reference a number of different segments in the repository. Each index file provides information about whether various data segments in the repository are likely to match a given test stream, although in the presence of noise there may be multiple possible matches. By consulting a number of the index files, a searching algorithm identifies a set of candidate data segments to test and thus reduces the number of streams that must be tested.

Description

CROSS REFERENCE TO RELATED APPLICATIONS [0001] This application is a divisional of U.S. application Ser. No. 10 / 830,962, filed Apr. 22, 2004, which claims the benefit of U.S. Provisional Application No. 60 / 563,076, filed Apr. 15, 2004, each of which is incorporated by reference in its entirety.BACKGROUND [0002] 1. Field of the Invention [0003] This invention relates generally to matching test data to data within a database, and in particular to efficient fuzzy matching of data sampled from a noisy environment to samples within a large repository. [0004] 2. Background of the Invention [0005] An important class of problems involves searching through a data repository for a match to particular item of test data, where the data repository contains a large number of data segments. The repository typically contains a set of sequenced data that reflects known events or items, and the test segment is a sample acquired from an unknown event or item. The test segment is often, but not necessa...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

Application Information

Patent Timeline
no application Login to View More
Patent Type & AuthorityApplications(United States)
IPC IPC(8): G06F17/30
CPCG06F17/3002Y10S707/99943Y10S707/99948Y10S707/99945Y10S707/99936G06F16/41
InventorCARUSO, JEFFREY L.
OwnerETSY INC