Method and system for reusing data item fingerprints in generating semantic maps

The method and system generate and reuse sparse distributed representations to enhance data clustering and analysis, addressing the limitations of conventional systems by enabling semantic projection and reuse of data representations.

JP7827716B2Active Publication Date: 2026-03-10CORTICAL IO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Conventional systems lack the ability to use self-organizing maps as distributed semantic projection maps for explicit semantic definition of data items and do not allow for the reuse of previously generated data representations, limiting their application in clustering and analysis of data documents.

Method used

A method and system that generates and reuses sparse distributed representations (SDRs) by clustering data documents in a two-dimensional metric space, using a reference map generator, parser, and sparsification module to create SDRs, which are stored in a database for subsequent semantic mapping and analysis.

Benefits of technology

Enables the reuse of data representations for generating subsequent semantic maps, improving data clustering and analysis by providing a system that can identify semantic similarity and enhance search functionality across various data types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007827716000001
    Figure 0007827716000001
  • Figure 0007827716000002
    Figure 0007827716000002
  • Figure 0007827716000003
    Figure 0007827716000003
Patent Text Reader

Abstract

A method for generating clusters of distributed representations in a second two-dimensional metric space using distributed representations of data items in a first set of data documents clustered in a first two-dimensional metric space includes clustering the set of data documents in the first two-dimensional metric space and generating a semantic map using a reference map generator. A parser generates a list of data items present in the set of data documents. The representation generator generates distributed representations using presence information for each data item. A sparsification module receives identification information for a maximum level of sparsity and reduces the total number of set bits in the distributed representations. The reference map generator clusters a set of SDRs retrieved from an SDR database and selected according to at least one second criterion in the second two-dimensional metric space to generate a second semantic map.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method and system for reusing data item fingerprints in the generation of semantic maps. Summary of the Invention [Problem to be solved by the invention]

[0002] In conventional systems, the use of self-organizing maps is typically limited to clustering data documents by type and predicting areas where unidentified data documents will be clustered, or analyzing the cluster structure of a set of used data documents. Such conventional systems typically do not provide the ability to use the resulting "clustering map" as a "distributed semantic projection map" for explicit semantic definition of the constituent data items of the data documents. Furthermore, conventional systems typically use conventional processor-based computing to implement methods for using self-organizing maps. In addition, conventional systems do not provide the ability to reuse previously generated data representations; for example, the system typically makes a selection regarding which subset of millions of data documents to use in generating the semantic map, forcing a decision between granularity and practical capabilities for generating data representations of data items across millions of data documents. Therefore, there is a need for a system that can generate data representations and reuse the generated data representations when generating subsequent semantic maps. [Means for solving the problem]

[0003] In one aspect, a method for generating clusters of distributed representations in a second two-dimensional metric space using distributed representations of data items in a first set of data documents clustered in a first two-dimensional metric space includes: clustering a set of data documents selected according to at least one criterion in the two-dimensional metric space by a reference map generator executing on a computing device to generate a semantic map. The method includes associating coordinate pairs with each of the set of data documents by the semantic map. The method includes generating a list of data items present in the set of data documents by a parser executing on the computing device. The method includes determining, by a representation generator executing on the computing device, presence information for each data item in the list, including (i) the number of data documents in which the data item occurs, (ii) the number of occurrences of the data item in each data document, and (iii) coordinate pairs associated with each data document in which the data item occurs. The method includes generating, by the representation generator, a distributed representation for each data item using the presence information. The method includes receiving, by a sparsification module executing on the computing device, an identification of a maximum level of sparsity. The method includes generating sparse distributed representations (SDRs) having a reference fill grade by using a sparsification module to reduce the total number of set bits in each distributed representation based on a maximum level of sparsity. The method includes storing each of the SDRs in an SDR database. The method also includes generating a second semantic map by using a reference map generator executing on a computing device to cluster a set of SDRs retrieved from the SDR database and selected according to at least one second criterion in a second two-dimensional metric space. [Brief explanation of the drawings]

[0004] The foregoing and other objects, aspects, features, and advantages of the present specification will become more apparent and will be better understood by referring to the following description taken in conjunction with the accompanying drawings. [Figure 1A] FIG. 1 is a block diagram illustrating an embodiment of a system for mapping data items to sparse distributed representations. [Figure 1B] FIG. 1 is a block diagram illustrating one embodiment of a system for generating a semantic map for use in mapping data items to sparse distributed representations. [Figure 1C] FIG. 1 is a block diagram illustrating one embodiment of a system for generating sparse distributed representations of data items in a set of data documents. [Figure 2] 1 is a flowchart illustrating one embodiment of a method for mapping data items to a sparse distributed representation. [Figure 3] FIG. 1 is a block diagram illustrating one embodiment of a system for performing operations on a sparse distributed representation of data items generated using data documents clustered on a semantic map. [Figure 4] 1 is a flow chart illustrating one embodiment of a method for identifying levels of semantic similarity between data items. [Figure 5] 1 is a flow chart illustrating one embodiment of a method for identifying a level of semantic similarity between a user-provided data item and a data item in a set of data documents. [Figure 6A] 1 is a block diagram illustrating one embodiment of a system for query expansion provided for use with a full-text search system. [Figure 6B] 1 is a flow chart illustrating one embodiment of a method for expanding a query provided for use with a full-text search system. [Figure 6C] 1 is a flow chart illustrating one embodiment of a method for expanding a query provided for use with a full-text search system. [Figure 7A] 1 is a block diagram illustrating one embodiment of a system for providing topic-based information to a full-text search system. [Figure 7B] 1 is a flow chart illustrating one embodiment of a method for providing topic-based documents to a full-text search system. [Figure 8A] 1 is a block diagram illustrating one embodiment of a system for providing keywords associated with documents to a full-text search system for improved indexing. [Figure 8B]1 is a flow chart illustrating one embodiment of a method for providing keywords associated with a document to a full-text search system that improves indexing. [Figure 9A] 1 is a block diagram illustrating one embodiment of a system for providing search functionality for text documents. [Figure 9B] 1 is a flow chart illustrating one embodiment of a method for providing search functionality for text documents. [Figure 10A] 1 is a block diagram illustrating one embodiment of a system for matching user expertise within a full-text search system. [Figure 10B] 1 is a block diagram illustrating one embodiment of a system for matching user expertise within a full-text search system. [Figure 10C] 1 is a flow chart illustrating one embodiment of a method for matching user expertise with a request for user expertise. [Figure 10D] 1 is a flow chart illustrating one embodiment of a method for user-profile-based semantic ranking of query results received from a full-text search system. [Figure 11A] FIG. 1 is a block diagram illustrating an embodiment of a system for providing medical diagnostic support. [Figure 11B] 1 is a flow chart illustrating one embodiment of a method for providing medical diagnostic support. [Figure 12A] FIG. 1 is a block diagram illustrating an embodiment of a computer useful in connection with the methods and systems described herein. [Figure 12B] FIG. 1 is a block diagram illustrating an embodiment of a computer useful in connection with the methods and systems described herein. [Figure 12C] FIG. 1 is a block diagram illustrating an embodiment of a computer useful in connection with the methods and systems described herein. [Figure 12D] 1 is a block diagram illustrating an embodiment of a system in which multiple networks provide data hosting and distribution services. [Figure 13]FIG. 1 is a block diagram illustrating one embodiment of a system for generating cross-lingual sparse distributed representations. [Figure 14A] 1 is a flow chart illustrating one embodiment of a method for determining similarity between cross-lingual sparse distributed representations. [Figure 14B] 1 is a flow chart illustrating one embodiment of a method for determining similarity between cross-lingual sparse distributed representations. [Figure 15] 1 is a block diagram illustrating one embodiment of a system for identifying a level of similarity between filtering criteria and data items within a set of streamed documents. [Figure 16] 1 is a block diagram illustrating one embodiment of a system for identifying a level of similarity between filtering criteria and data items within a set of streamed documents. [Figure 17A] FIG. 1 is a block diagram illustrating one embodiment of a method for identifying a level of similarity between multiple binary vectors. [Figure 17B] 1 is a flow chart illustrating one embodiment of a method for identifying a level of similarity between multiple binary vectors. [Figure 18A] 1 is a block diagram illustrating one embodiment of a system for identifying a level of similarity between multiple data representations. [Figure 18B] 1 is a flow chart illustrating one embodiment of a method for identifying a level of similarity between multiple data representations. [Figure 19] FIG. 1 is a flow diagram illustrating one embodiment of a method for lazy sparsification of composite data representations for use in identifying levels of similarity between multiple data representations. [Figure 20] FIG. 1 is a flow diagram illustrating one embodiment of a method for fractal fingerprinting of data items. DETAILED DESCRIPTION OF THE INVENTION

[0005] In some embodiments, the methods and systems described herein provide functionality for identifying a level of similarity between multiple data representations. In one of these embodiments, the identification information is based on a determined distance between sparse distributed representations (SDRs) or any other type of long binary vectors.

[0006] Referring to the block diagram in Figure 1A, a system that maps data items to sparse distributed representations is shown. One embodiment of a system is disclosed. Broadly speaking, the system 100 includes an engine 101, a machine 102A, a set of data documents 104 (a collection of data documents 104), a reference map generator 106, a semantic map 108, a parser and preprocessing module 110, a list of data items 112, a representation generator 114, a sparsification module 116, one or more sparse distributed representations (SDRs) 118, a sparse distributed representation (SDR) database 120, and a full-text search system 122. In some embodiments, the engine 101 exhibits all of the elements and functionality described in connection with FIGS. 1A-1C and 2.

[0007] 1A , the system includes a set of data documents 104. In one embodiment, the documents in the set of data documents 104 include text data. In another embodiment, the documents in the set of data documents 104 include various values ​​of a physical system. In yet another embodiment, the documents in the set of data documents 104 include patient medical records. In another embodiment, the documents in the set of data documents 104 include chemistry-based information (e.g., DNA sequences, protein sequences, chemical formulas). In yet another embodiment, each document in the set of data documents 104 includes musical scores. Data items in the data documents 104 may be words, numbers, medical analyses, medical measurements, and musical notes. The data items may be a series of any type (e.g., a string containing one or more numbers). The data items in the first set of data documents 104 may be in a language different from the language of the data items in the second set of data documents 104. In some embodiments, the set of data documents 104 includes historical log data. As used herein, a "document" may refer to a collection of data items, each of which corresponds to a system variable that exists from the same system. In some embodiments, the system variables of such documents are extracted simultaneously.

[0008] As mentioned above, the use of the term "data item" herein includes words as string data, scalar values ​​as numeric data, medical diagnoses and analyses as numeric or categorical data, musical notes, and any type of variable that all exist from the same "system." A "system" can be any physical system, natural or artificial, such as a river, a technological device, or a biological entity such as a living cell or a human organism. The system can also be a "concept system" such as a language or web server log data. The language can be a natural language such as English or Chinese, or an artificial language such as JAVA or C++ program code. As mentioned above, the use of the term "data document" includes a set of "data items." These data items can be interdependent depending on the meaning of the underlying "system." This grouping can be a time-based grouping if all data item values ​​are sampled simultaneously; for example, measurement data items coming from a car engine can be sampled every second and grouped into a single data document. The above classification can be made along the logical structure characterized by the "system" itself. For example, in the case of natural language, word data items can be classified as sentences, and in the case of music, data items corresponding to musical notes can be classified by meter. Based on these data documents, document vectors can be generated according to the methods described above (or according to other methods understood by those skilled in the art) to generate a semantic map for the "system," as described in more detail below. This "system" can be used to generate semantic map data item SDRs, as described in more detail below. All of the methods and systems described below may be used for all types of data item SDRs.

[0009] In one embodiment, a user selects the set of data documents 104 according to at least one criterion. For example, a user may select data documents to add to the set of data documents 104 based on whether the data documents relate to a particular subject matter. As another example, the set of data documents 104 represents a semantic set for which the system 100 is utilized. In one embodiment, the user is a human user of the system 100. In another embodiment, the machine 100 performs functions to select the data documents in the set of data documents 104.

[0010] The system 100 includes a reference map generator 106. In one embodiment, the reference map generator 106 is a self-organizing map. In another embodiment, the reference map generator 106 is a generated topological map. In yet another embodiment, the reference map generator 106 is an elastic map. In another embodiment, the reference map generator 106 is a neural gas map. In yet another embodiment, the reference map generator 106 is any type of competitive, learning-based, unsupervised, or dimensionality-reducing machine learning method. In another embodiment, the reference map generator 106 is any computational method capable of receiving a set of data documents 104 and generating from the set of data documents 104 a two-dimensional metric space in which clustered points representing documents reside. In yet another embodiment, the reference map generator 106 is any computer program capable of accessing a set of data documents 104 and generating from the set of data documents 104 a two-dimensional metric space in which clustered points representing documents reside. Although typically described herein as forming a two-dimensional metric space, in some embodiments, the reference map generator 106 forms an n-dimensional metric space. In some embodiments, the reference map generator 106 is implemented in software. In other embodiments, the reference map generator 106 is implemented in hardware.

[0011] The two-dimensional distance space may be referred to as a semantic map 108. The semantic map 108 may be any vector space with an associated distance measure.

[0012] In one embodiment, the parser and preprocessing module 110 generates the list of data items 112. In other embodiments, the parser and preprocessing module 110 forms part of the representation generator 114. In some embodiments, the parser and preprocessing module 110 is implemented as at least part of a software program. In other embodiments, the parser and preprocessing module 110 is implemented as at least part of a hardware module. In still other embodiments, the parser and preprocessing module 110 is implemented on the machine 102. In some embodiments, the parser and preprocessing module 110 may be specialized to suit the type of data. In other embodiments, multiple parser and preprocessing modules 110 are provided to suit the type of data.

[0013] In one embodiment, representation generator 114 generates distributed representations of data items. In some embodiments, representation generator 114 is implemented as at least part of a software program. In other embodiments, representation generator 114 is implemented as at least part of a hardware module. In yet other embodiments, representation generator 114 is implemented on machine 102.

[0014] In one embodiment, the sparsification module 116 generates a sparse distributed representation (SDR) of the data item. As one skilled in the art would understand, the SDR may be a large (binary) vector. For example, the SDR may have thousands of elements. In some embodiments, each element of the SDR generated by the sparsification module 116 has a particular semantic meaning. In one of these embodiments, vector elements with similar semantic meanings are closer to each other than vector elements that are semantically dissimilar, as measured by an associated distance metric.

[0015] In one embodiment, representation generator 114 provides the functionality of sparsification module 116. In other embodiments, representation generator 114 communicates with another sparsification module 116. In some embodiments, sparsification module 116 is implemented as at least part of a software program. In other embodiments, sparsification module 116 is implemented as at least part of a hardware module. In yet other embodiments, sparsification module 116 is implemented on machine 102.

[0016] In one embodiment, the sparse distributed representation (SDR) database 120 stores the sparse distributed representations 118 generated by the representation generator 114. In another embodiment, the sparse distributed representation database 120 stores the SDRs and the data items that the SDRs represent. In yet another embodiment, the SDR database 120 stores metadata associated with the SDRs. In another embodiment, the SDR database 120 has indices that identify the SDRs 118. In yet another embodiment, the SDR database 120 has indices that identify data items that are semantically close to a particular SDR 118. In one embodiment, the SDR database 120 may store, by way of example and not limitation, one or more of the following: a reference number for a data item; the data item itself; an indication of a data item frequency for the data item in the set of data documents 104; an abbreviated version of the data item; a compressed binary representation of the SDR 118 for the data item; one or more tags for the data item; an indication of whether the data item identifies a location (e.g., Vienna); or an indication of whether the data item identifies a person (e.g., Einstein). In other embodiments, SDR database 120 may be any type of database or any format of database.

[0017] Examples of SDR database 120 include, but are not limited to, structured storage (e.g., NoSQL-type databases, BigTable databases), HBase databases (available from The Apache Software Foundation (Forest Hill, Maryland)), MongoDB databases (available from 10Gen, Inc. (New York, New York)), Cassandra databases (available from The Apache Software Foundation), and document-based databases. In other embodiments, SDR database 120 is an ODBC-compliant database. For example, SDR database 120 may be provided as an ORACLE database manufactured by Oracle Corporation (Redwood City, California). In other embodiments, SDR database 120 may be a Microsoft ACCESS database or a Microsoft SQL Server database manufactured by Microsoft Corporation (Redmond, Washington). In yet other embodiments, SDR database 120 may be a custom-designed database based on an open-source database, such as the MYSQL family of freely available database products distributed by Oracle Corporation.

[0018] Referring now to the flowchart of FIG. 2 , one embodiment of a method 200 for mapping data items to sparse distributed representations is disclosed. In summary, the method 200 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional metric space by a reference map generator running on a computing device to generate a semantic map (202). The method 200 includes associating coordinate pairs with each of the set of data documents through the semantic map (204). The method 200 includes generating a list of data items present in the set of data documents by a parser running on the computing device (206). The method 200 includes determining, by a representation generator running on the computing device, presence information for each data item in the list, including: (i) the number of data documents in which the data item occurs, (ii) the number of occurrences of the data item in each data document, and (iii) coordinate pairs associated with each data document in which the data item occurs (208). The method 200 includes generating, by the representation generator, a distributed representation using the presence information (210). The method 200 includes receiving, by a sparsification module executing on a computing device, an identification of a maximum level of sparsity (212). The method 200 also includes, by the sparsification module, reducing the total number of set bits in the distributed representation based on the maximum level of sparsity to generate a sparse distributed representation having a reference fillgrade (214).

[0019] Referring now to FIG. 2 in more detail in conjunction with FIGS. 1A-1B, method 200 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional distance space by a reference map generator executed on a computing device to generate a semantic map (202). In one embodiment, the at least one criterion indicates that a data item in the set of data documents 104 occurs a threshold number of times. In another embodiment, the at least one criterion indicates that each data document in the set of data documents 104 should contain descriptive information related to the state of the system from which the data document originates. In the case of data documents, the at least one criterion indicates that each data document should represent a conceptual topic (e.g., an encyclopedic description). In another embodiment, the at least one criterion indicates that the feature list of the set of data documents 104 should uniformly fill a desired information space. In another embodiment, the at least one criterion indicates that the set of data documents 104 originate from the same system. In the case of data documents, the at least one criterion indicates that the data documents are all in the same language. In yet another embodiment, the at least one criterion indicates that the set of data documents 104 is in a natural (e.g., human) language. In yet another embodiment, the at least one criterion indicates that the set of data documents 104 is in a computer language (e.g., any type of computer code). In another embodiment, the at least one criterion indicates that the set of data documents 104 may include any type or form of professional or industry terms (e.g., medical, legal, scientific, automotive, military, etc.). In another embodiment, the at least one criterion indicates that the set of data documents 104 should have a maximum threshold number of documents in the set. In some embodiments, a human user selects the set of data documents 104, and the machine 120 receives the selected set of data documents 104 from the human user (e.g., via a user interface to a repository, directory, document database, or other data structure (not shown) that stores one or more data documents).

[0020] In one embodiment, the machine 102 preprocesses the set of data documents 104. In some embodiments, a parser and preprocessing module 110 provides the preprocessing functionality of the machine 102. In other embodiments, the machine 102 segments each of the set of data documents 104 into words or sentences, aligns punctuation, and removes or converts undesirable characters. In yet other embodiments, the machine 102 executes a tagging module (not shown) to associate one or more meta-information tags with any data items or portions of data items in the set of data documents 104. In other embodiments, the machine 102 standardizes the text size of basic conceptual units and divides the set of data documents 104 into equal-sized text amounts. In this embodiment, the machine 102 may apply one or more constraints when dividing the set of data documents 104 into snippets. For example, without limitation, the constraints may indicate that documents in the set of data documents 104 should contain only complete sentences, should contain a certain number of sentences, should have a limited number of data items, should have a minimum number of different nouns per document, the segmentation process should respect the verbatim text of the document's author, etc. In one embodiment, the application of the constraints is optional.

[0021] In some embodiments, to generate more useful document vectors, the system 100 provides functionality for identifying the most relevant data items for each document in the set of data documents 104 from a semantic perspective. In one of these embodiments, the parser and preprocessing module 110 provides this functionality. In other embodiments, the reference map generator 106 receives one or more document vectors and generates the semantic map 108 using the received document vectors. For example, the system 100 may be configured to identify and select nouns (e.g., based on part-of-speech tags assigned to each data item in the document during preprocessing). As another example, the system 100 may stem the selected nouns to aggregate all morphological variations (e.g., plural variations and case variations) present in a given main data item instance. As yet another example, a term frequency-inverse document frequency ("TF-IDF") statistic may be calculated for the selected nouns to reflect the importance of a data item to a given data document given a particular set of data documents 104. A coefficient may be calculated based on the number of data items in the document and the number of data items in the set of data documents 104. In some embodiments, the system 100 identifies a predetermined number of highest tf-idf index and stemmed nouns per document and generates a complete list of selected nouns to define document vectors (e.g., vectors indicating whether a particular data item is present in a document, as will be understood by those skilled in the art) that are used to train the semantic map 106. In other embodiments, the functionality for preprocessing and vectorizing the set of data documents 104 generates a vector for each document in the set of data documents 104. In one of these embodiments, an identifier and integer for each data item in the list of selected nouns represents each document.

[0022] In one embodiment, machine 102 provides the preprocessed documents to full-text search system 122. For example, parsing and preprocessing module 110 may provide this functionality. In other embodiments, full-text search system 122 may be utilized to enable cross-selection of documents. For example, full-text search system 122 may provide functionality that allows for the search of all documents containing a particular data item, or for the search of snippets of original documents that contain a particular data item, for example, using an exact character match. In yet other embodiments, each preprocessed document (or preprocessed document fragment) is associated with at least one of a document identifier, a fragment identifier, a document title, the document text, the number of data items in the document, the document byte length, and a classification identifier. In other embodiments, semantic map coordinate pairs may be assigned to the document, and such coordinate pairs may be associated with the preprocessed document in full-text search system 122, as described in more detail below. In such embodiments, full-text search system 122 may provide functionality that receives single and compound data items and returns coordinate pairs for all matching documents that contain the received data item. Full-text search systems 122 include, but are not limited to, Lucene-based systems (e.g., Apache SOLR, distributed by The Apache Software Foundation, Forest Hill, Maryland, and ELASTICSEARCH, distributed by Elasticsearch Global BV, Amsterdam, The Netherlands), open source systems (e.g., Slashdot Media, San Francisco, California, owned and operated by Dice Holdings, Inc. company, New York, New York, Indri, distributed by The Lemur Project through SourceForge, and MNOGOSEARCH, distributed by Lavtech.Com Corp.).Sphinx, distributed by Sphinx Technologies Inc.; Xapian, distributed by the Xapian Project; Swish-e, distributed by Swish-e.org; BaseX, distributed by BaseX GmbH (Konstanz, Germany); DataparkSearch Engine, distributed by www.dataparksearch.org; ApexKB, distributed by SourceForge, owned and operated by Slashdot Media; Searchdaimon, distributed by Searchdaimon AS (Oslo, Norway). Zettair, sold by RMIT University (Melbourne, Australia), and Commercial Systems (Autonomy IDOL, manufactured by Hewlett-Packard (Sunnyvale, California); COGITO product line, manufactured by Expert Systems SpA (Modena, Italy); Fast Search & Transfer, manufactured by Microsoft, Inc. (Redmond, Washington); ATTIVIO, manufactured by Attivio, Inc. (Newton, Massachusetts); BRS / Searc, manufactured by OpenText Corporation (Waterloo, Ontario, Canada); Perceptive Intelligent Capture (powered by Brainware), manufactured by Perceptive Software by Lexmark (Shawnee, Kansas); any product manufactured by Concept Searching, Inc. (McLean, Virginia); COVEO, manufactured by CoveoSolutions, Inc. (San Mateo, California); Dieselpoint SEARCH, manufactured by Dieselpoint, Inc. (Chicago, Illinois); dtSearch DTSEARCH is manufactured by DTSEARCH Company (Bethesda, Maryland).Oracle Endeca Information Discovery manufactured by Oracle Corporation (Redwood Shores, California). Products manufactured by Exalead (a subsidiary of Dassault Systemes (Paris, France)). Inktomi search engine provided by Yahoo!. ISYS Search (now Perceptive Enterprise Search) manufactured by Perceptive Software by Lexmark (Shawnee, Kansas). Locayta (now ATTRAQT FREESTYLE MERCHANDISING) manufactured by ATTRAQT, LTD. (London, England, UK). Lucid Imagination (now LUCIDWORKS) manufactured by LucidWorks (Redwood City, California). MARKLOGIC manufactured by MarkLogic Corporation (San Carlos, California). The Mindbreeze product line manufactured by Mindbreeze GmbH (Linz, Austria). Omniture (now Adobe SiteCatalyst) manufactured by Adobe Systems, Inc. (San Jose, California). The OpenText product line is manufactured by OpenText Corporation (Waterloo, Ontario, Canada). The PolySpot product line is manufactured by PolySpot SA (Paris, France). Thunder. Manufactured by Thunderstone Software LLC (Cleveland, Ohio) Product line: Vivisimo (now IBM Watson Explorer) manufactured by IBM Corporation (Armonk, NY). Full-text search systems may also be referred to herein as enterprise search systems.

[0023] In one embodiment, the reference map generator 106 accesses the document vectors of the set of data documents 104 and scatters each document across a two-dimensional metric space. In another embodiment, the reference map generator 106 accesses the preprocessed set of data documents 104 and scatters points representing each document across a two-dimensional metric space. In yet another embodiment, the scattered points are clustered. For example, the reference map generator 106 may calculate the locations of points representing documents based on the semantic content of the documents. The resulting scatters represent the semantic collection of the particular set of data documents 104.

[0024] In one embodiment, the reference map generator 106 is trained using document vectors from a preprocessed set of data documents 104. In another embodiment, the reference map generator 106 is trained using document vectors from a (e.g., non-preprocessed) set of data documents 104. A user of the system 100 may train the reference map generator 106 with the set of data documents 104 using training methods well understood by those skilled in the relevant art.

[0025] In one embodiment, the above training process leads to two results. First, for each document in the set of data documents 104, a coordinate pair that locates the document on the semantic map 108 is identified. These coordinates may be stored in each document entry in the full-text search system 122. Second, a weight map is generated that allows the reference map generator 106 to locate any new (unseen) document vectors on the semantic map 108. After training the reference map generator 106, the distribution of documents may be static. However, if the initial training set is sufficiently large and descriptive, the vocabulary can be expanded by adding new training documents. New documents may be located on the map to avoid time-consuming recalculation of the semantic map. This is done by transforming the document vectors of the new documents by the trained weights. The target semantic map 108 can be refined and improved by analyzing the distribution of points representing documents across the semantic map 108. If there are under- or over-represented topics, the set of data documents 104 can be adapted accordingly, and then the semantic map 108 can be recalculated.

[0026] Thus, the method 200 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional distance space by a reference map generator running on a computing device to generate a semantic map. As described above and as will be appreciated by those skilled in the art, various techniques may be used to cluster the data documents. For example, but not limited to, generative topological maps, growing self-organizing maps, elastic maps, neural gas, random mapping, latent semantic indexing, principal component analysis, or any other dimension reduction-based mapping method may be used.

[0027] Referring now to the block diagram of FIG. 1B , one embodiment of a system for generating a semantic map 108 for use in mapping data items to a sparse distributed representation is disclosed. As shown in FIG. 1B , a set of data documents 104 received by a machine 102 may be referred to as a language definition corpus. After preprocessing the set of data documents, the documents may be referred to as a reference map generator training corpus. The documents may also be referred to as a neural network training corpus. A reference map generator 106 accesses the reference map generator training corpus to generate as output a semantic map 108 in which the set of data documents are located. The semantic map 108 may extract coordinates for each document. The semantic map 108 may provide the coordinates to a full-text search system 122. As a non-limiting example, the corpus may include a corpus based on an application (e.g., a web application for content creation and management) that allows collaborative modification, expansion, or deletion of its content and structure. Such applications may be referred to, by way of example, as a "wiki" and the "Wikipedia" encyclopedia project provided and implemented by the Wikimedia Foundation (San Francisco, California). A corpus may include any kind or type of knowledge base.

[0028] As will be appreciated by those skilled in the art, any type or form of algorithm may be used to map high-dimensional vectors to a low-dimensional space (e.g., semantic map 108). For example, input vectors may be clustered such that similar vectors are located close to each other in the low-dimensional space, resulting in a topologically clustered low-dimensional map. In some embodiments, the size of the quadratic semantic map (map) defines a "semantic derivation," which computes a pattern of sparse distributed representations (SDRs) of data items, as described in more detail below. For example, a side length of 128 corresponds to a descriptive capacity of 16K features per data item SDR. In principle, the size of the map can be chosen freely, taking into account computational limitations, such as the longer it takes to train a reference map generator and the longer it takes to compare or process by any means. As another example, a 128x128 SDR of data items has been found to be useful when applied to a set of "general English" data documents 104.

[0029] Referring again to FIG. 2 , method 200 includes associating 204 coordinate pairs with each of the set of data documents through a semantic map. As described above, when generating semantic map 108, reference map generator 106 calculates the location of a point on semantic map 108, where the point represents a document in set of data documents 104. Semantic map 108 may then extract coordinates for the point. In some embodiments, semantic map 108 sends the extracted coordinates to full-text search system 122.

[0030] 1C, one embodiment of a system for generating sparse distributed representations for each of a plurality of data items in a set of data documents 104 is disclosed. As shown in FIG. 1C, a representation generator 114 sends a query to a full-text search system 122 and receives one or more data items that match the query. The representation generator 114 may generate sparse distributed representations of the retrieved data items from the full-text search system 122. In some embodiments, generating an SDR using data from the semantic map 108 may be said to include "folding" semantic information into the generated sparse distributed representations (e.g., sparsely populated vectors).

[0031] Referring again to FIG. 2 , method 200 includes generating 206 a list of data items present in the set of data documents by a parser executing on a computing device. In one embodiment, parser and preprocessing module 110 generates list of data items 112. In another embodiment, parser and preprocessing module 110 directly accesses the set of data documents 104 to generate list of data items 112. In yet another embodiment, parser and preprocessing module 110 accesses a full-text search system 122 that stores the set of data documents 104 in a preprocessed form (as described above). In another embodiment, parser and preprocessing module 110 expands list of data items 112 to include commonly useful combinations of data items, rather than just data items specifically contained in the set of data documents 104. For example, parser and preprocessing module 110 may access common combinations of data items (such as "bad weather" or "e-commerce") searched from publicly available collections.

[0032] In one embodiment, the parser and preprocessing module 110 defines the scope of the data item list 112, for example, using spaces and punctuation. In another embodiment, data items that occur multiple times in the list 112 under different part-of-speech tags are treated differently (e.g., the data item "fish" has a different SDR when used as a noun than when used as a verb, and therefore includes two entries). In another embodiment, the parser and preprocessing module 110 provides the data item list 112 to the SDR database 120. In yet another embodiment, the representation generator 114 accesses the stored data item list 112 to generate an SDR for each data item in the data item list 112.

[0033] Method 200 includes determining (208) presence information (including (i) through (iii) below) for each data item in the list by a representation generator executing on a computing device: (i) the number of data documents in which the data item occurs, (ii) the number of times the data item occurs in each data document, and (iii) coordinate pairs associated with each data document in which the data item occurs. In one embodiment, representation generator 114 accesses full-text search system 122 to search data stored in full-text search system 122 by semantic map 108 and parser and preprocessing module 110, and generates sparse distributed representations for the data items listed by parser and preprocessing module 110 using data from semantic map 108.

[0034] In one embodiment, the representation generator 114 accesses the full-text search system 122 to retrieve coordinate pairs for each document that contains a particular string (e.g., letters, numbers, or a combination of letters and numbers). The representation generator 114 counts the number of coordinate pairs retrieved to determine the number of documents in which the data item occurs. In another embodiment, the representation generator 114 retrieves a vector from the full-text search system 122 that represents each document that contains the string. In such an embodiment, the representation generator 114 determines the number of set bits in the vector (e.g., sets the number of bits in the vector to 1). The number of set bits indicates the number of times the data item occurs in a particular document. The representation generator 114 may add the number of set bits to determine an occurrence value.

[0035] The method 200 includes generating, by a representation generator, distributed representations using the presence information (210). The representation generator 114 may utilize known methods for generating distributed representations. In some embodiments, the distributed representations may be used to determine patterns that represent semantic contexts in which data items exist in a set of data documents 104. The spatial distribution of coordinate pairs in the patterns reflects the semantic domains of the contexts in which the data items exist. The representation generator 114 may generate a two-way mapping between the data items and their distributed representations. The SDR database 120 may be referred to as a pattern dictionary. The system 100 can use the pattern dictionary to identify data items based on the distributed representations, or vice versa. Those skilled in the art will appreciate that the system 100 can generate different pattern dictionaries by using different sets of data documents 104 (e.g., selecting documents on different types of subjects in different languages ​​based on different constraints), or by deriving them from different physical systems, different medical analysis methods, or different musical styles.

[0036] The method 200 includes receiving 212, by a sparsification module executing on a computing device, an identification of a maximum level of sparsity. In one embodiment, a human user provides the identification of the maximum level of sparsity. In other embodiments, the maximum level of sparsity is set to a predetermined threshold. In some embodiments, the maximum level of sparsity depends on the resolution of the semantic map 108. In other embodiments, the maximum level of sparsity depends on the type of reference map generator 106.

[0037] The method 200 includes generating a sparse distributed representation having a reference fill grade by reducing the total number of set bits in the distributed representation based on the maximum level of sparsity (214) using the sparsification module. In one embodiment, the sparsification module 116 sparsifies the distributed representation by setting a count threshold (e.g., by utilizing the identification information of the received maximum level of sparsity) that leads to a particular fill grade of the final SDR 118. The sparsification module 116 then generates the SDR 118, which may be said to provide a binary fingerprint of the semantic meaning of the data items in the set of data documents 104. Alternatively, the sparsification module 116 may be said to provide a binary fingerprint of the semantic values ​​of the data items in the set of data documents 104. The SDR 118 may also be referred to as a semantic fingerprint. The sparsification module 116 stores the SDR 118 in the SDR database 120.

[0038] In generating an SDR, system 100 populates a vector with ones and zeros. For example, system 100 populates a vector with ones if the data document uses the data item and zeros if the data document does not use the data item. A user may receive a graphical representation of the SDR showing points on a map that reflect the semantic meaning of the data items (the graphical representation may also be referred to as an SDR, a semantic fingerprint, or a pattern). The description herein may also refer to points and patterns. However, one skilled in the art will understand that references to "points" or "patterns" also refer to set bits that are set in an SDR vector—and to any data structure underlying such a graphical representation.

[0039] In some embodiments, the representation generator 114 and the sparsification module 116 may combine multiple data items into a single SDR. For example, if a phrase, sentence, paragraph, or other combination of data items needs to be converted into a single SDR that reflects the “union properties” of the individual SDRs, the system 100 may convert each individual data item into its SDR (either by dynamically generating it or by searching for a pre-generated SDR) and form a single composite SDR from the individual SDRs using a binary OR operation. Continuing with the above example, the number of set bits is added per position in the composite SDR. In one embodiment, the sparsification module 116 may proportionally reduce the total number of set bits using a threshold that results in a baseline fill grade. In other embodiments, the sparsification module 116 may utilize a weighting scheme to reduce the total number of set bits. Such a local weighting scheme emphasizes bits that are part of a set within the SDR. Therefore, a bit that is part of a set in an SDR is semantically more significant than a single isolated bit (e.g., with no surrounding set bits).

[0040] In some embodiments, the methods and systems described herein provide a system that not only generates a map that clusters a collection of data documents by context, but also continuously analyzes locations on the map that represent the clustered data documents, determines which data documents contain particular data items based on the analysis, and uses the analysis to provide detailed descriptions of each data item in each data document. A sparse distributed representation of the data items is generated based on data retrieved from the semantic map 108. The sparse distributed representation of the data items need not be limited to use in training other machine learning methods, but may also be used to determine relationships between data items (e.g., determining similarities between data items, ranking data items, or identifying data items that a user does not previously know to be similar for use in search and analysis in various environments). In some embodiments, by altering any portion of the information in the SDR using the methods and systems described herein, any data item becomes "semantically valid" (e.g., in its semantic collection), making it unambiguously comparable and computable without the use of machine learning, neural networks, or cortical algorithms.

[0041] In some embodiments, the generated SDRs may be used to generate additional semantic maps. For example, in embodiments where an initial semantic map is trained on a first corpus of data documents with a wide range, the SDRs may be used to avoid the need to generate new document vectors (and sparsify them) in a second, more technical corpus of data documents that includes data items that also appear in the first corpus. For example, without limitation, if the first corpus of data documents is a dictionary, encyclopedia, Wikipedia, or other corpus of general knowledge documents, and the second corpus includes a more technical set of documents (such as, without limitation, a set of medical, legal, scientific, or other specialized documents), it is likely that at least a subset of the data items in the second corpus appear in the first corpus. Using previously generated SDRs of data items common to both the first and second corpora may improve the speed and efficiency of the system, as the previously generated SDRs may be reused in the context of the second corpus. By extracting snippets from the second corpus, identifying which of them are already associated with SDRs in the SDR database, and relying on those SDRs associated with the second corpus, the system can provide enhanced functionality. For example, if the second corpus contains millions of data documents, reusing any previously generated SDRs reduces the need to regenerate those SDRs and improves the system's speed in dealing with the remaining millions of data documents. In situations where the system would otherwise have to select which subset of millions of data documents to use in generating the semantic map, a system that can reuse separately generated SDRs and focus new SDR generation in the second corpus on less common data items (e.g., generating SDRs for the more technical term "toxoplasmosis" instead of "cat") can provide improvements in both granularity and efficiency.Thus, as shown in Figure 20, in some embodiments, the methods and systems described herein may be used to generate a second semantic map based on use of a previously generated SDR. In some embodiments, a first-level semantic space may be used to generate a second-level semantic space, and the system may provide access to the second-level semantic space without providing access to the first-level semantic space, thereby providing the ability to generate multiple subsequent semantic maps from an initial corpus while maintaining privacy and / or restricting semantic map access to the initial corpus.

[0042] In some embodiments, generating a second semantic map based, at least in part, on the use of previously generated SDRs can enable the semantic map to be generated based on a collection of fewer references than would otherwise be required; as an example, the system may require only 10% of the data that would otherwise be required to generate the second semantic map. Furthermore, the system may provide functionality for identifying associations between subsequently generated semantic maps. For example, if a first semantic map is generated from a corpus of generic data items, a second semantic map is generated for a second corpus containing less commonly occurring data items (e.g., "toxoplasmosis"), and a third semantic map is generated for the third corpus (e.g., for illustrative purposes only, including data items with terms such as "hepatitis"), the system may identify commonalities between the second and third semantic maps, and, following the above example, identify correlations between the data items "hepatitis" and "toxoplasmosis." As a further example, if the first corpus is a large research collection, new topics may be identified within the collection without losing either the resolution of the second semantic map or the context of the first semantic map.

[0043] Referring to Figure 20, a flow diagram illustrates one embodiment of a method for generating clusters of distributed representations in a second two-dimensional metric space using distributed representations of data items in a first set of data documents clustered in a first two-dimensional metric space. Method 2000 includes clustering (2002) a set of data documents selected according to at least one criterion in the two-dimensional metric space with a reference map generator executing on a computing device to generate a semantic map. In one embodiment, the clustering is performed as described above in connection with Figure 2 (202). Method 2000 also includes associating (2004) coordinate pairs with each of the set of data documents with the semantic map. In one embodiment, the association is performed as described above in connection with Figure 2 (204). Method 2000 also includes generating (2006) a list of data items present in the set of data documents with a parser executing on the computing device. In one embodiment, the generation is performed as described above in connection with Figure 2 (206). The method 2000 includes determining (2008), by a representation generator executing on a computing device, presence information for each data item in the list, including (i) a number of data documents in which the data item occurs, (ii) a number of occurrences of the data item in each data document, and (iii) a coordinate pair associated with each data document in which the data item occurs. In one embodiment, the determining is performed as described above in connection with FIG. 2 (208). The method 2000 includes (2010), by the representation generator, generating a distributed representation for each data item using the presence information. In one embodiment, the generating is performed as described above in connection with FIG. 2 (210). The method 2000 includes (2012), by a sparsification module executing on the computing device, receiving an identification of a maximum level of sparsity. In one embodiment, the receiving is performed as described above in connection with FIG. 2 (212). The method 2000 includes (2014), by the sparsification module, reducing the total number of set bits in each distributed representation based on the maximum level of sparsity to generate sparse distributed representations (SDRs) having a benchmark fill grade. In one embodiment, the reduction is performed as described above in connection with Figure 2 (214).The method 2000 includes storing (2016) each of the SDRs in an SDR database.

[0044] The method 2000 includes clustering a set of SDRs retrieved from the SDR database and selected according to at least one second criterion in a second two-dimensional distance space by a reference map generator running on a computing device, executing the set of SDRs selected according to at least one second criterion, to generate a second semantic map (2018). In one embodiment, the set of SDRSs is selected based on receiving an indication from a full-text search system that an SDR is associated with a second set of data documents. In another embodiment, the method may include providing at least one snippet of at least one data document in the second set of data documents to the full-text search system, receiving from the full-text search system a list of coordinate pairs of matching data documents in the set of data documents containing the provided snippet, and retrieving at least one SDR associated with each of the coordinate pairs in the list of coordinate pairs from the SDR database. Upon generating the second semantic map initially populated with the retrieved SDRs, the system may generate additional SDRs for additional terms and add them to the second semantic map.

[0045] Referring now to the block diagram of FIG. 3 , one embodiment of a system for performing operations using sparse distributed representations of data items in data documents clustered on a semantic map is disclosed. In one embodiment, the system 300 includes functionality for determining semantic similarity between the sparse distributed representations. In another embodiment, the system 300 includes functionality for determining a relevance ranking of a data item converted to SDR by matching it with reference data items converted to SDR. In yet another embodiment, the system 300 includes functionality for determining a classification of a data item converted to SDR by matching it with reference text elements converted to SDR. In another embodiment, the system 300 includes functionality for performing topic filtering of the data item converted to SDR by matching it with reference data items converted to SDR. In yet another embodiment, the system 300 includes functionality for performing keyword extraction from the data item converted to SDR.

[0046] 1A-1C (engine 101 and SDR database 120 in FIG. 3) and provides the functionality described above in connection with those figures. System 300 also includes machine 102A, machine 102b, fingerprint module 302, similarity engine 304, disambiguation module 306, data item module 308, and expression engine 310. In one embodiment, engine 101 executes on machine 101A. In one embodiment, fingerprint module 302, similarity engine 304, disambiguation module 306, data item module 308, and expression engine 310 execute on machine 102b.

[0047] 3 in conjunction with FIGS. 1A-1C and 2, a system 300 includes a fingerprinting module 302. In one embodiment, fingerprinting Module 302 includes representation generator 114 and sparsification module 116, described above in connection with FIGS. 1A-1C and 2. In other embodiments, fingerprint module 302 forms part of engine 101. In one embodiment, fingerprint module 302 is implemented as at least part of a hardware program. In other embodiments, fingerprint module 302 is implemented as at least part of a software program. In yet other embodiments, fingerprint module 302 executes on machine 102. In some embodiments, fingerprint module 302 performs a post-production process that converts data item SDRs into semantic fingerprints in real time (e.g., via a sparsification process described herein) using SDRs that are not part of SDR database 120 but are dynamically generated (e.g., to create document semantic fingerprints from word semantic fingerprints), although such a post-production process is optional. In other embodiments, representation generator 114 may be accessed directly to generate sparsified SDRs for data items that do not have an SDR in SDR database 120. In such an embodiment, the representation generator 114 may automatically invoke the sparsification module 116 to automatically generate the sparsified SDR. The terms “SDR,” “fingerprint,” and “semantic fingerprint” are used interchangeably herein to refer to both SDRs generated by the fingerprint module 302 and SDRs generated by directly invoking the representation generator 114.

[0048] The system 300 includes a similarity engine 304. The similarity engine 304 provides functionality for calculating distances between SDRs and determining similarity levels. In other embodiments, the similarity engine 304 is implemented as at least part of a hardware module. In other embodiments, the similarity engine 304 is implemented as at least part of a software program. In yet other embodiments, the similarity engine 304 executes on the machine 102b.

[0049] The system 300 includes a disambiguation module 306. In one embodiment, the disambiguation module 306 identifies context subspaces that are incorporated into a single SDR of a data item. Thus, the disambiguation module 306 makes it easier for a user to understand different semantic contexts of a single data item. In some embodiments, the disambiguation module 306 is implemented as at least part of a hardware module. In some embodiments, the disambiguation module 306 is implemented as at least part of a software program. In other embodiments, the disambiguation module 306 executes on the machine 102b.

[0050] System 300 includes a data item module 308. In one embodiment, data item module 308 provides functionality for identifying the most distinctive data items from a set of received data items, i.e., data items whose SDR is less than a threshold distance from the SDR of the set of received data items, as described in more detail below. Data item module 308 is used in conjunction with or in place of keyword extraction module 802, described below in connection with FIG. 8A . In some embodiments, data item module 308 is implemented as at least part of a hardware module. In some embodiments, data item module 308 is implemented as at least part of a software program. In an embodiment, data item module 308 is implemented on machine 102b.

[0051] System 300 includes an expression engine 310. In one embodiment, as described in more detail below, expression engine 310 provides functionality for evaluating Boolean operators received from a user along with one or more data items. Evaluating Boolean operators provides flexibility to users in requesting analysis on one or more data items or combinations of data items. In some embodiments, expression engine 310 is implemented as at least part of a hardware module. In some embodiments, expression engine 310 is implemented as at least part of a software program. In one embodiment, expression engine 310 is implemented on machine 102b.

[0052] Referring to the flowchart of FIG. 4, one embodiment of a method for identifying similarity levels between data items is disclosed. In summary, method 400 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional metric space by a reference map generator running on a computing device to generate a semantic map (402). Method 400 also includes associating coordinate pairs with each of the set of data documents through the semantic map (404). Method 400 also includes generating a list of data items present in the set of data documents by a parser running on the computing device (406). Method 400 also includes determining, by a representation generator running on the computing device, presence information for each data item in the list, including: (i) the number of data documents in which the data item occurs, (ii) the number of occurrences of the data item in each data document, and (iii) coordinate pairs associated with each data document in which the data item occurs (408). Method 400 also includes generating, by the representation generator, a distributed representation using the presence information (410). The method 400 includes receiving, by a sparsification module executing on a computing device, an identification of a maximum level of sparsity (412). The method 400 includes generating, by the sparsification module, a sparse distributed representation (SDR) having a reference fill grade (fillgrAde) by reducing the total number of set bits in the distributed representation based on the maximum level of sparsity (414). The method 400 includes determining, by a similarity engine executing on the computing device, a distance between a first SDR of a first data item and a second SDR of a second data item (416). The method 400 includes providing, by the similarity engine, an identification of a semantic similarity level between the first data item and the second data item based on the determined distance (418).

[0053] 4 in more detail in conjunction with FIGS. 1A-1C and 2-3, a method 400 includes clustering, by a reference map presenter executing on a computing device, a set of selected data documents in a two-dimensional metric space according to at least one criterion to generate a semantic map (402). In one embodiment, the clustering is performed (202) as described above in conjunction with FIG.

[0054] The method 400 includes associating the coordinate pairs with each of the set of data documents with a semantic map 404. In one embodiment, the associating step is performed 204 as described above in connection with FIG.

[0055] The method 400 includes generating, by a parser executing on a computing device, a list of data items present in the set of data documents 406. In one embodiment, the generating step is performed as described above in connection with FIG.

[0056] Method 400 includes determining (408) for each data item in the list, by a representation generator executing on a computing device, presence information including (i) the number of data documents in which the data item occurs, (ii) the number of times the data item occurs in each data document, and (iii) the coordinate pairs associated with each data document in which the data item occurs. In one embodiment, the determining (208) is performed as described above in connection with FIG. 2.

[0057] The method 400 includes generating, by a representation generator, distributed representations using the presence information (410). In one embodiment, the generating step is performed as described above in connection with FIG. 2 (210).

[0058] The method 400 includes receiving, by a sparsification module executing on a computing device, identification information for a maximum level of sparsity 412. In one embodiment, the receiving step is performed 212 as described above in connection with FIG.

[0059] The method 400 includes reducing, by the sparsification module, the total number of set bits in the distributed representation based on the maximum level of sparsity to generate a sparse distributed representation (SDR) having a reference fill grade (fillgrAde) (414). In one embodiment, the reduction step is performed as described above in connection with FIG. 2.

[0060] The method 400 includes determining 416, by a similarity engine executing on a computing device, a distance between a first SDR of a first data item and a second SDR of a second data item. In one embodiment, the similarity engine 304 calculates the distance between at least two SDRs. Distance measures include, but are not limited to, Direct Overlap, Euclidean distance (e.g., determining the average distance between two points on an SDR, similar to how a human measures it with a ruler), Jacquard distance, and cosine similarity. The shorter the distance between two SDRs, the higher the similarity. Also, the higher the similarity, the higher the semantic relatedness of the data elements represented by the SDRs (in the case of semantically folded SDRs). In one embodiment, the similarity engine 304 counts the number of bits set in both the first SDR and the second SDR (e.g., points where both SDRs are set to 1). In another embodiment, the similarity engine 304 identifies a first point in the first SDR (e.g., an arbitrarily selected first bit set to 1), finds the same point in the second SDR, and determines the closest set bit in the second SDR. By determining the set bit in the second SDR that is closest to the set bit in the first SDR (each set bit in the first SDR), the similarity engine 304 can determine the total distance by summing the distances at each point and dividing by the number of points. Those skilled in the art will appreciate that other mechanisms can be used to determine the distance between SDRs. In some embodiments, similarity is not an absolute measure, but may vary based on different contexts possessed by data items. Thus, in one of these embodiments, the similarity engine 304 also analyzes the overlap relationship (aspect) between two SDRs. For example, the overlap relationship is used to add a weighting function in the similarity calculation. As another example, a similarity measure can be used.

[0061] The method 400 includes providing, by the similarity engine, an identification of a semantic similarity level between the first data item and the second data item based on the determined distance (418). The similarity engine 304 can determine that the distance between the two SDRs exceeds a maximum threshold of similarity, and therefore the represented data items are not similar. Alternatively, the similarity engine 304 can determine that the distance between the two SDRs does not exceed a maximum threshold of similarity, and therefore the represented data items are similar. The similarity engine 304 can identify the similarity level based on a range, threshold, or other calculation. In one embodiment, the semantic closeness between two data items can be determined because the SDR actually represents the meaning of the data item (as represented by multiple semantic features).

[0062] In some embodiments, system 100 provides a user interface (not shown) through which a user can input data items and receive similarity level identification information. The user interface may provide the functionality to a user accessing machine 100 directly. Alternatively, the user interface may provide the functionality to a user accessing machine 100 over a computer network. By way of example, and not by way of limitation, a user may input a set of data items, such as "music" and "apple," into the user interface. Similarity engine 304 receives the data items and generates SDRs for the data items as described above in connection with FIGS. 1A-1C and 2. Continuing with the example above, similarity engine 304 may then compare the two SDRs as described above. Although not required, the similarity engine 304 may provide the user with a graphical representation of each SDR via the user interface, allowing the user to visually see how each data item is semantically mapped (e.g., seeing clustered points in a semantic map representing the data item's use in the reference collection used to train the reference map generator 106). As mentioned above, some embodiments of the methods and systems described herein use associated presence information to apply a sparsification process at the time of generating a distributed representation for a data item in a list of data items, while other embodiments may prefer to delay the application of the sparsification step by the sparsification module. For example, in certain scenarios, such as when optimizing for a higher level of precision in search processing, it may be beneficial to create a mixed SDR for one or more data items (e.g., data items within a particular document) and then sparsify it. Sparsification typically involves removing granularity, opting to store and use smaller SDRs (e.g., when optimizing for faster collection without increasing latency, regardless of corpus size). However, when sparsifying, there may be a loss of granularity, or a variety of meanings, in the semantic meaning of the data items. For example, if the term “organ” in a particular corpus is more often associated with musical instruments than with animal bodies, then once unused semantic meanings (within that corpus) are removed, in SDR, “organ” can only mean “musical instrument” (again, within that corpus and for that SDR); if the 200 less common semantic meanings of a data item are removed, those semantic meanings will not be available after sparsification. Thus, in embodiments where search accuracy is optimized over size or speed, delaying sparsification until a later point (e.g., at least until after the generation of an SDR for each data item) may allow the system to substantially improve resolution. This does not require more effort, but requires a different goal for optimizing the system. Thus, as described in connection with FIG. 19 , in some embodiments of the methods and systems described herein, the system determines to apply a delayed sparsification process.

[0063] Referring to FIG. 19 , a flow diagram illustrates one embodiment of a method for lazy sparsifying a mixed distributed representation of multiple data items for use in identifying levels of similarity between the data items. Briefly, the method 1900 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional metric space by a reference map generator executing on a computing device to generate a semantic map (1902). The method 1900 also includes associating a coordinate pair with each of the set of data documents by the semantic map (1904). The method 1900 also includes generating, by a parser executing on the computing device, a list of data items present in the set of data documents (1906). The method 1900 also includes determining, by a representation generator executing on the computing device, presence information for each data item in the list, including (i) the number of data documents in which the data item occurs, (ii) the number of occurrences of the data item in each data document, and (iii) a coordinate pair associated with each data document in which the data item occurs (1908). The method 1900 includes generating, by a representation generator, a distributed representation for each data item in the list using the presence information (1910). The method 1900 also includes combining, by the representation generator, a first distributed representation for the first data item and a second distributed representation for the second data item to form a mixed distributed representation (1912). The method 1900 also includes adding, by the representation generator, a number of set bits at each position in the mixed distributed representation (1904). The method 1900 also includes receiving, by a sparsification module executing on a computing device, an identification of a maximum level of sparsity (1916). The method 1900 also includes, by the sparsification module, proportionally reducing the total number of set bits in the distributed representation based on the maximum level of sparsity to generate a mixed sparse distributed representation (SDR) having a baseline fill grade (1918). The method 1900 also includes storing the mixed SDR in an SDR database (1920).

[0064] 1A-1C and 2-4, a method 1900 includes clustering (1902) a set of data documents selected according to at least one criterion in a two-dimensional metric space by a reference map generator running on a computing device to generate a semantic map. In one embodiment, the clustering is performed as described above in connection with FIG. 2 (202).

[0065] The method 1900 includes associating 1904 the coordinate pairs with each of the set of data documents through a semantic map. In one embodiment, the association is performed as described above in connection with Figure 2 (204).

[0066] The method 1900 includes generating 1906, by a parser executing on a computing device, a list of data items present in the set of data documents. In one embodiment, the generating is performed as described above in connection with FIG. 2 (206).

[0067] The method 1900 includes determining, by a representation generator executing on a computing device, for each data item in the list, presence information including (i) the number of data documents in which the data item occurs, (ii) the number of occurrences of the data item in each data document, and (iii) a coordinate pair associated with each data document in which the data item occurs (1908). In one embodiment, the determining is performed as described above in connection with FIG. 2 (208).

[0068] The method 1900 includes generating, by a representation generator, a distributed representation for each data item in the list using the presence information (1910). In one embodiment, the generation is performed as described above in connection with FIG. 2 (210).

[0069] The method 1900 includes combining, by a representation generator, a first distributed representation of a first data item and a second distributed representation of a second data item to form a mixed distributed representation (1912). The method 1900 also includes adding, by the representation generator, a number of set bits at each position in the mixed distributed representation (1904). The representation generator may form the mixed distributed representation as described above in connection with FIG. 2, but instead of forming the mixed distributed representation after sparsifying each of the individual sparsified distributed representations, the combined distributed representation is not yet sparsified, preventing loss of granularity in the underlying vectors. By way of example and not limitation, the system may sparsify the mixed distributed representation after receiving a request to identify a similarity level. At that point, sparsification can occur without loss of granularity.

[0070] The method 1900 includes receiving, by a sparsification module executing on a computing device, an identification of a maximum level of sparsity (1916). In one embodiment, the receiving occurs as described above in connection with FIG. 2 (212).

[0071] The method 1900 includes, by a sparsification module, proportionally reducing the total number of set bits in the distributed representation based on the maximum level of sparsity to generate a mixed sparse distributed representation (SDR) having a baseline fill grade (1918). In one embodiment, the reduction is performed as described above in connection with FIG. 2 (214). The method 1900 includes storing (1920) the mixed SDR in a database of SDRs.

[0072] In some embodiments, the similarity engine 304 accepts only one data item from the user. Referring to FIG. 5, a flowchart illustrates one embodiment of such a method. In summary, the method 500 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional metric space by a reference map generator running on a first computing device to generate a semantic map (502). The method 500 also includes associating coordinate pairs with each of the set of data documents through the semantic map (504). The method 500 also includes generating a list of data items present in the set of data documents by a parser running on the first computing device (506). The method 500 also includes determining presence information (including (i) through (iii) below) for each data item in the list by a representation generator running on the first computing device (508).

[0073] (i) the number of data documents in which the data item exists; (ii) the number of times the data item exists in each data document; and (iii) coordinate pairs associated with each data document in which the data item exists. Method 500 includes generating, by a representation generator, a distributed representation for each data item in the list of data items using the presence information (510). Method 500 includes receiving, by a sparsification module executing on the first computing device, an identification of a maximum level of sparsity (512). Method 500 includes generating, by the sparsification module, a sparse distributed representation (SDR) having a reference fill grade by reducing the total number of set bits in the distributed representation based on the maximum level of sparsity (514). Method 500 includes storing each generated SDR in an SDR database (516). Method 500 includes receiving, by a similarity engine executing on a second computing device, the first data item from a third computing device (518). The method 500 includes determining, by the similarity engine, a distance between a first SDR of a first data item and a second SDR of a second data item retrieved from an SDR database (520). The method 500 also includes providing, by the similarity engine, to the third computing device, an identification of the second data item and a level of semantic similarity between the first and second data items based on the determined distance (522).

[0074] In some embodiments, steps 502-516 are performed according to the methods previously described in connection with Figure 2 and steps 202-214.

[0075] The method 500 includes receiving 518 the first data item from the third computing device by a similarity engine executing on the second computing device. In one embodiment, the system 300 includes a user interface (not shown). A user can input the first data item using the user interface. In another embodiment, the fingerprint module 302 generates an SDR of the first data item. In yet another embodiment, the representation generator 114 generates the SDR.

[0076] Method 500 includes determining (520) with the similarity engine a distance between a first SDR of the first data item and a second SDR of the second data item retrieved from the SDR database. In one embodiment, method 500 includes determining a distance between the first SDR of the first data item and the second SDR of the second data item, as described above in connection with FIG. 4 , (416). In some embodiments, similarity engine 304 retrieves the second data item from SDR database 120. In one of these embodiments, similarity engine 304 reviews each entry in SDR database 120 to determine whether a level of similarity exists between the retrieved item and the received first data item. In other of these embodiments, the system 300 implements current text indexing technology and text search libraries to (i) efficiently index the semantic fingerprint (i.e., SDR) collection and (ii) allow the similarity engine 304 to more efficiently identify a second SDR for a second data item than by "brute force" iterating through every item in the database 120 one by one.

[0077] Method 500 includes providing (522) to the third computing device, by the similarity engine, an identification of the second data item and an identification of the level of semantic similarity between the first and second data items based on the determined distance. In one embodiment, similarity engine 304 provides the identification via a user interface. In other embodiments, similarity engine 304 provides an identification of the level of semantic similarity between the first and second data items based on the determined distance, as described above in connection with FIG. 4 , (418). It will be appreciated that in some embodiments, similarity engine 304 repeats the process of (1) retrieving a third SDR for the third data item from the SDR database, (2) determining the distance between the first SDR of the first data item and the third SDR of the third data item, and providing an identification of the level of semantic similarity between the first and third data items based on the determined distance.

[0078] In one of these embodiments, the similarity engine 304 may return a list of other data items that are most similar to the received data item. By way of example, the similarity engine 304 may generate an SDR 118 for the received data item and then search the SDR database 120 for other SDRs similar to the SDR 118. In one embodiment, the data item module 308 provides the above functionality. By way of example, and not limitation, the similarity engine 304 (or alternatively, the data item module 308) may compare the SDR 118 for the received data item with each of multiple SDRs in the SDR database 120, as described above, and return a list of data items that meet similarity requirements (e.g., the distance between the data items is less than or equal to a predetermined threshold). In some embodiments, the similarity engine 304 returns the SDRs that are most similar to a particular SDR (as opposed to returning the data item itself).

[0079] In some embodiments, the method for receiving a data item (sometimes referred to as a keyword) and the method for identifying similar data items are performed as described above in connection with FIG. 2, (202)-(214). In some embodiments, the data item module 308 provides the functionality. In one of these embodiments, the method includes receiving a data item. The method includes receiving a request for a most similar data item that is not identical to the received data item. In another of these embodiments, the method includes generating a first SDR for the received data item. In yet another of these embodiments, the method includes determining a distance between the first SDR and each SDR in the SDR database 120. In yet another of these embodiments, the method includes providing a list of data items where the distance between the SDR of the listed data item and the first SDR is less than a threshold. Alternatively, the method includes providing a list of data items where each data item has a similarity level with the received data item that exceeds a threshold. In some embodiments, the method for identifying similar data items provides functionality for receiving a data item or an SDR for a data item, and for generating a list of SDRs as directed by increasing distance (e.g., Euclidean distance). In one of these embodiments, the system 100 provides functionality for returning all contextual data items, i.e., data items within the conceptual space in which the submitted data item resides.

[0080] The data item module 308 may return similar data items to either the user, another module, or an engine (eg, the disambiguation module 306) that provided the received data item.

[0081] In some embodiments, the system may generate a list of similar data items and send the list to a system (such as a system within system 300 or a third-party search system) that executes the query. For example, a user may enter data items into a user interface for executing a query (e.g., a search engine), which may then forward the data items to the query module 601. The query module 601 may automatically invoke a component of the system (e.g., the similarity engine 304) to generate a list of similar data items and provide the data items to the user interface for execution as an additional query to the user's initial query, thereby broadening the user's search results. As another example, and as described in more detail in connection with FIGS. 6A-6C , the system may generate a list of similar data items and provide the data items directly to a third-party search system that returns expanded search results to the user via the user interface. The third-party search system (sometimes referred to herein as an enterprise search system) may be of any type and form. As previously discussed in connection with full-text search system 122, a wide variety of such systems are available, the capabilities and performance of which may be enhanced using the methods and systems described herein.

[0082] Referring to the block diagram of Figure 6A, one embodiment of a system 300 for expanding queries in a full-text search system is disclosed. In general, the system 300 includes the elements previously described in connection with Figures 1A-1C and 3. The system 300 also provides the functionality previously described in connection with Figures 1A-1C and 3. The system 300 includes a machine 102d that executes a query module 601. The query module 601 executes a query expansion module 603, a ranking module 605, and a query input processing module 607.

[0083] In one embodiment, the query module 601 receives a query term, directs the generation of an SDR for the received term, and directs the identification of similar query terms. In another embodiment, the query module 601 communicates with an enterprise search system provided by a third party. For example, the query module 601 may include one or more interfaces (e.g., application program interfaces) that communicate with the enterprise search system. In some embodiments, the query module 601 executes as at least part of a software program. In one embodiment, the query module 601 executes as at least part of a hardware module. In yet another embodiment, the query module 601 executes on the machine 102d.

[0084] In one embodiment, the query input processing module 607 accepts query terms from a user of the client 102c. In another embodiment, the query input processing module 607 identifies the type of query term (e.g., individual word, group of words, sentence, paragraph, document, SDR, or other expression used to identify similar terms). In some embodiments, the query input processing module 607 executes as at least part of a software program. In one embodiment, the query input processing module 607 executes as at least part of a hardware module. In yet another embodiment, the query input processing module 607 executes on the machine 102d. In another embodiment, the query module 601 communicates with the query input processing module 607. Alternatively, the query module 601 provides the functionality of the query input processing module 607.

[0085] In one embodiment, the query expansion module 603 accepts a query term from a user of the client 102c. In another embodiment, the query expansion module 603 accepts a query term from the query input processing module 607. In yet another embodiment, the query expansion module 603 directs the generation of an SDR for the query term. In another embodiment, the query expansion module 603 directs the similarity engine 304 to identify one or more terms similar to the query term (based on the distance between the SDRs). In some embodiments, the query expansion module 603 executes as at least part of a software program. In one embodiment, the query expansion module 603 executes as at least part of a hardware module. In yet another embodiment, the query expansion module 603 executes on the machine 102d. In another embodiment, the query module 601 communicates with the query expansion module 603. Alternatively, the query module 601 provides the functionality of the query expansion module 603.

[0086] Referring to the flowchart of FIG. 6B, one embodiment of a method 600 for query expansion in a full-text search system is disclosed. In summary, the method 600 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional distance space by a reference map generator running on a first computing device to generate a semantic map (602). The method 600 includes associating coordinate pairs with each of the set of data documents through the semantic map (604). The method 600 includes generating a list of terms present in the set of data documents by a parser running on the first computing device (606). For each term in the list, the method 600 includes determining presence information (including the following (i) to (iii)) by an expression generator running on the first computing device (608): (i) the number of data documents in which the term occurs, (ii) the number of terms present in each data document, and (iii) the coordinate pairs associated with each data document in which the term occurs. The method 600 includes generating 610, by the representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information for each term. The method 600 includes storing 612 each generated SDR in an SDR database. The method 600 includes receiving 614 a first term from a third computing device by a query expansion module executing on a second computing device. The method 600 includes determining 616, by a similarity engine executing on a fourth computing device, a semantic similarity level between a first SDR for the first term and a second SDR for a second term retrieved from the SDR database. The method 600 includes submitting 618, by the query expansion module, to a full-text search system using the first and second terms to identify a set of documents each containing at least one term similar to at least one of the first and second terms. Method 600 includes sending, by the query expansion module, an identification of each of the set of documents to the third computing device (620).

[0087] In some embodiments, steps 602-612 are performed as previously described in connection with FIG. 2 and steps 202-214.

[0088] The method 600 includes receiving 614 the first term from the third computing device by a query expansion module executing on the second computing device. In one embodiment, the query expansion module 603 receives the first data item as described above in connection with FIG. 5 , (518). In another embodiment, the query input processing module 607 receives the first term. In yet another embodiment, the query input processing module 607 sends the first term to the fingerprint module 302 along with a request for SDR generation. In yet another embodiment, the query input processing module 607 sends the first term to the engine 101 for generation of the SDR by the representation generator 114.

[0089] The method 600 includes determining, with a similarity engine executing on a fourth computing device, a semantic similarity level between the first SDR of the first term and a second SDR of the second term retrieved from the SDR database (616). In one embodiment, the similarity engine 304 determines the semantic similarity level as described above in connection with FIG. 5, (520).

[0090] The method 600 includes sending a query to a full-text search system by the query expansion module using the first term and the second term to identify a set of documents each containing at least one term similar to at least one of the first term and the second term (618). In some embodiments, the similarity engine 304 provides the second term to the query module 601. It will be appreciated that the similarity engine may provide multiple terms having an similarity to the first term that exceeds a similarity threshold. In other embodiments, the query module 601 may include one or more application program interfaces that send a query including one or more search terms to a third-party enterprise search system.

[0091] The method 600 includes sending, by the query expansion module, an identification of each of the set of documents to the third computing device (620). Referring to the flowchart of Figure 6C, one embodiment of a method 650 for query expansion in a full-text search system is disclosed. In summary, the method 650 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional distance space by a reference map generator executing on a first computing device to generate a semantic map (652). The method 650 also includes associating coordinate pairs with each of the set of data documents through the semantic map (654). The method 600 also includes generating a list of terms present in the set of data documents by a parser executing on the first computing device (656). For each term in the list, the method 650 also includes determining presence information (including the following (i)-(iii)) by an expression generator executing on the first computing device (658): (i) the number of data documents in which the term occurs, (ii) the number of times the term occurs in each data document, and (iii) the coordinate pairs associated with each data document in which the term occurs. The method 650 includes generating, by a representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information for each term (660). The method 650 includes storing each generated SDR in an SDR database (662). The method 650 includes receiving, by a query expansion module executing on a second computing device, a first term from a third computing device (664). The method 650 includes determining, by a similarity engine executing on a fourth computing device, a semantic similarity level between a first SDR for the first term and a second SDR for a second term retrieved from the SDR database (666). The method 650 includes sending, by the query expansion module, the second term to the third computing device (668).

[0092] In one embodiment, steps 652-666 are performed as described above in connection with steps 602-612. However, instead of providing one or more of the terms identified by the similarity engine directly to the enterprise search system, method 650 includes sending the second terms to the third computing device by the query expansion module (668). In such a method, a user of the third computing device can review and modify the second terms before the query is sent to the enterprise search system. In some embodiments, a user desires additional control over the query. In one embodiment, a user prefers to execute the query themselves. In another embodiment, a user desires to modify the terms identified by the system before submitting the query. In yet another embodiment, providing the identified terms to the user allows the system to request feedback from the user regarding the identified terms. In one of these embodiments, for example, the accuracy of the similarity engine in identifying the second terms can be evaluated. In another of these embodiments, for example, the user provides an indication that the second term is a type of term that interests the user (e.g., the second term is a type related to an area of ​​expertise that the user is currently researching or developing).

[0093] In some embodiments, a method for evaluating at least one Boolean operator includes receiving, by expression engine 310, at least one data item and at least one Boolean operator. The method includes performing the functions described above in connection with FIG. 2 and steps (202) through (214). In one embodiment, expression engine 310 receives multiple data items combined by a user using Boolean operators and parentheses. For example, a user may submit a phrase such as "Jaguar SUB Porsche," and expression engine 310 evaluates the phrase and generates a modified version of the SDR for the expression. Thus, in another embodiment, expression engine 310 generates a first SDR 118 for a first data item in the received phrase. In yet another embodiment, expression engine 310 identifies the Boolean operator in the received phrase (e.g., by determining that the second data item in a three-data item phrase is the Boolean operator, or by comparing each data item in the received phrase to a list of Boolean operators to determine whether the data item is a Boolean operator). Expression engine 310 evaluates the identified Boolean operator to determine how to modify the first data item. For example, the expression evaluator 310 may determine that the received phrase contains the Boolean operator "SUB." The expression engine 310 then determines to generate a second SDR for the data item following the Boolean operator (e.g., "Porsche" in the example phrase above) and generates a third SDR by removing from the first SDR the points that appeared in the second SDR. The third SDR is then an SDR for the first data item, but does not include the SDR for the second data item. Similarly, if the expression engine 310 determines that the Boolean operator was "AND," the expression engine 310 generates a third SDR using only the points common to the first and second SDRs. Thus, the expression engine 310 receives data items, compound data items, and SDRs combined using Boolean operators and parentheses, and returns an SDR that reflects the Boolean result of the formulated expression.The resulting modified SDR may be returned to the user or provided to other engines in system 200, such as similarity engine 304. Those skilled in the art will appreciate that Boolean operators include, but are not limited to, AND, OR, XOR, NOT, and SUB.

[0094] In some embodiments, a method for identifying multiple subcontexts for a data item includes receiving the data item by a disambiguation module 306. The method includes performing the functions described above with reference to FIG. 2, (202)-(214). In one embodiment, the method includes generating a first SDR for the received data item. In another embodiment, the method includes generating a list of data items having SDRs similar to the first SDR. For example, the method includes providing the first SDR to the similarity engine and requesting a list of similar SDRs as described above. In yet another embodiment, the method includes analyzing one of the listed SDRs that is similar to but does not match the first SDR, and removing (e.g., via binary subtraction) from the first SDR any points (e.g., set bits) that are also present in the listed SDR to generate a modified SDR. In another embodiment, the method includes repeating the process of removing points that are present in both the first SDR and the similar (but not identical) SDR. This step continues until all points in each list of similar SDRs have been removed from the first SDR. For example, upon receiving a request for data items similar to the data item "Apple," the system may return data items such as "Macintosh," "iPhone," and "Operating System." If a user provides the expression "Apple SUB Macintosh" and requests similar data items from the remaining points, the system may return data items such as "Fruit," "Plum," "Orange," and "Banana." Continuing with the example, if a user next provides the expression "Apple SUB Macintosh SUB Fruit" and requests similar data items, the system may return data items such as "Records," "The Beatles," and "Pop Music." In some embodiments, the method includes extracting similar SDR points from the largest clusters of the first SDR, rather than from the entire SDR. This provides a more optimal solution.

[0095] In some embodiments, as described above, a data item may refer to an item other than a word. By way of example, the system 300 (e.g., the similarity engine 304) may generate an SDR for a number, compare the SDR to reference SDRs generated from other numbers, and provide a list of similar data items to the user. For example, and without limitation, the system 300 (e.g., the similarity engine 304) may generate an SDR for the data item "100.1" and determine that the SDR has a similar pattern to an SDR for a data item associated with a patient diagnosed with a fever caused by an infection. (For example, in an embodiment in which a doctor or healthcare provider implements the above-described method and system, for a data item generated based on a patient's physical characteristics (e.g., body temperature or other characteristics), the system may store an association between the data item (100.1) and its identification as a reference data item for a patient with a fever.) Determining that a data item has a similar pattern provides the ability to identify commonalities between the dynamically generated SDR and the reference SDR, allowing the user to better understand the capture of a particular data item. Thus, in some embodiments, the reference SDRs are linked to eligible diagnoses, allowing a new patient's SDR profile to be matched to the diagnosed pattern and from which a set of possible diagnoses for the new patient can be inferred. In one of these embodiments, the collection of possible diagnoses allows a user to "see" where points (e.g., semantic features of certain data items) overlap or match. In such an embodiment, the diagnosis that most closely resembles the new patient's SDR pattern is the predicted diagnosis.

[0096] As another example, without limitation, the set of data documents 104 may include a log of captured flight data generated by sensors on an airplane (as opposed to, e.g., an encyclopedic entry about flights). The log of captured data may include alphanumeric or primarily numeric data items. In such an example, the system 100 may provide functionality for generating an SDR for a variable (e.g., a variable related to any type of flight data) and compare the generated SDR to a reference SDR (e.g., an SDR for a data item used as a reference item known to have certain characteristics, such as facts about the flight during which the data item was generated, e.g., flight at a particular altitude, altitude characteristics such as too high or too low). As another example, the system 100 may generate a first SDR for “500 degrees” and determine that the first SDR is similar to a second SDR for “28,000 feet.” Subsequently, the system 100 determines that the second SDR is a reference SDR for a data item that indicates a characteristic of the flight (e.g., too high, too low, too fast, etc.), and the system 100 may prompt the user who initiated the data item “500 (degrees)” to acknowledge the capture of the data item.

[0097] In some embodiments, a method is provided for dividing a document into portions (also referred to herein as slicing) while considering the topic structure of the presented text. In one embodiment, a data item module 308 receives a document to be divided into topics. In another embodiment, the data item module 308 identifies a location in the document having a semantic fingerprint different from a second location and divides the document into two slices, one slice including the first location and the other slice including the second location. The method includes performing the functionality described above with reference to FIG. 2, (202)-(214). In one embodiment, the method includes generating an SDR 118 for each sentence (e.g., a string separated by a period) in the passage. In another embodiment, the method includes comparing a first SDR 118A for a first sentence with a second SDR 118b for a second sentence. For example, the method includes sending the two SDRs to a similarity engine 304 for comparison. In yet another embodiment, the method includes inserting a break in the sentence after the first sentence when the distance between two SDRs exceeds a predetermined threshold. In another embodiment, the method includes determining not to insert a break in the sentence after the first sentence when the distance between two SDRs does not exceed a predetermined threshold. In yet another embodiment, the method includes repeatedly comparing the second sentence with subsequent sentences. In another embodiment, the method includes repeatedly comparing sentences until the end of the document is reached. In yet another embodiment, the method includes using the inserted break to generate slices of the document (e.g., extending a section of the document back to the first inserted break as the first slice). In some embodiments, having multiple small slices is preferable to a single document. However, arbitrary division of a document (e.g., by length or word count) may be less efficient or useful than topic-based division of a document.In one of these embodiments, by comparing the composite SDRs of sentences, system 300 can determine where the topic of a document changes, forming logical division points. In other of these embodiments, system 300 may provide a semantic fingerprint index in addition to a traditional index. Further examples of topic slicing are described in connection with Figures 7A-7B below.

[0098] With reference to FIG. 7B and the block diagram of FIG. 7A, one embodiment of a system 700 for providing topic-based documents to a full-text search system is disclosed. In general, the system 700 includes the elements previously described with reference to FIGS. 1A-1C and 3. The system 700 also provides the functionality previously described with reference to FIGS. 1A-1C and 3. The system 700 further includes a topic slicing module 702. In one embodiment, the topic slicing module 702 receives a document, directs the generation of an SDR for the received document, and directs the generation of a subdocument in which sentences with a similarity below a threshold are located in a different document, i.e., a different data structure. In another embodiment, the topic slicing module 702 communicates with an enterprise search system provided by a third party. In some embodiments, the topic slicing module 702 executes as at least part of a software program. In one embodiment, the topic slicing module 702 executes as at least part of a hardware module. In yet another embodiment, the topic slicing module 702 executes on machine 102b.

[0099] 7A and 7B, an embodiment of a system 750 for providing topic-based documents to a full-text search system is disclosed. The method 750 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional distance space by a reference map generator running on a first computing device to generate a semantic map (752). The method 750 includes associating coordinate pairs with each of the set of data documents through the semantic map (754). The method 750 includes generating a list of terms present in the set of data documents by a parser running on the first computing device (756). For each term in the list, the method 750 includes determining presence information (including the following (i) through (iii)) by an expression generator running on the first computing device (758): (i) the number of data documents in which the term occurs, (ii) the number of times the term occurs in each data document, and (iii) the coordinate pairs associated with each data document in which the term occurs. The method 750 includes generating, by a representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information (760). The method 750 includes storing each generated SDR in an SDR database (762). The method 750 includes receiving, by a topic slicing module executing on a second computing device, a second set of documents from a third computing device associated with an enterprise search system (764). The method 750 includes generating, by the representation generator, a composite SDR for each sentence in each of the second set of documents (766). The method 750 includes determining, by a similarity engine executing on the second computing device (768), a distance between a first composite SDR for a first sentence and a second composite SDR for a second sentence. The method 750 includes generating, by the topic slicing module, a second document including the first sentence and a third document including the second sentence based on the determined distance (770). The method 750 includes sending, by a topic slicing module, the second document and the third document to the third computing device (772).

[0100] In one embodiment, steps 752-762 are performed as previously described in connection with FIG. 2 and steps 202-214.

[0101] The method 750 includes receiving 764, by a topic slicing module executing on a second computing device, a second set of documents from a third computing device associated with an enterprise search system. In one embodiment, the topic slicing module 702 receives the second set of documents for processing to generate a version of the second set of documents optimized for indexing by the enterprise search system, which may be a conventional search system. In other embodiments, the topic slicing module 702 receives the second set of documents for processing to generate a version of the second set of documents optimized for indexing by the search system provided by the system 700. This is described in more detail below with reference to FIGS. 9A and 9B. In some embodiments, the received second set of documents includes one or more XML documents. For example, the third computing device may have converted one or more enterprise documents into XML documents for improved indexing.

[0102] Method 750 includes generating, by a representation generator, a composite SDR for each sentence in each of the second set of documents (766). As discussed above in connection with FIG. 2, for example, if a phrase, sentence, paragraph, or other combination of data items needs to be converted into an SDR that reflects the "union properties" of the individual SDRs (e.g., the combination of SDRs for each word in a sentence), system 100 may convert the individual data items into their SDRs (either by dynamically generating them or by retrieving pre-generated SDRs) and form the composite SDR from the individual SDRs using a binary OR operation. The result may be sparsified by sparsification module 116.

[0103] The method 750 includes determining, with a similarity engine executing on the second computing device, a distance between the first composite SDR of the first sentence and the second composite SDR of the second sentence (768). In one embodiment, the similarity engine determines the distance as described above in connection with FIG. 4 and step (416).

[0104] The method 750 includes generating 770, by a topic slicing module, a second document including the first sentence and a third document including the second sentence based on the determined distance. The topic slicing module determines that the distance determined by the similarity engine exceeds a similarity threshold and therefore the second sentence relates to a different topic than the first sentence and should therefore be placed in a different document (or other data structure). In another embodiment, the similarity engine provides the topic slicing module 702 with an identification of a level of similarity between the first sentence and the second sentence based on the determined distance (as described above in connection with FIG. 4). The topic slicing module 702 then determines that the level of similarity does not meet a threshold level of similarity and decides to place the second sentence in a different document from the first sentence. However, in another embodiment, the topic slicing module 702 determines that the determined distance (and / or similarity) meets a similarity threshold. The topic slicing module 702 then determines that the first sentence and the second sentence are topically similar and should be left together in one document.

[0105] In yet another embodiment, the method includes repeatedly comparing the second sentence with subsequent sentences. In another embodiment, the method includes repeatedly going through the document and repeating the comparison of sentences until the end of the document is reached.

[0106] The method 750 sends (772) the second document and the third document to a third computing device by the topic slicing module.

[0107] Referring to Figure 8B in conjunction with Figure 8A, a flowchart illustrates one embodiment of a method 850 for extracting keywords from text documents. The method 850 includes clustering a set of data documents selected according to at least one criterion in a two-dimensional distance space by a reference map generator executing on a first computing device to generate a semantic map (852). The method 850 includes associating coordinate pairs with each of the set of data documents by the semantic map (854). The method 850 includes generating, by a parser executing on the first computing device, a list of terms present in the set of data documents (856). The method 850 includes determining, by a representation generator executing on the first computing device, presence information for each term in the list, including: (i) the number of data documents in which the term occurs, (ii) the number of occurrences of the term in each data document, and (iii) coordinate pairs associated with each data document in which the term occurs (858). The method 850 includes generating, by a representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information (860). The method 850 includes storing each generated SDR in an SDR database (862). The method 850 includes receiving, by a keyword extraction module executing on a second computing device, a document from a third computing device associated with a full-text search system (864). The method 850 includes generating, by the representation generator, at least one SDR for each term in the received document (866). The method 850 includes generating, by the representation generator, a composite SDR for the received document based on the generated at least one SDR (868). The method 850 includes selecting, by the keyword extraction module (870), a plurality of term SDRs that, when combined, generate a composite SDR that has a semantic similarity level with the composite SDR for the document that meets a threshold. The method 850 includes modifying, by the keyword extraction module, a keyword field of the received document to include the plurality of terms (872).The method 850 includes sending the modified document by the keyword extraction module to a third computing device (874).

[0108] In one embodiment, steps 852 through 862 are performed as described above in connection with FIG. 2 and steps 202 through 214.

[0109] 1A-1C and 3 and provides the functionality described above in connection with those figures. System 800 further comprises a keyword extraction module 802. In one embodiment, keyword extraction module 802 receives documents, directs the generation of SDRs for the received documents, identifies keywords for the received documents, and modifies the received documents to include the identified keywords. In other embodiments, keyword extraction module 802 communicates with an enterprise search system provided by a third party. In some embodiments, keyword extraction module 802 is implemented at least in part as a software program. In other embodiments, keyword extraction module 802 is implemented at least in part as a hardware module. In yet other embodiments, keyword extraction module 802 executes on machine 102b.

[0110] The method 850 includes receiving, by a keyword extraction module executing on the second computing device, a document of the second set of documents from a third computing device associated with the full-text search system (864). In one embodiment, the keyword extraction module 802 receives the document in association with the topic slicing module 702 as shown in FIG. 7 and step (764).

[0111] The method 850 includes generating 866, by the representation generator, at least one SDR for each term in the received document. In one embodiment, the keyword extraction module 802 sends each term in the received document to the representation generator 114 to generate at least one SDR. In another embodiment, the keyword extraction module 802 sends each term in the received document to the fingerprint module 302 to generate the at least one SDR.

[0112] In some embodiments, the keyword extraction module 802 sends the document to the fingerprint module 302 along with a request to generate a composite SDR for each sentence in the document. In other embodiments, the keyword extraction module 802 sends the document to the representation generator 114 along with a request to generate a composite SDR for each sentence in the document.

[0113] The method 850 includes generating 868, by the representation generator, a composite SDR for the received document based on the generated at least one SDR. In one embodiment, the keyword extraction module 802 requests generation of the composite SDR from the representation generator 114. In another embodiment, the keyword extraction module 802 requests generation of the composite SDR from the fingerprint module 302.

[0114] The method 850 includes having the keyword extraction module select 870 a plurality of term SDRs that, when combined, generate a composite SDR that has a semantic similarity level with the composite SDR for the document that meets a threshold. In one embodiment, the keyword extraction module 802 instructs the similarity engine 304 to compare the composite SDR for the document with SDRs for a plurality of terms ("term SDRs") and generate an identification of a similarity level between the plurality of terms and the document itself. In some embodiments, the keyword extraction module 802 identifies a plurality of terms that meets the threshold by causing the similarity engine 304 to iterate through combinations of term SDRs, generate comparisons with the composite SDR for the document, and return a list of the semantic similarity level between the document and each combination of terms. In another of these embodiments, the keyword extraction module 302 identifies a plurality of terms that meets the threshold and has a level of semantic similarity with a document that includes as few terms as possible.

[0115] The method 850 includes modifying, by the keyword extraction module, a keyword field of the received document to include the plurality of terms 872. As noted above, the received document may be a structured document, such as an XML document, and may have portions into which the plurality of terms are inserted by the keyword extraction module 802.

[0116] The method 850 includes sending the modified document by the keyword extraction module to a third computing device (874).

[0117] As noted above, an enterprise search system may include an implementation of a conventional search system. Conventional search systems include those described in connection with full-text search system 122 above (e.g., Lucene-based systems, open-source systems such as Xapian, commercial systems such as Autonomy IDOL or COGITO, and other systems detailed above). In this specification, the terms "enterprise search system" and "full-text search system" may be used interchangeably. The methods and systems described in FIGS. 6 through 8 illustrate enhancements to such enterprise systems. That is, by implementing the methods and systems described herein, entities selling such enterprise systems can enhance the functionality they are selling—adding keywords or expanding query terms for users and automatically providing them to the current system, making indexing more efficient, and so on. However, entities selling search systems to their users may desire more than just enhancing certain aspects of their current systems by replacing them entirely, or by implementing an improved search system in the first place. Accordingly, in some embodiments, an improved search system is provided.

[0118] Referring to FIG. 9A, a diagram illustrates one embodiment of a system 900 implementing a full-text search system 902. In one embodiment, the system 900 has the functionality described above in connection with FIGS. 1A-1C, 3, 6A, 7A, and 8A. The search system 902 includes a query module 601, which may be provided as described above in connection with FIGS. 6A and 6B. The search system 902 includes a document fingerprint index 920. The document fingerprint index 920 may be a version of the SDR database 120. The document fingerprint index 920 may include metadata (e.g., tags). The search system 902 may include a document similarity engine 304b. For example, the document similarity engine 304b may be a copy of the similarity engine 304 that has been refined over time to work with the search system 902. The search system 902 includes an indexer 910. The indexer 910 may be provided as a hardware module or a software module.

[0119] Referring to FIG. 9B in conjunction with FIG. 9A , a method 950 includes clustering a set of data documents selected along at least one criterion in a two-dimensional metric space by a reference map generator running on a first computing device to generate a semantic map (952). The method 950 also includes associating coordinate pairs with each of the set of data documents via the semantic map (954). The method 950 also includes generating a list of terms present in the set of data documents by a parser running on the first computing device (956). For each term in the list, the method 950 also includes determining presence information (including (i) through (iii) below) by a representation generator running on the first computing device (958): (i) the number of data documents in which the term occurs, (ii) the number of times the term occurs in each data document, and (iii) coordinate pairs associated with each data document in which the term occurs. The method 950 also includes generating, by the representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information (960). The method 950 includes storing each of the generated SDRs in an SDR database (962). The method 950 includes receiving a second set of documents by a full-text search system executing on a second computing device (964). The method 950 includes generating at least one SDR for each document in the second set of documents by an expression generator (966). The method 950 includes storing each of the generated SDRs in a document fingerprint index by an indexer in the full-text search system (968). The method 950 includes receiving at least one search term from a third computing device by a query module in the search system (970). The method 950 includes querying the document fingerprint index by the query module for at least one term in the document fingerprint index that has an SDR similar to the SDR of the at least one received search term (972). The method 950 includes providing, by the query module, results of the query to a third computing device (974).

[0120] The method 950 includes clustering the set of data documents selected along at least one criterion in a two-dimensional metric space by a reference map generator executing on a first computing device to generate a semantic map (952). In some embodiments, the set of data documents is selected and clustering occurs as described above in connection with FIG. 2 and step (202). As described above in connection with FIGS. 1 and 2, a training process is performed when starting a system for use with the methods and systems described herein. As described above, the reference map generator 106 is trained using at least one set of data documents (more specifically, using the document vectors of each document in the set of data documents). As also described above, the semantic resolution of a set of documents refers to how many locations are available based on the training data. This, in some embodiments, reflects the nature of the training data (in a more general sense, this refers to how many "places" are available on a map). Different or additional training documents may be used to increase the semantic resolution. Accordingly, there are several different approaches to training the reference map generator 106. In one embodiment, a general training corpus may be used to generate an SDR for each received term (e.g., terms in corporate documents). The advantage of such an approach is that such a corpus is likely to have been selected to meet one or more training criteria, but the disadvantage is that such a corpus may, but may not, contain enough words to support a specialized corporate corpus (e.g., a highly technical corpus containing many terms with specific meanings in a particular domain or practice). In other embodiments, therefore, a set of corporate documents may be used as the training corpus.An advantage of this approach is that the documents used for training will likely include any highly technical or specialized terms shared within the company; a disadvantage is that the company's documents may not meet the training criteria (e.g., there may not be a sufficient number of documents, they may not be of sufficient length or diversity, etc.). In yet another embodiment, a general training corpus and a company's corpus are combined for training. In yet another embodiment, a specific set of technical documents is identified and processed to serve as the training corpus. For example, such documents may include important medical papers, technical specifications, or other important reference materials in a field related to the company's documents that will be used. For example, a reference corpus may be processed and used for training, after which the resulting trained database is used by engine 101. This trained database may be separately licensed to the company implementing the methods and systems described herein. These embodiments are applicable to the embodiments discussed in connection with FIGS. 6-8, as well as to the embodiments discussed in connection with FIGS. 9A and 9B.

[0121] Still referring to FIG. 9B, in some embodiments, steps (954) through (962) may be performed as described above in connection with FIGS.

[0122] The method 950 includes receiving 964 a second set of documents by a full-text search system executing on a second computing device. In one embodiment, the second set of documents includes enterprise documents (e.g., documents created by, maintained by, accessed by, or otherwise associated with the enterprise implementing the full-text search system 902). In another embodiment, the search system 902 makes one or more enterprise documents searchable. To do so, the search system 902 indexes one or more enterprise documents. In one embodiment, the search system 902 directs preprocessing of the enterprise documents (e.g., by having the topic slicing module 702 and / or the keyword extraction module 802 process the documents as described above in connection with FIGS. 7B and 8B). In another embodiment, the search system 902 directs generating an SDR for each document based on a training corpus (as described above in connection with FIGS. 1 and 2). In yet another embodiment, after generating an SDR for each document, the search system 902 enables the search process. In the search process, a query term is received (eg, by the query input processing module 607), an SDR is generated for the query term, and the query SDR is compared to the indexed SDRs.

[0123] The method 950 includes generating 966 at least one SDR for each document in the second set of documents with a representation generator. In one embodiment, the search system 902 is operable to send the document to the fingerprint module 302 to generate the at least one SDR. In another embodiment, the search system 902 is operable to send the document to the representation generator 114 to generate the at least one SDR. Examples of the at least one SDR include, but are not limited to, an SDR for each term in the document, a composite SDR for a subsection of the document (e.g., a sentence or paragraph), and a composite SDR for the document itself.

[0124] Method 950 includes storing each generated SDR in a document fingerprint index by an indexer in the full-text search system 968. In one embodiment, the generated SDRs are stored in document fingerprint index 920 in substantially the same manner as the SDRs described above are stored in SDR database 120.

[0125] The method 950 includes receiving 970 at least one search term from a third computing device by a query module in the search system. In one embodiment, the query module receives the search term as described above in connection with Figures 6A and 6B.

[0126] The method 950 includes querying, by a query module, the document fingerprint index for at least one term having an SDR similar to an SDR of the received at least one search term (972). In one embodiment, the query module 601 queries the document fingerprint index 920. In another embodiment, the system 900 includes a document similarity engine 304b, and the query module 601 directs the document similarity engine 304b to identify an SDR of at least one term in the document fingerprint index 920. In yet another embodiment, the query module 601 directs the similarity engine 304 executing on the machine 102b to identify the term. In some other embodiments, the query module 601 performs the search as described above in connection with FIGS. 6A and 6B, except that instead of sending the query to an external enterprise search system, the query module 601 sends the query to a component within the system 900.

[0127] The method 950 includes providing 974, by the query module, results of the query to a third computing device. In some embodiments, there are one or more results (e.g., one or more similar terms), and the query module 601 first ranks the results or directs another module to rank the results. The ranking may implement traditional ranking techniques. Alternatively, the ranking may include performing the methods described below in connection with FIGS. 11A and 11B.

[0128] In some embodiments, the full-text search system 902 provides a user interface (not shown) that allows a user to provide feedback on the query results. In one of these embodiments, the user interface includes a user interface element that allows a user to indicate whether the results were helpful. In other of these embodiments, the user interface includes a user interface element that allows a user to direct the query module 601 to perform a new search using one of the query results. In still other of these embodiments, the user interface may include a user interface element that allows a user to indicate that they are interested in a topic related to one of the query results and that they would like the query result identifier and / or the related topic to be saved by the user or the system 900 for future reference.

[0129] In one embodiment, the system is capable of monitoring the types of searches a user performs and creating a profile for that user based on an analysis of the SDRs of the search terms provided by that user. In such an embodiment, the profile identifies the user's level of expertise and may be provided to other users.

[0130] 10A and 10B, block diagrams illustrate embodiments of a system that matches a user's expertise with a request for user expertise based on previous search results. FIG. 10A illustrates an embodiment in which functionality for creating a user's expertise profile (e.g., user expertise profile module 1010) is provided in association with a conventional full-text search system. FIG. 10B illustrates an embodiment in which functionality for creating a user's expertise profile (e.g., user expertise profile module 1010) is provided in association with a full-text search system 902. Both of the modules illustrated in FIG. 10A and 10B may be provided as hardware or software modules.

[0131] Referring to FIG. 10C , a flowchart illustrates an embodiment of a method 1050 for matching user expertise with a request for user expertise based on previous search results. The method 1050 includes clustering a set of data documents selected along at least one criterion in a two-dimensional distance space by a reference map generator running on a first computing device to generate a semantic map (1052). The method 1050 includes associating coordinate pairs with each of the set of data documents through the semantic map (1054). The method 1050 includes generating a list of terms present in the set of data documents by a parser running on the first computing device (1056). For each term in the list, the method 1050 includes determining (1058) presence information (including (i) through (iii) below) by an expression generator running on the first computing device: (i) the number of data documents in which the term occurs, (ii) the number of times the term occurs in each data document, and (iii) the coordinate pairs associated with each data document in which the term occurs. The method 1050 includes generating, by a representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information (1060). The method 1050 includes storing each of the generated SDRs in an SDR database (1062). The method 1050 includes receiving, by a query module executing on a second computing device, at least one term from a third computing device (1064). The method 1050 includes storing, by a user expertise profile module executing on the second computing device, an identifier of a user of the third computing device and the at least one term (1066). The method 1050 includes generating, by the representation generator, an SDR for the at least one term (1068). The method 1050 includes receiving, by the user expertise profile module executing on the second computing device, a request for identification information of users associated with the second term and similar terms (1070). The method 1050 includes determining, with a similarity engine, a level of semantic similarity between the SDR of the at least one term and the SDR of the second term (1072).The method 1050 includes providing, by the user expertise profile module, to the fourth computing device an identifier of the user of the third computing device (1074).

[0132] In one embodiment, steps 1052 through 1062 are performed as described above with respect to FIG. 2 and steps 202 through 214.

[0133] The method 1050 includes receiving 1064, by a query module executing on the second computing device, at least one term from the third computing device. In one embodiment, the query module 601 receives the at least one term and executes a query as described above with respect to Figures 6A-6C and 9A-9B.

[0134] The method 1050 includes storing 1066, by a user expertise profile module executing on a second computing device, the identifier of the user of the third computing device and the at least one term. In one embodiment, the user profile module 1002 receives the identifier of the user and the at least one term from the query input processing module 607. In another embodiment, the user expertise profile module 1010 receives the identifier of the user and the at least one term from the query input processing module 607. In yet another embodiment, the user expertise profile module 1010 stores the identifier of the user and the at least one term in a database. For example, the user expertise profile module 1010 stores the identifier of the user and the at least one term in a user expertise SDR database 1012 (e.g., along with an SDR for the at least one term). In some embodiments, the method includes recording queries received from users along with the user identifier and an SDR for each query term. In some embodiments, the user profile module 1002 is also capable of receiving identification of search results that the querying user has indicated are relevant or of interest to the user.

[0135] The method 1050 includes generating 1068, by an expression generator, an SDR for the at least one term. In one embodiment, the user expertise profile module 1010 sends the at least one data item to the fingerprint module 302 to generate the SDR. In another embodiment, the user expertise profile module 1010 sends the at least one term to the expression generator 114 to generate the SDR.

[0136] In some embodiments, the user expertise profile module 1010 receives multiple data items as the user continues to query over time. In one of these embodiments, the user expertise profile module 1010 directs the creation of a composite SDR that combines an SDR for a first query term with an SDR for a second query term. The resulting composite SDR more accurately reflects the types of queries the user makes, and the more term SDRs that can be added to the composite SDR over time, the more accurately the composite SDR reflects the user's area of ​​expertise.

[0137] The method 1050 includes receiving, by the user expertise profile module, a request from the fourth computing device for identities of users associated with the second term and similar terms (1070). In some embodiments, the request for identities of users associated with similar data items is explicit. In other embodiments, the user expertise profile module 1010 automatically provides such identities as a service to the user of the fourth computing device. For example, a user of the fourth computing device searching for documents similar to a query term in a white paper the user is creating can request (or be given the option to receive) identities of other users who have developed expertise in topics similar to the selected query term. For example, this functionality allows a user to identify users who have developed expertise in a particular topic, regardless of whether such expertise is part of the user's official title, job description, or role. This makes available information that would previously be difficult to identify based solely on official sources, word of mouth, or personal connections. Multiple areas of expertise (e.g., multiple SDRs based on one or more query terms) may be relevant to a single user, providing information about not only the first area of ​​expertise but also the second area of ​​expertise. For example, an individual may formally focus on a first area of ​​research, but may make a series of queries over the course of a week that explore the potential expansion of work into a second area of ​​research. Even expertise gained in such a limited time may be useful to other users. As another example, an individual seeking to build a team or build (rebuild) an organization based on actual areas of interest may use the functionality of the user expertise profile module 1010 to identify users with expertise relevant to the individual's needs.

[0138] The method 1050 includes determining 1072, with a similarity engine, a level of semantic similarity between the SDR of the at least one term and the SDR of the second term. In one embodiment, the similarity engine 304 executes on the second machine 102b. In another embodiment, the similarity engine 304 is provided by and executes within the search system 902. Upon receiving a query term from a user seeking to identify individuals with an area of ​​expertise, the user expertise profile module 101 may direct the similarity engine 304 to identify other users from the user expertise SDR database 1012 who fulfill the request.

[0139] The method 1050 includes providing, by the user expertise profile module, to the fourth computing device an identifier of the user of the third computing device (1074).

[0140] In some embodiments, users of the methods and systems described herein may identify preferences for query terms. As an example, a first user conducting a search on a query term may be interested in documents related to legal aspects of the query term (e.g., use of terms such as the query term in court cases, patent applications, published licenses, or other legal documents), while a second user conducting a search on the same query term may be interested in documents related to scientific aspects of the query term (e.g., use of the query term or use of terms such as the query term in white papers, research publications, grant applications, or other scientific documents). In some embodiments, the systems described herein provide the ability to identify such preferences and rank search results according to which search results most closely match the searcher's preferred document types (based on SDR analysis).

[0141] 10A and 10B, block diagrams illustrate embodiments of a system for semantically ranking query results received from an enterprise search system based on user preferences. FIG. 10A illustrates an embodiment in which semantic ranking functionality is provided in conjunction with results from a traditional enterprise search system. FIG. 10B illustrates an embodiment in which semantic ranking functionality is provided in conjunction with results from search system 902.

[0142] Referring to FIG. 10D , a flowchart illustrates one embodiment of a method 1080 for semantically ranking query results received from a full-text search system based on a user profile. The method 1080 includes clustering a set of data documents selected along at least one criterion in a two-dimensional distance space by a reference map generator running on a first computing device to generate a semantic map (1081). The method 1080 includes associating coordinate pairs with each of the set of data documents through the semantic map (1082). The method 1080 includes generating a list of terms present in the set of data documents by a parser running on the first computing device (1083). For each term in the list, the method 1080 includes determining presence information (including (i) through (iii) below) by an expression generator running on the first computing device (1084): (i) the number of data documents in which the term occurs, (ii) the number of times the term occurs in each data document, and (iii) the coordinate pairs associated with each data document in which the term occurs. The method 1080 includes generating, by a representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information (1085). The method 1080 includes storing each generated SDR in an SDR database (1086). The method 1080 includes receiving, by a query module executing on a second computing device, from a third computing device (1087), a first term and a plurality of preferred documents. The method 1080 includes generating, by the representation generator, a composite SDR using the plurality of preferred documents (1088). The method 1080 includes sending, by the query module, a query of identification information for each of a set of result documents similar to the first term to a full-text search system (1089). The method 1080 includes generating, by the representation generator, an SDR for each document identified in the set of result documents (1090). The method 1080 includes determining 1091, with a similarity engine, a level of semantic similarity between each SDR generated for each of the set of result documents and the composite SDR.The method 1080 includes reordering at least one document in the set of result documents by a ranking module executing on a second computing device based on the determined semantic similarity level (1092). The method 1080 includes providing identification information for each of the set of result documents in the reordered order by a query module to a third computing device (1093).

[0143] In one embodiment, steps 1081 through 1086 are performed as described above with respect to FIG. 2 and steps 202 through 214.

[0144] The method 1080 includes receiving 1087, by a query module executing on the second computing device, from the third computing device a first term and a plurality of preferred documents. In one embodiment, the query input processing module 607 receives the first term as described above with respect to FIGS. 6A, 6B, 9A, and 9B. In another embodiment, the query input processing module 607 provides a user interface element (not shown) that enables a user of the third computing device to provide (e.g., upload) one or more preferred documents. The preferred documents may be any type or form of data structure that includes one or more data items that indicate the types of documents that are of interest to the searching user. As an example, a scientific researcher may provide a number of research documents that reflect the style and / or content of the types of documents that he or she considers relevant or preferable for a given search purpose. As another example, a lawyer may provide a number of legal documents that reflect the style and / or content of the types of documents that he or she considers relevant or preferable for a given search purpose. Additionally, the system provides functionality that allows users to perform different searches against different sets of preferred documents, or to create different preference profiles for use with different searches at different times. For example, a different preference profile might be more relevant to a scientific search focused on a first topic of research and less relevant to a scientific search focused on a second, different topic.

[0145] The method 1080 includes generating 1088, by a representation generator, a composite SDR using the plurality of preferred documents. In one embodiment, the user preference module 1004 directs the generation of the composite SDR. For example, the user preference module 1004 sends the preferred documents to the fingerprint module 302 to generate the composite SDR. As another example, the user preference module 1004 sends the preferred documents to the representation generator 114 to generate the composite SDR. A composite SDR that combines the SDRs of individual preferred documents may be generated in the same manner as composite SDRs of individual documents are generated from term SDRs. The user preference module 1004 may store the generated composite SDR in the user preferred SDR database 1006.

[0146] The method 1080 includes sending 1089, by a query module, a query of identities of each of the set of result documents similar to the first term to a full-text search system. The query module 601 may send the query to an external enterprise search system, as described in connection with Figures 6A and 6B. Alternatively, the query module 601 may send the query to a search system 902, as described in connection with Figures 9A and 9B.

[0147] The method 1080 includes generating 1090 an SDR for each document identified in the set of result documents using a representation generator. In one embodiment, the user preference module 1004 receives the set of result documents from a search system (the search system 902 or a third-party enterprise search system). In another embodiment, the user preference module 1004 directs the similarity engine 304 to generate an SDR for each received result document.

[0148] The method 1080 includes determining 1091, with a similarity engine, a semantic similarity level between each SDR generated for each of the set of result documents and the composite SDR. In one embodiment, the similarity engine 304 executes on the second machine 102b. In another embodiment, the similarity engine 304 is provided by and executes within the search system 902. In one embodiment, the user preference module 1004 instructs the similarity engine 304 to identify the similarity level. In another embodiment, the user preference module 1004 receives the similarity level from the similarity engine 304.

[0149] The method 1080 includes reordering at least one document in the set of result documents by a ranking module executing on a second computing device based on the determined semantic similarity level (1092). In one embodiment, by way of example and not limitation, the similarity engine 304 may indicate that a result included as the fifth document in the set of result documents has a higher level of similarity to the composite SDR of the multiple preferred documents than the first through fourth documents. The user preference module 1004 may then move the fifth document (or an identification of the fifth document) to the first position.

[0150] The method 1080 includes providing, by the query module, identification information for each of the set of result documents in the permuted order to a third computing device (1093). In one embodiment, by analyzing the search results relative to the preferred documents, the system may personalize the search results by considering the context of the search and selecting search results that are likely to be most relevant to the searcher. As another example, instead of returning a traditional, ranked number of results (e.g., the first ten, or the first page, or any other number of results), the system may analyze thousands of documents and provide only semantically relevant documents to the searcher.

[0151] In some embodiments, a patient's symptoms of illness may appear very early, allowing a medical professional to make a definitive medical diagnosis. However, in other embodiments, a patient may only exhibit some of the symptoms, and a definitive medical diagnosis cannot yet be made. For example, a patient may provide a blood sample in which ten different measurements are measured, with only one showing a pathological value and the other nine measurements being within the normal range, albeit near the threshold. In such cases, making a definitive medical diagnosis may be difficult, and the patient may undergo further testing, additional monitoring, or a delayed diagnosis while the medical professional waits to see if the remaining symptoms develop. In such instances, the inability to make an early diagnosis may result in delayed treatment, adversely affecting the patient's medical outcomes. Some embodiments of the methods and systems described herein address such embodiments and provide functionality to support medical diagnosis.

[0152] As noted above, the systems described herein may generate and store SDRs for items of numeric data as well as for text-based items, and then identify a level of similarity between the generated SDR for a received document and one of the stored SDRs. In some embodiments, if the received document is associated with other data or metadata (e.g., a medical diagnosis), the system may provide an identification of that data or metadata as a result of identifying the level of similarity (e.g., identifying a medical diagnosis associated with a document that includes an item of numeric data).

[0153] Referring to FIG. 11B in conjunction with FIG. 11A , a flowchart illustrates one embodiment of a method 1150 for providing medical diagnosis support. The method 1150 includes clustering a set of data documents associated with a medical diagnosis, selected along at least one criterion, in a two-dimensional distance space by a reference map generator executing on a first computing device to generate a semantic map (1152). The method 1150 includes associating coordinate pairs with each of the set of data documents through the semantic map (1154). The method 1150 includes generating a list of measurements present in the set of data documents by a parser executing on the first computing device (1156). The method 1150 includes determining, by a representation generator executing on the first computing device, presence information for each measurement in the list, including: (i) the number of data documents in which the measurement occurs; (ii) the number of times the measurement occurs in each data document; and (iii) coordinate pairs associated with each data document in which the measurement occurs. Method 1150 includes generating 1160, by a representation generator, a sparse distributed representation (SDR) for each of the listed measurements using the presence information. Method 1150 also includes storing 1162 each of the generated SDRs in an SDR database. Method 1150 also includes receiving 1164, by a diagnostic support module executing on a second computing device, from a third computing device, a document including a plurality of measurements and associated with a diseased patient. Method 1150 also includes generating 1166, by the representation generator, at least one SDR for the plurality of measurements. Method 1150 also includes generating 1168, by the representation generator, a composite SDR for the document based on the at least one SDR generated for the plurality of measurements. Method 1150 also includes determining 1170, by a similarity engine executing on the second computing device, a semantic similarity level between the composite SDR generated for the document and an SDR retrieved from the SDR database. The method 1150 includes providing 1172, by a diagnostic assistance module, to a third computing device, identification information of a medical diagnosis associated with the SDR retrieved from the SDR database based on the determined semantic similarity level.

[0154] The method 1150 includes clustering 1152 a set of data documents selected along at least one criterion and associated with a medical diagnosis in a two-dimensional distance space using a reference map generator running on a first computing device to generate a semantic map. In one embodiment, the clustering is performed as described above in connection with FIG. 2. In some embodiments, each document in the set of documents includes multiple data items, as described above. In one of these embodiments, however, the multiple data items are a set of laboratory values ​​taken from a single sample at a time (e.g., a blood sample from a medical patient). As an example, the multiple data items in a document may be provided as a comma-separated list of numeric values. As an example, the system may receive 500 documents (one document for each of the 500 patients), each document including five measurements (five values ​​of one measurement from a single blood sample provided by each patient) associated with a medical diagnosis. The system may use the measurements as data items to generate document vectors as described above in connection with FIG. 2. In one embodiment, the system of Figure 11A may have the functionality described in connection with Figures 1A-1C and 3. However, the system of Figure 11A may have a different parser 110 (shown as experimental document parser and preprocessing module 110b) that is optimized for parsing documents containing test values, and the system may have a binning module 150 that optimizes the generation of a list of measurements present in a set of data documents, as described in more detail below.

[0155] Method 1150 includes associating 1154 coordinate pairs with each of the set of data documents through a semantic map. In one embodiment, the generation of semantic map 108, the distribution of document vectors to semantic map 108, and the association of coordinate pairs are performed as described above in connection with FIG. 2. By way of example, and without limitation, each point in semantic map 108 may represent one or more documents containing multiple lab values ​​for a single type of measurement, including, but not limited to, any type of measurement identified in a metabolic panel (e.g., calcium per liter). While the specific examples described herein refer to lab values ​​derived from blood tests, one skilled in the art will recognize that any type of medical data associated with a medical diagnosis may be used with the methods and systems described herein.

[0156] The method 1150 includes generating 1156 a list of measurements present in the set of data documents using a parser running on a first computing device. In one embodiment, the measurements are listed as described above in connection with FIG. 2. However, in some embodiments, the system includes a binning module 150 that provides an optimized process for generating the list. Each received document may contain multiple values, each value identifying a measurement type. For example, a document may contain a blood calcium level value. The value is the numeric value in the document, and "calcium" is the measurement type. However, the values ​​of each type may vary from document to document. For example, in a set of 500 documents, the measurements of "calcium" may range from 0.0 to 5.2 mg / liter. This is not limiting. When dealing with text-based documents, if multiple documents each contain a word, the word is the same in each document—for example, if two documents contain the word "fast," the text forming the word "fast" is the same in each document. In contrast, when dealing with laboratory values, two documents may each contain values ​​for the same type of measurement (e.g., a "calcium" measurement or a "glucose" measurement) but show very different values ​​(e.g., 0.1 and 5.2). Both values ​​are valid for that type of measurement. Therefore, to optimize the system, the system provides the user with the ability to identify a range of values ​​for each type of measurement contained in a set of documents and divide that range into substantially equal subgroups. Such a process is sometimes referred to as binning. Binning groups together significant overlaps in measurements. As an example, the system may indicate that there are 5,000 values ​​for "calcium" measurements in a set of documents, with the range of values ​​ranging from 0.01 to 5.2, and give the user the option to specify how to distribute the values. For example, a user may specify that values ​​from 0.01 to 0.3 are grouped into a first subdivision (also referred to herein as a "bin"), values ​​from 0.3 to 3.1 are grouped into a second subdivision, and values ​​from 3.1 to 5.2 are grouped into a third subdivision.The system may then tabulate how many of the 5000 values ​​fall into the three bins and use the presence information to generate an SDR for each value. Binning module 150 may provide this functionality.

[0157] The method 1150 includes determining (1158) for each listed measurement, by a representation generator executing on the first computing device, presence information including: (i) the number of data documents in which the measurement occurs; (ii) the number of times the measurement occurs in each data document; and (iii) coordinate pairs associated with each data document in which the measurement occurs. In one embodiment, the presence information is the information described above in connection with FIG. 2.

[0158] The method 1150 includes generating 1160, by a representation generator, a sparse distributed representation (SDR) for each measurement in the list using the presence information. In one embodiment, the SDR is generated as described above in connection with FIG. 2.

[0159] The method 1150 includes storing each of the generated SDRs in an SDR database 1162. In one embodiment, the generated SDRs are stored in the SDR database 120, as described above in connection with FIG.

[0160] The method 1150 includes receiving 1164, by a diagnostic assistance module executing on the second computing device, from a third computing device a document that includes the plurality of measurements and that is associated with the ill patient. In one embodiment, the diagnostic assistance module 1100 receives the document from the client 102c.

[0161] Method 1150 includes generating 1166, by a representation generator, at least one SDR for the plurality of measurements. In one embodiment, diagnostic assistance module 1100 directs fingerprint module 302 to generate an SDR as described above with respect to Figures 1-3. In one embodiment, diagnostic assistance module 1100 directs representation generator 114 to generate an SDR as described above with respect to Figures 1-3.

[0162] The method 1150 includes generating 1168, by a representation generator, a composite SDR for the document based on at least one SDR generated for the plurality of measurements. In one embodiment, the diagnostic assistance module 1100 directs the fingerprint module 302 to generate the composite SDR as described above with respect to Figures 1-3. In one embodiment, the diagnostic assistance module 1100 directs the representation generator 114 to generate the composite SDR as described above with respect to Figures 1-3.

[0163] The method 1150 includes determining 1170, with a similarity engine executing on a second computing device, a level of semantic similarity between the composite SDR generated for the document and the SDR retrieved from the SDR database. In one embodiment, the diagnostic assistance module 1100 directs the similarity engine 304 to determine the level of semantic similarity as described above in connection with Figures 3-5.

[0164] The method 1150 includes providing (1172) to a third computing device, by a diagnostic support module, the identification of a medical diagnosis associated with the SDR retrieved from the SDR database based on the determined semantic similarity level. Such a system can detect an impending medical diagnosis even if individual measurements have not yet reached a pathological level. By inputting multiple SDRs and analyzing patterns among them, the system can identify changes in a patient's patterns, thereby capturing even dynamic processes. For example, a pre-cancer detection system may identify small changes in some values, but with the ability to compare the patterns with other patients' SDRs and analyze time-based sequences, a medical diagnosis can be made.

[0165] In one embodiment, the diagnostic assistance module 1100 can direct the generation of an SDR for an incomplete parameter vector, e.g., when the diagnostic assistance module 1100 receives multiple measurements for a document and the multiple measurements are missing a type of measurement relevant to the diagnosis, without compromising the quality of the results. For example, as described above, even if two SDRs are not identical, the two SDRs can be compared to identify a similarity level that meets a threshold level of similarity. Thus, even if an SDR generated for a document with incomplete measurements is missing one or two locations (e.g., a location in the semantic map 108 where a value for the measurement would be present in a more complete document), the comparison can still be performed using the stored SDR. In such an embodiment, the diagnostic assistance module 1100 can identify at least one parameter relevant to the medical diagnosis for which a value has not been received and can recommend providing that value (e.g., suggesting a subsequent procedure or analysis for the missing parameter).

[0166] In some embodiments, a received document may include associations to metadata in addition to a medical diagnosis. For example, a document may also be associated with an identification of the patient's gender. Such metadata may be used to verify the level of similarity between two SDRs and an identified medical diagnosis. By way of example, the diagnostic assistance module 1100 may determine that two SDRs are similar and identify a medical diagnosis associated with one of the two SDRs from which the document was generated. The diagnostic assistance module 1100 may then apply rules based on the metadata to verify the accuracy of the identification of the medical diagnosis. By way of example, and not by way of limitation, if the metadata indicates that the patient is male and the identified medical diagnosis indicates a risk of ovarian cancer, the diagnostic assistance module 1100 may use rules to provide that, rather than providing the identified medical diagnosis to the user at client 102c, an error (men cannot have ovarian cancer because they do not have ovaries) is reported.

[0167] 13, 14A, and 14B, various embodiments of methods and systems for generating and using cross-language sparse distributed representations are disclosed. In some embodiments, a system 1300 may receive a translation of some or all of a set of documents from a first language to a second language and use the translation to identify corresponding SDRs in a second SDR database generated from a corpus of translated documents. Generally, the system 1300 includes an engine 101, which includes a second representation generator 114b, a second parser and preprocessing module 110c, a set of translated data documents 104b, a second full-text search system 122b, a second list of data items 112b, and a second SDR database 120b. The engine 101 may be the engine 101 described above in connection with FIG. 1A.

[0168] 14A , method 1400 includes clustering a set of data documents in a first language, selected according to at least one criterion, in a two-dimensional metric space by a reference map generator running on a first computing device to generate a semantic map (1402). Method 1400 also includes associating coordinate pairs with each of the set of data documents through the semantic map (1404). Method 1400 also includes generating a list of terms present in the set of data documents by a first parser running on the first computing device (1406). For each term in the list, method 1400 also includes determining (1408) presence information (including (i) through (iii) below) by a first representation generator running on the first computing device: (i) the number of data documents in which the term occurs; (ii) the number of times the term occurs in each data document; and (iii) the coordinate pairs associated with each data document in which the term occurs. The method 1400 includes generating 1410, by a first representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information. The method 1400 includes storing 1412, by the first representation generator, each of the generated SDRs in a first SDR database. The method 1400 includes receiving 1414, by a reference map generator, a translation of each of the set of data documents into a second language. The method 1400 includes 1416 associating coordinate pairs from each of the set of data documents with corresponding documents in the translated set of data documents via a semantic map. The method 1400 includes generating 1418, by a second parser, a second list of terms present in the translated set of data documents. The method 1400 includes, for each term in the second list based on the set of translated data documents, determining (1420) by a second representation generator presence information including: (i) the number of translated data documents in which the term occurs; (ii) the number of occurrences of the term in each translated data document; and (iii) a coordinate pair associated with each translated data document in which the term occurs.The method 1400 includes generating an SDR for each term in the second list based on the translated set of data documents by a second representation generator (1422). The method 1400 includes storing each of the generated SDRs for each term in the second list in a second SDR database by a second representation generator. The method 1400 includes generating a first SDR for the first document in a first language by a first representation generator (1426). The method 1400 includes generating a second SDR for the second document in a second language by a second representation generator (1428). The method 1400 includes determining a distance between the first SDR and the second SDR (1430). The method 1400 includes providing an identification of a similarity level between the first document and the second document (1432).

[0169] In one embodiment, steps 1402-1412 are performed as previously described in connection with FIG. 2 and steps 202-214.

[0170] Method 1400 includes receiving 1414, by a reference map generator, a translation into the second language of each of the set of data documents. In one embodiment, the translation is provided to reference map generator 106 by a translation process performed by machine 102A. In another embodiment, the translation is provided to engine 101 by a human translator. In yet another embodiment, the translation is provided to engine 101 by a machine translation process. The machine translation process may be performed by a third party, or the machine translation process may provide the translation to engine 101 directly or over a network. In yet another embodiment, the translation is uploaded to machine 102A by a user of system 1300.

[0171] Method 1400 includes associating 1416 coordinate pairs from each of the set of data documents with each corresponding document in the translated set of data documents via a semantic map. In one embodiment, the association is performed via semantic map 108. In another embodiment, the association is performed as described above in connection with FIG. 2 and (204).

[0172] The method 1400 includes generating 1418, by a second parser, a second list of terms present in the translated set of data documents. In one embodiment, the generating is performed as described above in connection with FIG. 2 and (206). In another embodiment, the second parser is configured (e.g., includes a configuration file) to optimize the second parser 110c for parsing documents in the second language.

[0173] The method 1400 includes determining (1420) by a second representation generator, for each term in the second list based on the set of translated data documents, presence information including: (i) the number of translated data documents in which the term occurs; (ii) the number of times the term occurs in each translated data document; and (iii) coordinate pairs associated with each translated data document in which the term occurs. In one embodiment, determining the presence information is performed as described above in connection with FIG. 2 and (208).

[0174] The method 1400 includes generating 1422, for each term in the second list, an SDR based on the translated set of data documents with a second representation generator. In one embodiment, generating the term SDRs is performed as described above in connection with Figure 2 and (210)-(214).

[0175] Method 1400 includes storing 1424, by a second representation generator, each of the SDRs generated for each term in the second list in a second SDR database. In one embodiment, storing the SDRs in the second database is performed as described above in connection with FIG. 1A.

[0176] The method 1400 includes generating 1426, by a first representation generator, a first SDR of the first document in the first language. In one embodiment, generating the first SDR is performed as described above in connection with FIG.

[0177] The method 1400 includes generating 1428, by a second representation generator, a second SDR of the second document in the second language. In one embodiment, generating the second SDR is performed as described above in connection with FIG.

[0178] The method 1400 includes determining 1430 a distance between the first SDR and the second SDR. The method 1400 includes determining 1432 a similarity level between the first document and the second document. In one embodiment, steps 1430-1432 are performed as described above with respect to Figures 3-4.

[0179] In one embodiment, the methods and systems described herein may be used to measure the quality of a translation system. For example, a translation system may translate text from a first language to a second language and provide both the first language text and the second language translation to the system described herein. If the system determines that the SDR of the first language text is similar to the SDR of the translated text (second language) (e.g., above a threshold level of similarity), the quality of the translation may be high. Continuing with the example above, if the system determines that the SDR of the first language text is not sufficiently similar to the SDR of the translated text (second language) (e.g., not above a predetermined threshold level of similarity), the quality of the translation may be low.

[0180] Referring now to the flowchart of FIG. 14B in conjunction with FIG. 13 and FIG. 14A, one embodiment of a method 1450 is disclosed. Generally, in FIG. 14B, the method 1450 includes clustering a set of data documents in a first language, selected according to at least one criterion, in a two-dimensional metric space to generate a semantic map (1452) by a reference map generator executing on a first computing device. The method 1450 also includes associating coordinate pairs with each of the set of data documents through the semantic map (1454). The method 1450 also includes generating, by a first parser executing on the first computing device, a list of terms present in the set of data documents (1456). The method 1450 also includes determining, for each term in the list, presence information (including (i) through (iii) below) by a first representation generator executing on the first computing device (1458). (i) the number of data documents in which the term occurs, (ii) the number of times the term occurs in each data document, and (iii) coordinate pairs associated with each data document in which the term occurs. Method 1450 includes generating 1460, by a first representation generator, a sparse distributed representation (SDR) for each term in the list using the presence information. Method 1450 includes storing 1462, by the first representation generator, each of the generated SDRs in a first SDR database. Method 1450 includes receiving 1464, by a reference map generator, a translation of each of the set of data documents into a second language. Method 1450 includes associating 1466, by a semantic map, a coordinate pair from each of the set of data documents with each of the translated set of data documents. Method 1450 includes generating 1468, by a second parser, a second list of terms present in the translated set of data documents. The method 1450 includes, for each term in the second list based on the set of translated data documents, determining (1470) by a second representation generator presence information including: (i) the number of translated data documents in which the term occurs; (ii) the number of occurrences of the term in each translated data document; and (iii) a coordinate pair associated with each translated data document in which the term occurs.The method 1450 includes generating an SDR for each term in the second list based on the translated set of data documents by a second representation generator (1472). The method 1450 includes storing each of the generated SDRs for each term in the second list in a second SDR database by the second representation generator (1474). The method 1450 includes generating a first SDR for a first term received in a first language by the first representation generator (1476). The method 1450 includes determining a distance between the first SDR and a second SDR for a second term in a second language retrieved from a second SDR database (1478). The method 1450 includes providing an identification of the second term in the second language and an indication of a similarity level between the first term and the second term based on the determined distance (1480).

[0181] In one embodiment, steps 1452-1474 are performed as described above in connection with Figure 14A (steps 1402-1424).

[0182] The method 1450 includes generating 1476, with a first representation generator, a first SDR of the received first term in the first language. In one embodiment, generating the first SDR is performed as described above in connection with FIG.

[0183] Method 1450 includes determining 1478 a distance between the first SDR and a second SDR of a second term in a second language retrieved from a second SDR database. Method 1450 includes providing 1480 an identification of the second term in the second language and an indication of a similarity level between the first term and the second term based on the determined distance. In one embodiment, steps 1478-1478 are performed as described above in connection with FIGS. 3 and 4.

[0184] In other embodiments, the methods and systems described herein may be used to enhance a search system. For example, system 1300 may receive a first term in a first language (e.g., a term a user desires to use in a query on the search system). System 1300 may generate an SDR for the first term and use the first SDR to identify a second SDR in a second SDR database that meets a threshold similarity level. System 1300 may then provide the first SDR, the second SDR, or both to the search system, as described above with reference to FIGS. 6A-6C, to enhance the user's search query.

[0185] In some embodiments, the methods and systems described herein may be used to provide streaming data filtering capabilities. For example, an entity may desire to review streaming social media data to identify sub-streams of social media data associated with the entity (e.g., for brand management purposes or competitive monitoring). As another example, an entity may desire to review streaming network packets traversing a network device, e.g., for security purposes.

[0186] Referring now to Figure 16 in conjunction with Figure 15, a system 1500 provides functionality for performing a method 1600 for identifying similarities between filtering criteria and data items in a set of streaming data documents. The system 1500 includes an engine 101, a fingerprint module 302, a similarity engine 304, a disambiguation module 306, a data item module 308, an expression engine 310, an SDR database 120, a filtering module 1502, a reference SDR database 1520, streaming data documents 1504, and a client agent 1510. The engine 101, the fingerprint module 302, the similarity engine 304, the disambiguation module 306, the data item module 308, the expression engine 310, and the SDR database 120 may be provided as previously described in conjunction with Figures 1A through 14.

[0187] The method 1600 includes a step 1602 of clustering a set of data documents selected along at least one criterion in a two-dimensional metric space by a reference map generator running on a first computing device to generate a semantic map. The method 1600 also includes a step 1604 of associating coordinate pairs with each of the set of data documents by the semantic map. The method 1600 also includes a step 1606 of generating a list of terms present in the set of data documents by a parser running on the first computing device. For each term in the list, the method 1600 also includes a step 1610 of determining presence information (including (i) to (iii) below) by a representation generator running on the first computing device, including: (i) the number of data documents in which the term occurs; (ii) the number of times the term occurs in each data document; and (iii) coordinate pairs associated with each data document in which the term occurs. The method 1600 also includes a step 1610 of generating a sparse distributed representation (SDR) for each term in the list using the presence information by the representation generator. The method 1600 includes storing (1612) each of the generated SDRs in an SDR database. The method 1600 includes receiving (1614) filtering criteria from a third computing device by a filtering module executing on a second computing device. The method 1600 includes generating (1616) at least one SDR for the filtering criteria by a representation generator. The method 1600 includes receiving (1618) a plurality of stream documents from a data source by the filtering module. The method 1600 includes generating (1620) a composite SDR for a first document of the plurality of stream documents by the representation generator for a first document of the plurality of stream documents. The method 1600 includes determining (1622) a distance between the filtering criteria SDR and the generated composite SDR for the first document of the plurality of stream documents by a similarity engine executing on the second computing device. The method 1600 includes processing 1624 the first stream document with a filtering module based on the determined distance.

[0188] In one embodiment, steps 1602-1612 are performed as described in connection with FIG. 2, steps 202-214.

[0189] The method 1600 includes receiving (1614) filtering criteria from the third computing device by a filtering module executing on the second computing device. The filtering criteria may be any term by which the filtering module 1502 can narrow down the plurality of stream documents. By way of example, as noted above, an entity may desire to review streaming social media data to identify substreams of social media data related to the entity (e.g., for brand management purposes or competitive monitoring). As another example, an entity may desire to review streaming network packets traversing a network device, e.g., for security purposes. Thus, in one embodiment, the filtering module 1502 receives at least one brand-related term. For example, the filtering module 1502 may receive a name, such as a company, product, or individual name (for an entity related to the third machine or for an entity not associated with the third machine, such as a competitor). In another embodiment, the filtering module 1502 receives security-related terms. For example, the filtering module 1502 receives terms related to computer security exploits (e.g., terms related to hacking, malware, or other exploits of security vulnerabilities) or physical security exploits (e.g., terms related to acts of violence or terrorism). In yet another embodiment, the filtering module 1502 receives at least one virus signature (e.g., a computer virus signature, as will be understood by those skilled in the art).

[0190] In some embodiments, the filtering module 1502 receives at least one SDR. For example, a user of the machine 102c may have already been able to interact with the system 1500 for a particular purpose and developed one or more SDRs that can be used in connection with filtering the browsing data.

[0191] In some embodiments, the filtering module 1502 communicates with the query expansion module 603 to identify additional filtering criteria (e.g., as described above with respect to FIGS. 6A-6C). For example, the filtering module 1502 can send the filtering criteria to the query expansion module 603 (not shown, but running on machine 102b or another machine 102g). The query expansion module 603 can instruct the similarity engine 304 to identify a level of semantic similarity between the first SDR of the filtering criteria and a second SDR of a second term retrieved from the SDR database 120. In such an example, the query expansion module 603 repeats the identification process for each term in the SDR database 120 and instructs the similarity engine 304 to return terms having a level of semantic similarity above a threshold. The query expansion module 603 can provide the resulting terms identified by the similarity engine 304 to the filtering module 1502, which filters them. The filtering module 1502 can then use the resulting terms when filtering a streaming set of documents.

[0192] The method 1600 includes generating (1616) at least one SDR for the filtering criteria by the representation generator. In one embodiment, the filtering module 1502 sets the filtering criteria to the engine 101 for generating the at least one SDR by the representation generator 114. In another embodiment, the filtering module 1502 sets the filtering criteria to the fingerprint module SDR. The filtering module 1502 can store the at least one SDR in a reference SDR database 1520.

[0193] In some embodiments, generating at least one SDR is optional. In one embodiment, the representation generator 114 (or fingerprint module 302) determines whether the received filtering criteria is or includes an SDR and determines whether to generate an SDR based on that determination. For example, the representation generator 114 (or fingerprint module 302) may determine that the filtering criteria received by the filtering module 1502 are an SDR and therefore decide not to generate another SDR. Alternatively, the representation generator 114 (or fingerprint module 302) may determine that the SDR for the filtering criteria already exists in the SDR database 120 or the reference SDR database 1520. However, as another example, the representation generator 114 (or fingerprint module 302) may determine that the filtering criteria is not an SDR and generate an SDR based on that determination.

[0194] The method 1600 includes receiving (1618), by the filtering module, a plurality of stream documents from a data source. In one embodiment, the filtering module 1502 receives a plurality of social media text documents, such as documents of any length or type generated within computer-mediated tools that allow users to create, share, or exchange any type of data (audio, video, and / or text-based). Examples of such social media include, but are not limited to, blogs; wikis; consumer review sites such as YELP, offered by Yelp, Inc. (San Francisco, CA); microblogging sites such as TWITTER®, offered by Twitter, Inc. (San Francisco, CA); and combination microblogging and social networking sites such as FACEBOOK®, offered by Facebook, Inc. (Menlo Park, CA), or GOOGLE+, offered by Google, Inc. (Mountain View, CA). In another embodiment, the filtering module 1502 receives a plurality of network traffic documents. For example, the filtering module 1502 can receive a plurality of network packets, each of which can be referred to as a document.

[0195] In one embodiment, the filtering module 1502 receives identification information of a data source having filtering criteria from a third computing device. In another embodiment, the filtering module 1502 utilizes an application programming interface provided by the data source to initiate reception of the plurality of stream documents. In yet another embodiment, the filtering module receives the plurality of stream documents from a third machine 102c. As an example, the data source may be a third-party data source, and the filtering module 1502 is programmed to contact the third-party data source and initiate reception of the plurality of stream documents. For example, the third party may provide a social media platform and stream documents that are regenerated and downloadable on the platform. In another example, the data source may be provided by the third machine 102c, and the filtering module 1502 may directly retrieve the stream documents from the third machine 102c. For example, the machine 102c may be a router that retrieves network packets from another machine on the network 104 (not shown). As described in more detail below, the filtering module can receive multiple stream documents from one or more data sources and compare them to each other, to a reference SDR, or to an SDR retrieved from the SDR database 120 .

[0196] The method 1600 includes generating (1620), by the representation generator, for a first document of the plurality of stream documents, a composite SDR for the first document of the plurality of stream documents. The filtering module 1502 may provide the first document of the plurality of stream documents directly to the representation generator 114. Alternatively, the filtering module 1502 may provide the first document of the plurality of stream documents to the fingerprint module 302. The composite SDR may be generated as described above in connection with FIG. 2. In some embodiments, the representation generator 114 (or the fingerprint module 302) generates the composite SDR for the first document of the plurality of stream documents before receiving the second document of the plurality of stream documents.

[0197] The method 1600 includes determining (1622) by a similarity engine executing on a second computing device a distance between the filtered reference SDR and the generated composite SDR for a first document of the plurality of stream documents. The filtering module 1502 can provide the filtered reference SDR and the generated composite SDR to the similarity engine 304. Alternatively, the filtering module 1502 can provide an identification of the reference SDR database 1520 to the similarity engine 304 so that the similarity engine 304 can directly retrieve the filtered reference SDR.

[0198] The method 1600 includes processing (1624) the first stream document by a filtering module based on the determined distance. In one embodiment, the filtering module 1502 forwards the stream document to the third computing device 102c. In another embodiment, the filtering module 1502 determines not to forward the stream document to the third computing device 102c. In yet another embodiment, the filtering module 1502 determines whether to send an alert to the third computing device based on the determined distance. In yet another embodiment, the filtering module 1502 determines whether to send an alert to the third computing device based on the determined distance and the filtering criteria. For example, if the stream document and the filtering criteria have a similarity level based on the determined distance that exceeds a predetermined threshold, the filtering module 1502 can determine that the stream document contains malicious content (e.g., has an SDR that is substantially similar to the SDR in the case of a virus signature). The filtering module 1502 can access policies, rules, or other instruction sets to determine that in such cases, an alert should be sent to one or more users or machines (e.g., paging a network administrator).

[0199] In one embodiment, the filtering module 1502 forwards the first document of the plurality of stream documents to a client agent 1510 executing on the third machine 102c. The client agent 1510 may execute on a router. The client agent 1510 may execute on any type of network device. The client agent 1510 may execute on a web server. The client agent 1510 may execute in any form or type of manner described herein.

[0200] In one embodiment, the filtering module 1502 adds a first document of the multiple stream documents to a substream of the stream document. In another embodiment, the filtering module 1502 stores the substream in a database (not shown) accessible by the client agent 1510 (e.g., by polling the database or subscribing to update notifications or other mechanisms known to those skilled in the art and then downloading all or part of the substream). In yet another embodiment, the filtering module 1502 responds to a polling request received from the client agent 1510 by sending the substream to the client agent 1510.

[0201] In some embodiments, the filtering module 1502 receives a second plurality of stream documents from a second data source. The filtering module 1502 directs the generation of a composite SDR for a first document of the second plurality of stream documents (e.g., as described above in connection with the generation of a composite SDR for a first document of the first plurality of stream documents). The similarity engine 304 determines a distance between the composite SDR generated for the first document of the second plurality of stream documents and the composite SDR generated for the first document of the first plurality of stream documents. The filtering module 1502 determines whether to forward the second plurality of stream first document to a third computing device based on the determined distance. In one embodiment, the filtering module 1502 may determine whether to forward a first stream document of the second plurality of streaming documents based on determining that the compared SDR is below a predetermined similarity threshold; for example, the filtering module 1502 may determine to discard a first stream document of the second plurality of streaming documents if it is sufficiently different from the first stream document of the first plurality of streaming documents (e.g., below a predetermined similarity threshold), while determining that the first stream document of the second plurality of streaming documents is too similar to the first stream document of the first plurality of streamed documents (e.g., above a predetermined similarity threshold), so that the first stream document of the second plurality of streaming documents may be deemed cumulative, overlapping, or too similar to the first streaming document. In this way, the filtering module 1502 may determine that documents from different data sources (e.g., posted on different social media sites, or posted from different accounts on a single social media site, or included in different network packets) are sufficiently similar to provide an improved sub-stream that surpasses a sub-stream with overlap information with a single document available.

[0202] In some embodiments, steps 1606-1610 are customized to address data documents containing virus signatures. In one of these embodiments, the parser generates a list of virus signatures present in a set of data documents. In another of these embodiments, the representation generator determines presence information for each virus feature in the list, including: (i) the number of data documents in which the virus signature is present, (ii) the number of occurrences of the virus signature in each data document, and (iii) a coordinate pair associated with each data document in which the virus signature is present. In yet another of these embodiments, the representation generator generates an SDR, which may be a composite SDR, for each virus signature in the list. In another embodiment, the system parses each virus signature in the list into multiple subunits (e.g., phrases, sentences, or other portions of the virus signature document) based on a protocol (e.g., a network protocol). In yet another embodiment, the system parses each subunit in the list into at least one value (e.g., a word). In yet another embodiment, the system determines presence information for each value of a plurality of subunits of a virus signature in the inventory, including: (i) the number of data documents in which the value occurs, (ii) the number of occurrences of the value in each data document, and (iii) each data document coordinate pair in which the value occurs; for each value in the inventory, the system generates an SDR using the value's presence information. However, in another embodiment, the system generates a composite SDR for each subunit in the inventory using the SDR(s) for the value. In a further embodiment, the system generates a composite SDR for each virus signature in the SDR based on the generated plurality of subunit SDRs. The plurality of virus signature SDRs, the plurality of subunit SDRs, and the plurality of value SDRs may be stored in the SDR database 120.

[0203] The method 1600 includes generating (1606), by a parser executing on a first computing device, a list of terms present in a set of data documents. The method 1600 includes determining (1608), by a representation generator executing on the first computing device, presence information for each term in the list, including: (i) the number of data documents in which the term occurs, (ii) the number of occurrences of the term in each data document, and (iii) a coordinate pair associated with each data document in which the term occurs. The method 1600 includes generating (1610), by the representation generator, for each term in the list, a sparse distributed representation (SDR) using the presence information.

[0204] In some embodiments, the client agent 1510 invokes the functions of the filtering module 1502, the fingerprint module 302 to generate multiple SDRs, and interacts with the similarity engine 304 to receive a determination of the similarity level between the SDR of the stream document and a reference SDR; the client agent 1510 can make a decision regarding whether to save or discard the stream document based on the similarity level.

[0205] In some embodiments, the elements described herein may perform one or more functions automatically, i.e., without human intervention. For example, system 100 may receive a set of data documents 104 and automatically, without human intervention, perform one or more of the following: preprocessing the data documents, training reference map generator 106, or generating an SDR 118 for each data item in the set of data documents 104. As another example, system 300 may receive at least one data item and automatically, without human intervention, perform one or more of the following: identifying a level of similarity between the received data item and data items in SDR database 120, creating a list of similar data items, or performing other functions described above. As yet another example, system 300 may be part of, or include elements of, a so-called "Internet of Things," in which autonomous entities perform, communicate, and provide functions as described herein. For example, automatic, autonomous processes may generate queries, receive answers from system 300, and provide answers to other users (human, computer, or otherwise). In some cases, this includes a speech-to-text or text-to-speech based interface, whereby, for example, but not limited to, a user may generate voice commands that the interface recognizes and generates computer-processable instructions.

[0206] As described above in connection with FIG. 5 , the similarity engine 304 can receive a first data item from a user, determine a distance between a first SDR of the first data item and a second SDR of a second data item retrieved from the SDR database, and provide an identification of the second data item and an identification of the level of semantic similarity between the first data item and the second data item based on the determined distance. In some embodiments, the first data item is a profile description. For example, the profile may be a profile of an ideal job candidate. As another example, the profile may be a profile of an individual the user is interested in meeting (e.g., for networking, dating, or other relationship-building purposes). In some embodiments, the profile is a free-text description of one or more characteristics of the ideal individual, and the similarity engine 304 searches the SDR database 120 for one or more profiles of the individual, where the SDR generated from the individual's profile overlaps with (e.g., has a minimum distance from) the SDR of the provided ideal candidate profile. In some embodiments, the profile is a profile of products (e.g., goods or services) about which the user is interested in learning more or acquiring. By way of example, a user may provide a text-based description of needed or desired product attributes, and in contrast to systems that recommend products based on the user's previous purchasing history or other user-based attributes, the methods and systems described herein may search an SDR database 120 containing SDRs generated from the product descriptions and provide results where descriptions of the products themselves (and not the user or the user's habits) are semantically similar to the desired need or function.

[0207] In some embodiments, the methods and systems described herein provide functionality for searching for and generating SDRs of web pages (e.g., documents stored on a computer and made available for search by other computers over one or more computer networks according to any number of computer networking protocols) and populating SDR database 120 with the SDRs of the web pages. Thus, in one of these embodiments, the set of data documents is a plurality of web pages retrieved by the system (e.g., a web crawler communicating with the system) or by a user of the system. In another of these embodiments, similarity engine 304 receives a data item including a description of a web search the user wishes to perform and performs functions such as those described above in connection with FIG. 5; similarity engine 304 can receive the data item directly from the user (e.g., via a query module provided by the system such as those described above in connection with FIGS. 6A, 9A-B, and 10A-B) or from a third-party system that forwards user input to similarity engine 304.

[0208] In some embodiments, the methods and systems described herein can benefit from training on a specific document corpus to provide more accurate search results. As an example, a system may be customized to provide improved results when providing fraud detection functionality in a particular topic or area of ​​specialized knowledge or industry knowledge. As another example, a system may be customized to provide improved results when providing forensic analysis. In such embodiments, a user may not have a specific description of the features or attributes they are searching for (unlike, for example, a user seeking job candidates with specific skills or a user seeking to purchase a product with specific features), but they may have one or more documents that are examples of the type of documents they want to find, and these documents may be used as described above in connection with FIG. 5. For example, a user may have examples of email messages that trigger financial reporting requirements or ethical violations, and by leveraging semantic fingerprinting, the user can find semantically similar documents even if the specific words used in the emails (or the nature of the documents or communications) vary from scenario to scenario.

[0209] In some embodiments, the similarity engine 304 determines the distance (as described above in connection with FIG. 5 ) between an SDR generated based on a user-provided data item and a previously generated SDR retrieved from the SDR database 120 and determines that this distance exceeds a predetermined threshold. In one of these embodiments, the similarity engine 304 determines that such an SDR is an outlier that may merit further analysis. For example, if a business document contains information that is substantially different from the information contained in other business documents of its type, this can signal to an analyzer (human or computer, or both working together) that the outlier document should undergo further analysis. Continuing with this example, if multiple documents are being analyzed to attempt to identify instances of insider trading, the document author may have used coded words not traditionally used in business documents in that industry; regardless of the specific words used, such documents would generate SDRs that are sufficiently different from conventional documents to be flagged for further analysis. In some embodiments, such additional analysis is performed by the systems described herein or by a human analyst; for example, an investigator may link the creation date of a document to a timeline to determine whether there are other anomalous characteristics of the document. In such instances, use of the systems and methods described herein benefits the investigator by narrowing down the group of documents that may merit additional analysis.

[0210] In some embodiments, the methods and systems described herein can communicate with other third-party artificial intelligence algorithms to provide additional functionality. For example, in providing anomaly detection functionality, an artificial intelligence system can be trained to predict the data item that will follow a data item (e.g., as a result of identifying a pattern in a sequence of data items). Continuing with that example, when an artificial intelligence system is provided with multiple SDRs generated as described above, the artificial intelligence system can identify a pattern within the SDR and determine what should come next within that pattern; if the next SDR breaks that pattern, the system can identify an anomaly. An anomaly can include a new topic in a stream of data; for example, in a stream of data items related to news, an anomaly can indicate breaking news about a different topic.

[0211] In some embodiments, the systems and methods described herein may be used to replace language models when providing functionality supporting machine translation (including speech-to-text translation, optical character recognition, and other uses of machine translation). Conventional systems use language models that can calculate the probability that one language (word, sentence, etc.) is associated with another language (e.g., one word follows another word in a sentence). In one embodiment, a similarity engine 304 may be utilized to replace a language model. The similarity engine 304 may receive a data item (e.g., a word or phrase in a sentence, or a sentence in a paragraph), generate an SDR for the data item, and identify SDRs for words, phrases, or data items that are typically found in association with the received data item. The similarity engine 304 may also receive a document or portion of a document that includes the received word, generate a composite SDR for comparison with composite SDRs of other documents to identify similar documents, and then determine which data items typically follow the received data item. In such an embodiment, the system may also utilize the topic slicing functionality described above.

[0212] In some embodiments, the data item is a managed document (e.g., a document in a system in which at least one item of metadata is associated with the document and the SDR of the document can also be metadata associated with the document). In one of these embodiments, the managed document is a document that is being edited and in which an SDR is generated and updated throughout the time the user is writing or editing the document. In another of these embodiments, the systems and methods described herein provide functionality for providing user feedback while the document is still being written, updated, or otherwise edited. By way of example, the data item may be a managed document at a first point in time, an initial SDR may be generated for the managed document at the first point in time, and at a subsequent point in time (e.g., at a point predetermined by the user or administrator, or when the user requests an update, or when the system is programmed to ask if the user wishes to generate an update), the system may generate an updated SDR. When generating the updated SDR, the similarity engine 304 can compare the SDR to an SDR generated from a previously generated managed document (e.g., an SDR retrieved from the SDR database 120). Based on this comparison, the system can provide feedback to the user creating or modifying the managed document, for example, the system can identify the type of managed document, ask the user if they would like access to other previously generated documents of a similar type (e.g., other letters, other contracts, other documents containing similar keywords, or other documents containing similar sections), and then provide access to the requested documents. The system can also provide other guidance to the user (e.g., reminding them that other managed documents whose SDRs are substantially similar to the SDR of the managed document being created or modified typically contain particular sections or text or attachments).

[0213] In some embodiments, the methods and systems described herein provide functionality for routing documents. In one of these embodiments, the similarity engine 304 receives an SDR of a document to be routed to one of a plurality of users (e.g., an email sent to a specific individual among a plurality of email recipients, or a document to be reviewed by one of a plurality of users) and compares the received SDR with an SDR retrieved from the SDR database 120. In one embodiment, the SDRs of user profiles within the system are entered into the SDR database 120. For example, a user profile associated with a user who reviews tax documents may have a different SDR than a user profile associated with a user who reviews financial documents. By comparing the SDR of the incoming document with the SDRs of the profiled user, the similarity engine 304 can identify an SDR of the user's profile that has a substantial similarity to the SDR of the incoming document. The system can then decide to provide the incoming document to the user identified in the profile. In some embodiments, the SDR database 120 is populated with SDRs of previously routed documents, allowing the system to determine where other documents have been routed. For example, based on analyzing the metadata of documents with similar SDRs to the received SDR, the system can determine that the document being sent should go to a contract attorney or corporate accountant, or an individual responsible for reviewing work by interns.

[0214] Semantic sentiment analysis In some embodiments, the methods and systems described herein provide functionality for performing sentiment analysis. In one of these embodiments, the order of multiple data items during analysis affects the results of the analysis, and the order in which the data items appear varies to determine the sentiment intended by a sentence containing multiple data items (e.g., "man bites dog" vs. "dog bites man"). In the particular embodiment described above, the system generates the same SDR regardless of the word order in the sentence. Thus, to improve the functionality provided in determining whether the SDR of a data item (or group of data items) is substantially similar to the SDR of a data item (or group of data items) that conveys one or more sentiments, the methods and systems described herein may include functionality for interfacing with an artificial intelligence system that provides sequence learning functionality. As will be appreciated by those skilled in the art, a sequence learner can be exposed to a sequence of patterns, predict what the next pattern will be, and present words (or data items) in a particular order. The sequence learning function may include or communicate with a hierarchical temporal memory that identifies an order as a related group of data items (e.g., a sentence) and generates an output SDR representing the sentence in that particular order; if the order of words in the sentence is changed, the hierarchical temporal memory generates a second, different output SDR for the sentence with the words in the changed order. Thus, the similarity engine 304 can receive from the artificial intelligence system an output SDR that reflects the order of data items within the plurality of data items and compare such output SDR with other output SDRs. Furthermore, the artificial intelligence system may be trained based on data associated with particular emotions (e.g., positive, negative, neutral, anger, anxiety, etc.), resulting in a classifier for use in sentiment analysis.

[0215] In some embodiments, the data item includes the inclusion rate of an advertisement. As an example, an advertising company may seek to identify ad placement opportunities for advertisements (e.g., advertisements placed on behalf of customers) and may use the systems and methods described herein to improve the placement of the advertisements. As another more specific example, the systems and methods described herein may improve the placement of advertisements by comparing an SDR of an internet user's shopping content (e.g., the inclusion rate of the internet user's shopping-related cookies, which may include identification of items the user recently searched for or acquired) with a previously generated SDR from an advertisement in a catalog of advertisements available for placement (e.g., an SDR previously generated and stored in SDR database 120). In one of these embodiments, the similarity engine 304 receives the SDR of the shopping content (the data item received in this embodiment) and compares the SDR with the SDR in SDR database 120 to determine which advertisements in the catalog of advertisements are relevant to the user's shopping content. The system can then recommend placing the identified advertisements in locations (e.g., websites) where users with shopping content will see the advertisements. In contrast to conventional systems that can only match keywords to keywords, the use of an affinity engine allows for the rapid identification of related topics, even if the keywords are different. In some embodiments, the methods and systems described herein can provide for the identification of data items having SDRs that are substantially similar to the SDR of shopping content, and can do so quickly enough to satisfy constraints in, for example, an Internet advertising environment where advertisements are identified and placed in milliseconds to prevent delivery from exceeding acceptable time limits (e.g., milliseconds).

[0216] In some embodiments, unlike conventional systems, the systems and methods described herein incorporate semantic context into individual representations. For example, even if it is unknown how a particular SDR was generated, the system can compare that SDR with other SDRs and use the semantic context of the two SDRs to provide insight to the user. In other embodiments, unlike conventional systems that traditionally focus on document-level clustering, the systems and methods described herein use document-level context to provide term-level contextual insight, allowing users to identify the contextual meaning of individual terms within a corpus of documents.

[0217] It should be understood that the above-described systems may include multiples of any of the above elements, or multiples of each of the above elements, and that the above elements may be included on stand-alone machines or, in some embodiments, on multiple machines in the described system. Generally, phrases such as "in one embodiment," "in another embodiment," or the like, mean that the following particular configuration, structure, step, or characteristic is included in at least one embodiment herein, and may be included in multiple embodiments herein. Furthermore, such phrases may refer to the same embodiment, but are not necessarily limited to it.

[0218] Each of the components referred to herein as an engine, generator, module, or element may be implemented as software, hardware, or a combination of the two, and may be executed on one or more machines 100. While components may be shown as separate components in this specification for ease of explanation, it should be understood that this does not limit the components to a particular implementation. For example, some or all of the functionality of the components may be included in a single circuit or software function. As another example, the functionality of one or more components may be distributed across multiple components.

[0219] The machine 102 providing the functionality described herein may be any type of workstation, desktop computer, laptop or notebook computer, server, portable computer, cellular phone, portable smart phone, or other portable communication device, media playback device, gaming system, handheld computing device, or any other type and / or type of computing, communication, or media device capable of communicating over a network and having sufficient processor power and storage capacity to perform the operations described herein. The machine 102 may execute, operate, or provide any type and / or type of software, program, or application, such as executable instructions, including, but not limited to, any type and / or type of web browser, web-based client, client-server application, ActiveX control, JAVA applet, or any other type and / or type of executable instructions executable on the machine 102.

[0220] The machines 100 can communicate with each other over a network, which may be any type and / or variety of network and may include any of the following: a point-to-point network, a broadcast network, a wide area network, a local area network, a telecommunications network, a data communications network, a computer network, an ATM (Asynchronous Transfer Mode) network, a SONET (Synchronous Optical Network) network, an SDH (Synchronous Digital Hierarchy) network, a wireless network, and a wired network. In some embodiments, the network may include a wireless link, such as an infrared channel or a satellite band. The network topology may be a bus network topology, a star network topology, or a ring network topology. The network topology may be any network topology that is compatible with the operations described herein, as known to those skilled in the art. The network may include a cellular network using any one or more protocols used for communication between mobile devices (typically including tabletop or handheld devices), including AMPS, TDMA, CDMA, GSM, GPRS, UMTS, or LTE.

[0221] Machine 102 may include a network interface for connecting to a network via various connections, including, but not limited to, a standard telephone line, an LLAN or WAN link (e.g., 802.11, T1, T3, 56kb, X.25, SNA, DECNET), a broadband connection (e.g., ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethernet-over-SONET), a wireless connection, or a combination of any or all of the above. Connections can be established using a variety of communication protocols (e.g., TCP / IP, IPX, SPX, NetBIOS, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), RS232, IEEE 802.11, IEEE 802.11b, IEEE 802.11gg, IEEE 802.11n, 802.15.4, BLUETOOTH, ZIGBEE, CDMA, GSM, WiMAx, and direct asynchronous connections). In one embodiment, computing devices 100 communicate with other computing devices 100' via any type and / or form of gateway or tunneling protocol, such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS). The network interface may comprise a built-in network adapter, a network interface card, a PCMCIA network card, a card bus network adapter, a wireless network adapter, a USB network adapter, a modem, or any other device suitable for connecting computing device 100 to any type of network capable of communicating and performing the operations described herein.

[0222] The systems and methods described above may be implemented as a method, apparatus, or article of manufacture using programming and / or engineering techniques to produce software, firmware, hardware, or any combination thereof. The techniques may be implemented in one or more computer programs executed on a programmable computer having a processor, a processor-readable storage medium (e.g., including volatile or non-volatile memory and / or storage elements), at least one input device, and at least one output device. Program code may be applied to inputs implemented using the input devices to perform the functions described above and to generate outputs. The outputs may be provided to one or more output devices.

[0223] Each computer program within the scope of the claims below may be implemented in any programming language, such as assembly language, machine language, a high-level procedural programming language, or an object-oriented programming language, such as LISP, PROLOG, PERL, C, C++, C#, JAVA, or any compiled or interpreted programming language.

[0224] Each such computer program may be embodied in a computer program product tangibly embodied in a machine-readable storage device for execution by a computer processor. The steps of the methods of the present invention may be performed by a computer processor executing a program tangibly embodied in a computer-readable medium to process input and generate output to perform the functions of the present invention. Suitable processors include, by way of example, both general-purpose and special-purpose microprocessors. Typically, such processors receive instructions and data from read-only memory and / or random-access memory. Suitable storage devices tangibly embodied with computer program instructions include, for example, all forms of computer-readable devices, firmware, programmable logic, and hardware (e.g., integrated circuit chips, electronic devices, computer-readable non-volatile storage units, non-volatile memory (e.g., semiconductor memory devices) including EPROM, EEPROM, and flash memory devices, magnetic disks such as internal and removable hard disks, magneto-optical disks, and CD-ROMs). Any of the foregoing may be supplemented by or incorporated in specially designed ASICs (application-specific integrated circuits) or FPGAs (field-programmable gate arrays). The computer also typically receives programs and data from a storage medium such as an internal disk (not shown) or a removable disk. These elements are also found in conventional desktop or workstation computers and other computers suitable for executing computer programs that implement the methods described herein. The computer may be used with any digital print or marking engine, display monitor, or other raster output device capable of producing color or grayscale pixels on paper, film, a display screen, or other output medium.The computer may also receive programs and data from a second computer that allows access to the programs via network transmission lines, wireless transmission media, signals propagated over the air, radio waves, infrared signals, etc.

[0225] More specifically, an embodiment of a network environment is disclosed in Figure 12A. Broadly speaking, the network environment includes one or more clients 1202A-1202n in communication with one or more remote machines 1206A-1206n (commonly referred to as servers 1206 or computing devices 1206) via one or more networks 1204. The machine 102 described above may be implemented as machine 1202, machine 1206, or any type of machine 1200.

[0226] While FIG. 12A shows network 1204 between client 1202 and remote machine 1206, client 1202 and remote machine 1206 may be on the same network 1204. Network 1204 may be a local area network (LAN), such as a corporate intranet, a metropolitan area network (MAN), or a wide area network (WAN), such as the Internet or the World Wide Web. In other embodiments, there are multiple networks 1204 between client 1202 and remote machine 1206. In one of these embodiments, network 1204′ (not shown) may be a private network, and network 1204 may be a public network. In another of these embodiments, network 1204 may be a private network, and network 1204′ may be a public network. In yet another embodiment, network 1204 and network 1204′ may both be private networks.

[0227] The network 1204 may be any type and / or form of network, including a point-to-point network, a broadcast network, a wide area network, a local area network, a telecommunications network, a data communications network, a computer network, an ATM (Asynchronous Transfer Mode) network, a SONET (Synchronous Optical Network) network, an SDH (Synchronous Digital Hierarchy) network, a wireless network, and a wired network. In some embodiments, the network 1204 may include a wireless link, such as an infrared channel or a satellite band. The topology of the network 1204 may be a bus network topology, a star network topology, or a ring network topology. The topology of the network 1204 may be any network topology that is compatible with the operations described herein, as known to those skilled in the art. The network may include a cellular network using any one or more protocols used for communication between mobile devices, including AMPS, TDMA, CDMA, GSM, GPRS, or UMTS. In some embodiments, different types of data may be transmitted via different protocols. In other embodiments, the same type of data may be transmitted via different protocols.

[0228] The client 1202 and remote machine 1206 (generally referred to as computing device 1200) may be any workstation, desktop computer, laptop or notebook computer, server, portable computer, mobile phone or other portable communication device, media playback device, gaming system, handheld computing device, or any other type and / or form of computing, communication, or media device capable of communication and having sufficient processor power and storage capacity to perform the operations described herein. In some embodiments, the computing device 1200 may have a different processor, operating system, and input devices compatible with the computing device 1200. In other embodiments, the computing device 1200 may be a handheld device, a digital audio player, a digital media player, or a combination of such devices. Computing device 1200 may execute, operate, or provide applications such as any type and / or form of software, programs, or executable instructions, including, but not limited to, any type and / or form of web browsers, web-based clients, client-server applications, ActiveX controls, JAVA applets, or any other type and / or form of executable instructions executable on computing device 1200.

[0229] In one embodiment, computing device 1200 provides the functionality of a web server. In some embodiments, web server 1200 comprises an open source web server, such as the Apache server (maintained by the Apache Software Foundation, Delaware). In other embodiments, web server 1200 executes proprietary software, such as the INTERNET INFORMATION SERVICE product (offered by Microsoft Corporation, Redmond, Washington), the ORACLE IPLANET webserver product (offered by Oracle Corporation, Redwood Shores, California), or the BEA WEBLOGIC product (offered by BEA Systems, Santa Clara, California).

[0230] In some embodiments, the system may include a plurality of logically grouped computing devices 1200. In one of these embodiments, the logical grouping of computing devices 1200 may be referred to as a server farm. In another of these embodiments, the server farm may be managed as a single entity.

[0231] 12B and 12C are block diagrams illustrating computing devices 1200 useful for implementing embodiments of a client 1202 or a remote machine 1206. As shown in FIGS. 12B and 12C, each computing device 1200 includes a central processing unit 1221 and a main memory 1222. As shown in FIG. 12B, computing device 1200 may also include a storage device 1228, an installation device 1216, a network interface 1218, an input / output controller 1223, display devices 1224A-n, a keyboard 1226, a pointing device 1227 such as a mouse, and one or more other input / output devices 1230A-n. Storage device 1228 may include, but is not limited to, an operating system and software. As shown in FIG. 12C, each computing device 1200 may include additional optional elements such as a memory port 1203, a bridge 1270, one or more input / output devices 1230A-1230n (generally referred to using the reference numeral 1230), and a cache memory 1240 in communication with the central processing unit 1221.

[0232] Central processing unit 1221 is any logic circuitry that responds to and processes instructions retrieved from main memory 1222. In many embodiments, central processing unit 1221 is provided by a microprocessor unit, such as those manufactured by Intel Corporation (Mountain View, Calif.), Motorola Corporation (Schaumburg, Ill.), Transmeta Corporation (Santa Clara, Calif.), International Business Machines (White Plains, New York), and Advanced Micro Devices (Sunnyvale, Calif.). Other examples include PARC processors, ARM processors, processors used for building UNIX / LINUX "white" boxes, and processors for portable devices. Computing device 1200 may be based on any of these processors, or any other processor capable of operating as described herein.

[0233] Main memory 1222 may be one or more memory chips capable of storing data and allowing microprocessor 1221 direct access to any storage location. Main memory 1222 may be based on any available memory chip capable of operating as described herein. In the embodiment shown in FIG. 12B, processor 1221 communicates with main memory 1222 via system bus 1250. FIG. 12C illustrates an embodiment of computing device 1200 in which the processor communicates directly with main memory 1222 via memory port 1203. FIG. 12C also illustrates an embodiment in which main processor 1221 communicates directly with cache memory 1240 via a secondary bus (also known as a backside bus). In other embodiments, main processor 1221 communicates with cache memory 1240 using system bus 1250.

[0234] In the embodiment shown in FIG. 12B, the processor 1221 communicates with various input / output devices 1230 via a local system bus 1250. Various buses may be used to connect the central processing unit 1221 to any of the input / output devices 1230, including a VESA VL bus, an ISA bus, an EISA bus, a MicroChannel Architecture (MCA) bus, a PCI bus, a PCI-X bus, a PCI-Express bus, or a NuBus. For embodiments in which the input / output device is a video display 1224, the processor 1221 may communicate with the display 1224 using an Advanced Graphics Port (AGP). FIG. 12C shows an embodiment of a computer 1200 in which the main processor 1221 communicates directly with an input / output device 1230b via, for example, HYPERTRANSPORT, RAPIDIO, or INFINIBAND communications technologies.

[0235] The computing device 1200 may include a wide variety of input / output devices 1230A-1230n. Input devices include keyboards, mice, trackpads, trackballs, microphones, scanners, cameras, and drawing tablets. Output devices include video displays, speakers, inkjet printers, laser printers, and dye-sublimation printers. The input / output devices may be controlled by an input / output controller 1223, as shown in FIG. 12B. The input / output devices may also include storage and / or installation devices 1216 for the computing device 1200. In some embodiments, the computing device 1200 may include a USB connection (not shown) for accepting handheld USB storage devices, such as the USB Flash Drive series (manufactured by Twintech Industry, Inc., Los Alamitos, California).

[0236] Continuing with reference to FIG. 12B, computing device 1200 may support any suitable installation device 1216, such as a disk drive that accepts 3.5-inch disks, 5.25-inch disks, ZIP disks, or other floppy disks, a CD-ROM drive, a CD-R / RW drive, a DVD-ROM drive, various formats of tape drives, a USB device, a hard drive, or any other device suitable for installing software and programs. In some embodiments, computing device 1200 may provide functionality for installing software over network 1204. Additionally, computing device 1200 may include storage devices, such as one or more hard disk drives or RAID, for storing the operating system and other software.

[0237] The computing device 1200 may also include a network interface 1218 for connecting to the network 1204 via various connections, including, but not limited to, standard telephone lines, LAN or WAN links (e.g., 802.11, T1, T3, 56kb, X.25, SNA, DECNET), broadband connections (e.g., ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethernet-over-SONET), wireless connections, or combinations of any or all of the above. Connections can be established using a variety of communication protocols (e.g., TCP / IP, IPX, SPX, NetBIOS, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), RS232, IEEE 802.11, IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, IEEE 802.15.4, BLUETOOTH, ZIGBEE, CDMA, GSM, WiMax, and direct asynchronous connections). In one embodiment, computing devices 1200 communicate with other computing devices 1200' via any type and / or form of gateway or tunneling protocol, such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS). The network interface 1218 may comprise an internal network adapter, a network interface card, a PCMCIA network card, a card bus network adapter, a wireless network adapter, a USB network adapter, a modem, or any other device suitable for connecting the computing device 1200 to any type of network capable of communicating and performing the operations described herein.

[0238] In other embodiments, input / output device 1230 may be a bridge between system bus 1250 and an external communications bus, such as a USB bus, an Apple Desktop Bus, an RS-232 serial connection, a SCSI bus, a FireWire® bus, a FireWire® 800 bus, an Ethernet® bus, an AppleTalk bus, a GigabitEthernet® bus, an Asynchronous Transfer Mode bus, a HIPPI bus, a Super HIPPI bus, a SerialPlus bus, an SCI / LAMP bus, a FibreChannel bus, or a Serial Attached Small Computer System Interface bus.

[0239] 12B and 12C typically operate under the control of an operating system that controls task scheduling and access to system resources. Computing device 1200 may run any operating system, such as any version of the MICROSOFT WINDOWS® operating system, different releases of the UNIX® and LINUX® operating systems, any version of MACOS for Macintosh computers, any embedded operating system, any real-time operating system, any open source operating system, any proprietary operating system, any operating system for portable computing devices, or any other operating system capable of running on a computing device and performing the operations described herein.Representative operating systems include, but are not limited to, WINDOWS® 3.x, WINDOWS® 95, WINDOWS® 98, WINDOWS® 2000, WINDOWS® NT 3.51, WINDOWS® NT 4.0, WINDOWS® CE, WINDOWS® XP, WINDOWS® 7, WINDOWS® 8, and WINDOWS® VISTA (all manufactured by Microsoft Corporation, Redmond, Washington), MAC OS (manufactured by Apple Inc., Cupertino, California), OS / 2 (manufactured by International Business Machines, Armonk, New York), and the freely available operating system LINUX® (sold by Caldera Corp., Salt Lake City, Utah), a Linux-based operating system, Red Hat Enterprise Linux® (sold by Red Hat, Inc., Raleigh, North Carolina), the freely available operating system Ubuntu (Canonical Ltd., London, England), or in particular any kind and / or form of UNIX operating system.

[0240] As mentioned above, computing device 1200 may be a portable communications device, a media device, a gaming system, a handheld computing device, or any type and / or form of computing, communications, or media device capable of communications and having sufficient processor power and storage capacity to perform the operations described herein. Computing device 1200 may be, by way of example and not limitation, a handheld device such as those manufactured by Apple Inc. (Cupertino, California), Google / Motorola Div. (Fort Worth, Texas), Kyocera (Kyoto, Japan), Samsung Electronics Co., Ltd. (Seoul, South Korea), Nokia (Finland), Hewlett-Packard Development Company, LP, and / or Palm, Inc. (Sunnyvale, California), Sony Ericsson Mobile Communications AB (Lund, Sweden), or Research In Motion Limited (Waterloo, Ontario, Canada). In yet another embodiment, computing device 1200 is a smartphone, pocket PC, pocket PC phone, or other portable handheld device that is compatible with Microsoft Windows Mobile Software.

[0241] In some embodiments, computing device 1200 is a digital audio player. In one of these embodiments, computing device 1200 is a digital audio player, such as the Apple IPOD®, IPOD® Touch, IPOD® NANO, and IPOD® SHUFFLE series (manufactured by Apple Inc.). In other of these embodiments, the digital audio player may function as both a portable media player and a mass storage device. In other embodiments, computing device 1200 is a digital audio player, such as those manufactured by Samsung Electronics America (Ridgefield Park, New Jersey) or Creative Technologies Ltd. (Singapore), by way of example and not limitation. In yet another embodiment, computing device 1200 is a portable media player or digital audio player that supports file formats including, but not limited to, MP3, WAV, M4A / AAC, WMA Protected AAC, AEFF, Audible audiobooks, Apple Lossless audio file formats, and .mov, .m4v, and .mp4 MPEG-4 (H.264 / MPEG-4 AVC) video file formats.

[0242] In some embodiments, computing device 1200 comprises a combination device, such as a mobile phone combined with a digital audio player or portable media player. In one of these embodiments, computing device 1200 is a combination digital audio player and mobile phone device from the Google / Motorola series. In another of these embodiments, computing device 1200 is an IPHONE® smartphone series (manufactured by Apple Inc.). In yet another of these embodiments, computing device 1200 is a device running the ANDROID® open source mobile phone platform (marketed by the Open Handset Alliance); for example, computing device 1200 may be a device such as those offered by Samsung Electronics (Seoul, Korea) or HTC Headquarters (Republic of China, Taiwan). In other embodiments, the computing device 1200 is a tablet device, such as, by way of example and not limitation, the IPAD® series (manufactured by Apple Inc.), the PLAYBOOK (manufactured by Research In Motion), the CRUZ series (manufactured by Velocity Micro, Inc. (Richmond, Virginia)), the FOLIO and THRIVE series (manufactured by Toshiba America Information Systems Inc. (Irvine, California)), the GALAXY series (manufactured by Samsung), the HPSLATE series (manufactured by Hewlett-Packard), and the STREAK series (manufactured by Dell, Inc. (Round Rock, Texas)).

[0243] 12D, one embodiment of a system in which hosting and delivery services are provided by multiple networks is disclosed. Broadly, the system includes a cloud services hosting infrastructure 1280, a service provider data center 1282, and an information technology (IT) network 1284.

[0244] In one embodiment, data center 1282 includes computing devices such as, but not limited to, servers (e.g., including application servers, file servers, databases, and backup servers), routers, switches, and communications equipment. In other embodiments, cloud services hosting infrastructure 1280 provides access to services for accessing remotely located hardware and software platforms, including, but not limited to, storage systems, databases, application servers, desktop servers, directory services, web servers, and other services. In yet other embodiments, cloud services hosting infrastructure 1280 includes data center 1282. However, in other embodiments, cloud services hosting infrastructure 1280 relies on services provided by a third-party data center 1282. In some embodiments, IT network 1204c may provide local services such as mail services and web services. In other embodiments, IT network 1204c may provide local versions of remotely located services, such as locally cached versions of remotely located print servers, databases, application servers, desktop servers, directory services, and web servers. In other embodiments, the additional servers may reside across cloud service hosting infrastructure 1280, data center 1282, or other networks, such as those provided by third party service providers, including, but not limited to, infrastructure service providers, application service providers, platform service providers, tool service providers, and desktop service providers.

[0245] In one embodiment, a user of client 1202b accesses services provided by remote server 1206A. For example, an administrator of corporate IT network 1284 may decide that a user of client 1202A has access to an application running on a virtual machine running on remote server 1206A. As another example, an individual user of client 1202b may use resources provided to consumers by remote server 1206 (such as email, fax, voice, or other communication services, data backup services, or other services).

[0246] 12D , data center 1282 and cloud service hosting infrastructure 1280 are located remotely from the individuals or organizations they serve. For example, data center 1282 may be on a first network 1204A, cloud service hosting infrastructure 1280 may be on a second network 1204b, and IT network 1284 may be a separate, third network 1204c. In other embodiments, data center 1282 and cloud service hosting infrastructure 1280 are on the first network 1204A, and IT network 1284 is a separate, second network 1204c. In yet other embodiments, cloud service hosting infrastructure 1280 is on the first network 1204A, and data center 1282 and IT network 1284 form second network 1204c. While Figure 12D shows only one server 1206A, one server 1206b, one server 1206c, two clients 1202, and three networks 1204, it should be understood that the system may include multiples of any of the elements, or multiples of each of the elements. Servers 1206, clients 1202, and networks 1204 may be included as described above in connection with Figures 12A-12C.

[0247] Thus, in some embodiments, an IT infrastructure may extend from a first network, such as a network owned and managed by an individual or company, to a second network, the owner or administrator of which may be separate from the first network. The resources provided by the second network may be said to be “on the cloud.” Elements in the cloud may include, but are not limited to, storage devices, servers, databases, computing environments (including virtual machines, servers, and desktops), and applications. For example, the IT network 1284 may use a remote data center 1282 to store servers (including, e.g., application servers, file servers, databases, and backup servers), routers, switches, and communications equipment. The data center 1282 may be owned and managed by the IT network 1284, or a third-party service provider (including, e.g., a cloud service hosting infrastructure provider) may provide access to another data center 1282. As another example, machine 102A described above in connection with FIG. 3 may be owned or managed by a first entity (e.g., cloud service hosting infrastructure provider 1280), and machine 102b described above in connection with FIG. 3 may be owned or managed by a second entity (e.g., service provider data center 1282) to which client 1202 connects directly or indirectly (e.g., using resources provided by any of 1280, 1282, or 1284).

[0248] In some embodiments, one or more networks that provide computing infrastructure for consumers may be referred to as a cloud. In one of these embodiments, a system in which users of a first network access at least a second network (including a pool of abstracted, scalable, managed computing resources capable of hosting resources) may be referred to as a cloud computing environment. In other of these embodiments, resources may include, but are not limited to, virtualization technology, data center resources, applications, and management tools. In some embodiments, Internet-based applications (which may be provided via a "software-as-a-service" model) may be referred to as cloud-based resources. In other of these embodiments, a network that provides users with computing resources, such as remote servers, virtual machines, or blades on blade servers, may be referred to as a compute cloud or "infrastructure-as-a-service" provider. In yet other embodiments, a network that provides storage resources, such as a storage area network, may be referred to as a storage cloud. In other embodiments, resources may be cached on a local network and stored in the cloud.

[0249] In some embodiments, some or all of the remote machines 1206 may be leased or rented from third-party companies, such as, by way of example and not limitation, Amazon Web Services LLC (Seattle, WA), Rackspace US, Inc. (San Antonio, TX), Microsoft Corporation (Redmond, WA), and Google Inc. (Mountain View, CA). In other embodiments, all of the hosts 1206 are owned and managed by third-party companies, including, but not limited to, Amazon Web Services LLC, Rackspace US, Inc., Microsoft, and Google.

[0250] As noted above, many types of hardware may be used in conjunction with the above-described systems and methods to provide the functionality described above, although in some embodiments the hardware itself may be modified to provide improved execution of the above-described methods and systems.

[0251] 17A is a block diagram illustrating one embodiment of a system for identifying a level of similarity between multiple binary vectors. System 1700 includes machine 102a executing engine 101 as described above. System 1700 includes machine 102b including processor 1221, data bus 1722 (each of which may be part of system bus 1250 or memory port 1203 described above), address bus 1724 (each of which may be part of system bus 1250 or memory port 1203 described above), and multiple memory cells 1702a-n (which may be referred to herein as memory cells 1702). Memory cells 1702 each include a first register 1704a, a second register 1704b, and a bitwise comparison circuit 1710.

[0252] Referring now to FIG. 17B, a flowchart illustrating one embodiment of a method 1750 for identifying a level of similarity between multiple binary vectors is shown in relation to FIG. 17A. The method 1750 includes storing, by a processor on a computing device, one of multiple binary vectors in each of multiple memory cells on the computing device, each of the multiple memory cells including a bitwise comparison circuit (1752). The method 1750 includes receiving, by the computing device, a binary vector for comparison with each of the stored multiple binary vectors (1754). The method 1750 also includes providing, by the processor, via a data bus, the received binary vector to each of the multiple memory cells (1756). The method 1750 includes determining, by each bitwise comparison circuit, a level of overlap between the received binary vector and the binary vector stored in the memory cell associated with the bitwise comparison circuit (1758). The method 1750 also includes determining, by each of the multiple bitwise comparison circuits, whether the level of overlap satisfies a threshold value provided by the processor (1760). The method 1750 includes, by each comparison circuit that determines that the level of overlap meets the threshold, providing to the processor an identification of the stored binary vector having a satisfactory level of overlap (1762). The method 1750 includes providing, by the processor, an identification of each stored binary vector that meets the threshold and the level of similarity between the stored binary vector and the received binary vector (1764).

[0253] The method 1750 includes storing, by a processor on the computing device, one of a plurality of binary vectors in each of a plurality of memory cells on the computing device, each of the plurality of memory cells including a bitwise compare circuit (1752). The processor 1221 can receive the plurality of binary vectors from the engine 101 for storage. In one embodiment, the processor 1221 uses the address bus 1724 to identify the memory cell in which the one of the plurality of binary vectors is stored. In another embodiment, the processor 1221 uses the data bus 1722 to send the binary vector to the memory cell for storage in the first register 1704a. The machine 102b can implement cell selector logic, chip selector logic, and board selector logic to address the particular memory cell in which the binary vector is stored.

[0254] The method 1750 includes receiving 1754, by a computing device, a binary vector for comparison with each of the stored plurality of binary vectors. In one embodiment, the processor 1221 receives the binary vector. In another embodiment, a user (e.g., a user of machine 102b or a different machine 100) provides the binary vector. In some embodiments, the processor 1221 also receives a request for identification of similar binary vectors.

[0255] The method 1750 includes providing 1756, by the processor, via the data bus to each of the plurality of memory cells. In one embodiment, the processor 1221 transmits the same binary vector to all of the memory cells. In another embodiment, the processor 1221 transmits an instruction to store the received binary vector (e.g., in the second register 1704b). In another embodiment, the processor 1221 transmits an instruction to compare a previously stored binary vector (e.g., the binary vector in the first register 1704a) with the received binary vector (e.g., stored in the second register 1704b).

[0256] The method 1750 includes determining, by each bitwise comparison circuit, a level of overlap between the received binary vector and the binary vector stored in the memory cell associated with the bitwise comparison circuit (1758). In one embodiment, the bitwise comparison circuit in the memory cell instructs a shift register to compare a bit in the first register 1704a with a corresponding bit in the second register 1704b (e.g., both bits in a first position in the register, both bits in a second position in the register, etc.). In another embodiment, the bitwise comparison circuit instructs the shift register to return a 1 if both bits in a particular position are set to 1 (e.g., the bits are the same). In yet another embodiment, the bitwise comparison circuit adds the number of received 1s to calculate a number representing a level of overlap between the received binary vector and the binary vector stored in the memory cell associated with the bitwise comparison circuit.

[0257] Method 1750 includes determining, by each of a plurality of bitwise comparison circuits, whether the level of overlap satisfies a threshold provided by the processor (1760). In one embodiment, processor 1221 transmits a number representing a certain percentage of overlap to the memory cells (e.g., for a memory cell capable of storing 16,000 pieces of information in a register, processor 1221 transmits more than 16,000 only if it wishes to receive the identification of memory cells where there was 100% overlap between the binary vectors); the bitwise comparison circuit determines whether the calculated overlap level matches the number transmitted from processor 1221. For example, without limitation, if the bitwise comparison circuit determines that there are 16,000 instances where two registers each contain the same data, the bitwise comparison circuit can transmit an indication to the processor that the memory cells meet the threshold level of overlap; if the bitwise comparison circuit determines that there are only 14,000 instances where two registers each contain the same data, the bitwise comparison circuit does not respond to processor 1221. In some embodiments, processor 1221 uses a number representing a threshold level of overlap previously specified by a user. In other embodiments, processor 1221 uses the highest number available (e.g., the highest number of bits a register can store) as a counter, decrementing the number it sends to the memory cells, and recursing until it receives a response from the bitwise comparison circuit indicating there is a memory cell storing a binary vector that has a level of overlap with the received binary vector that meets the threshold received from processor 1221.

[0258] The method 1750 includes, for each comparison circuit that determines that the level of overlap meets the threshold, providing to the processor an identification of the stored binary vector having a satisfactory level of overlap (1762). For example, the identifier may be stored in a third register.

[0259] The method 1750 includes providing, by the processor, an identification of each stored binary vector that meets the threshold and level of similarity between the stored binary vector and the received binary vector (1764). The processor may return the identification information directly to the user (e.g., via a user interface). Alternatively, the processor may return the identification information to any executing process in which a comparison between the two binary vectors was originally requested.

[0260] In this way, the comparison and sorting are both accomplished simultaneously and in the memory cells rather than in the processor. In contrast to the methods and systems described herein, conventional systems for utilizing memory cannot feasibly store large binary vectors because typical techniques for size reduction (e.g., hashing) are ineffective for comparisons between very large binary vectors.

[0261] In some embodiments, the systems and methods described in connection with Figures 17A and 17B can be utilized to improve the efficiency and speed of comparison. Thus, in one of these embodiments, the systems and methods can be used to replace or augment similarity engine 304 (effectively implementing similarity engine 304 in hardware rather than software). The systems and methods described above in connection with Figures 1A through 16 can be combined with the systems and methods described in connection with Figures 17A and 17B.

[0262] FIG. 18A is a block diagram illustrating one embodiment of a system for identifying a level of similarity between multiple data representations. In addition to the components explicitly described above in connection with FIG. 17A, the methods and systems described herein may include the components described in FIG. 18A. For example, a separate register may be used to store a document reference identifying the document containing the data item from which the stored SDR ("fingerprint") was generated. As another example, the bitwise comparison circuit 1710a may include additional subcomponents, such as an overlap adder and a comparator, that provide the functionality described herein. FIG. 18B is a flowchart illustrating one embodiment of a method for identifying a level of similarity between multiple data representations.

[0263] 18A-18B in conjunction with FIGS. 17A-17B. A method 1850 for identifying a level of similarity between a first data item and data items in a set of data documents includes clustering 1852, by a reference map generator executing on a first computing device, the selected set of data documents in a two-dimensional metric space according to at least one criterion, to generate a semantic map. The method 1850 also includes associating 1854, by the semantic map, with each of the set of data documents a coordinate pair. The method 1850 also includes generating 1856, by a parser executing on the first computing device, a list of data items present in the set of data documents. The method 1850 also includes determining 1858, by a representation generator executing on the first computing device, presence information for each data item in the list, including: (i) the number of data documents in which the data item occurs; (ii) the number of occurrences of the data item in each data document; and (iii) a coordinate pair associated with each data document in which the data item occurs. The method 1850 includes generating, by a list generator, a sparse distributed representation (SDR) for each data item in the list using the presence information, resulting in a plurality of SDRs (1860). The method 1850 includes storing, by a processor on a second computing device, one of the plurality of generated SDRs in each of a plurality of memory cells on the second computing device, each of the plurality of memory cells including a bitwise comparison circuit (1862). The method 1850 includes receiving, by the second computing device, a first data item from a third computing device (1864). The method 1850 includes providing, by the processor, the SDR of the first data item to each of the plurality of memory cells via a data bus (1866). The method 1850 includes determining, by each of the plurality of bitwise comparison circuits, a level of overlap between the SDR of the first data item and the generated SDR stored in the memory cell associated with the bitwise comparison circuit (1868). The method 1850 includes determining, by each of the plurality of bitwise comparison circuits, whether the level of overlap satisfies a threshold value provided by the processor (1870).The method 1850 includes, by each comparison circuit that determines that the level of overlap meets the threshold, providing to the processor (1872) the document reference number stored in the associated memory cell, the document reference number identifying the document containing the data item for which the SDR stored in the memory cell was generated. The method 1850 also includes, by the second computing device, providing to a third computing device (1874) the identity of each data item for which the SDR stored in the memory cell meets the threshold, and the level of similarity between the data item for which the stored SDR was generated and the received data item.

[0264] In one embodiment, steps 1852-1860 are performed as described above in connection with steps 202-214 of Figure 2.

[0265] The method 1850 includes storing, by a processor on a second computing device, one of the plurality of generated SDRs in each of a plurality of memory cells on the second computing device (1862), each of the plurality of memory cells including a bitwise comparison circuit. The processor can store each of the plurality of generated SDRs in a plurality of memory cells (1752), as described above in connection with FIG. 17(B).

[0266] The method 1850 includes receiving, by the second computing device, the first data item from the third computing device 1864. The processor may receive the first data item as described above (for example, but not limited to, with respect to FIG. 5).

[0267] The method 1850 includes providing, by the processor, via the data bus, an SDR of the first data item to each of a plurality of memory cells 1866. The processor first directs generation of the SDR as described above, and then provides the generated SDR to the memory cells for comparison with a previously stored SDR.

[0268] The method 1850 includes determining, by each of the plurality of bitwise comparison circuits, a level of overlap between the SDR of the first data item and the generated SDR stored in a memory cell associated with the bitwise comparison circuit (1868). The bitwise comparison circuits may perform the determination as described above in connection with Figures 17A and 17B. The method 1850 includes determining, by each of a plurality of bitwise comparison circuits, whether the level of overlap satisfies a threshold value provided by the processor (1870). The bitwise comparison circuits may perform the determination as described above in connection with Figures 17A and 17B.

[0269] The method 1850 includes, by each comparison circuit that determines that the level of overlap meets a threshold, providing the processor with the document reference number stored in the associated memory cell (the document reference number identifies the document containing the data item from which the SDR stored in the memory cell was generated) 1872. The bitwise comparison circuit can provide the determined level of overlap as described above in connection with Figures 17A and 17B.

[0270] Method 1850 includes providing, by the second computing device, to a third computing device, an identification of each data item for which an SDR stored in a memory cell satisfies a threshold value and a level of similarity between the data item for which the stored SDR was generated and the received data item (1874). The processor can provide the identification information, as described above in connection with Figures 17A-B, either directly to the third computing device or to another processor that performs any of the methods for performing a comparison between SDRs described above in connection with Figures 1A-16.

[0271] Having described an embodiment of a method and system for recursively generating data item fingerprints, it will be apparent to those skilled in the art that other embodiments incorporating the concepts herein may be used. Accordingly, the specification should not be limited to any one embodiment, but rather should be limited only by the spirit and scope of the following claims.

Claims

1. 1. A method for generating clusters of distributed representations in a second two-dimensional metric space using distributed representations of data items in a first set of data documents clustered in a first two-dimensional metric space, the method comprising: clustering the set of data documents selected according to at least one criterion in a two-dimensional distance space by a reference map generator running on a computing device to generate a semantic map, wherein the reference map generator calculates a position on the semantic map of each point representing each document of the set of data documents; associating with each document of the set of data documents by said semantic map a coordinate pair that locates each document on said semantic map; generating, by a parser executing on said computing device, a list of data items present in said set of data documents; determining, by a representation generator executing on the computing device, presence information for each data item in the list, including: (i) the number of data documents in which the data item occurs; (ii) the number of occurrences of the data item in each data document; and (iii) the coordinate pair associated with each data document in which the data item occurs; generating, by the representation generator, a distributed representation for each data item using the presence information; receiving, by a sparsification module executing on the computing device, identification information of a maximum level of sparsity; reducing the total number of set bits in each distributed representation based on the maximum level of sparsity using the sparsification module to generate sparse distributed representations (SDRs) with a reference fill grade; storing each of said SDRs in an SDR database; and clustering, by the reference map generator running on the computing device, a set of SDRs retrieved from the SDR database and selected according to at least one second criterion, in a second two-dimensional distance space to generate a second semantic map. method.

2. The method of claim 1 , wherein the set of SDRs is selected based on receiving an indication from a full-text search system that the SDRs are associated with a second set of data documents.

3. providing at least one snippet of at least one data document in the second set of data documents to a full-text search system; receiving from the full-text search system a list of coordinate pairs of matching data documents in the set of data documents that contain the provided snippet; and The method of claim 1 , further comprising retrieving from the SDR database at least one SDR associated with each of the coordinate pairs in the list of coordinate pairs.

Citation Information

Patent Citations

  • Method and system for mapping data items to sparse distributed representation

    JP2017524200A

  • Method and system for determining similarity between filtering criteria and data items in a set of stream documents

    JP2018530047A