Semiotic Indexing for Digital Resources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for indexing and classifying large amounts of domain-specific documents or digital content face challenges such as the need for extensive training, poor resolution in heterogeneous corpora, and increased instances of semantic ambiguities like synonymies, meronymies, and polysemies, which hinder rapid and precise retrieval of relevant documents.

Innovation Solution

The use of an externally managed, semantically disambiguated nomenclature or terminology, combined with persistent identifiers, allows for efficient classification and retrieval by generating vectors that act as signatures or fingerprints for documents, overcoming limitations of existing methods by providing a tunable classifier and resolving semantic ambiguities within specific fields.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If natural language processing and vector space models are used to automatically extract meaning from documents, then the automation of document classification is improved, but measurement precision deteriorates due to semantic ambiguities like synonymies, polysemies, and the inherent ambiguity of natural language

Engineering Contradiction:
Improveautomation of document classificationVSAvoidprecision of semantic interpretation
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary layer between natural language processing and document classification. This intermediary uses controlled vocabularies, taxonomies, and ontologies to mediate the semantic interpretation process, translating ambiguous natural language into precise structured representations that resolve synonymies and polysemies while maintaining automation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter of semantic representation from raw natural language vectors to structured semantic annotations based on controlled vocabularies and ontologies. This parameter transformation allows automated processing while improving measurement precision by constraining interpretations to predefined semantic frameworks

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If extensive training is conducted to improve document classification accuracy, then measurement precision is improved, but loss of time increases due to the training requirement

Engineering Contradiction:
Improveaccuracy of document classificationVSAvoidtime required for training
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-defining controlled vocabularies, taxonomies, and ontologies before the actual classification task. These semantic frameworks are prepared in advance and can be reused across multiple classification tasks, eliminating the need for extensive re-training while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates universal semantic frameworks (controlled vocabularies and ontologies) that can be applied across different domains and classification tasks. These multi-functional resources serve multiple purposes and can be reused without extensive re-training, reducing time loss while maintaining precision

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If domain-specific terminologies and nomenclatures are used to improve classification accuracy, then measurement precision is improved, but device complexity increases due to the need to manage multiple terminologies and their relationships

Engineering Contradiction:
Improveaccuracy of domain-specific classificationVSAvoidcomplexity of terminology management
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a nested structure where controlled vocabularies are embedded within taxonomies, which are in turn embedded within ontologies. This nested organization allows domain-specific terminologies to be structured hierarchically, managing complexity by organizing multiple terminologies within a unified framework rather than treating them as separate systems

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent adds a new dimension to terminology management by introducing structured semantic relationships (hypernymy, synonymy, polysemy) as an organizational dimension. This transforms flat terminology lists into multi-dimensional semantic networks, enabling precise classification while managing complexity through structured relationships

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Adaptability or versatility

If the volume of digital material is increased to provide more comprehensive information, then adaptability is improved, but productivity deteriorates because knowledge workers must search through increasingly large volumes of material

Engineering Contradiction:
Improvecomprehensiveness of information coverageVSAvoidefficiency of information retrieval
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent extracts key semantic features and concepts from large volumes of digital material and represents them in compact structured forms using controlled vocabularies and ontologies. This extraction process creates concise semantic representations that capture the essential information without requiring workers to search through the entire volume of material

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical process of manual searching and filtering through large volumes of material with automated semantic processing. The system automatically applies controlled vocabularies and ontologies to classify and organize information, substituting automated semantic analysis for manual information retrieval efforts

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8903825B2Semiotic indexing of digital resources
Publication Date: 2014.12.02 RGT UNIV OF CALIFORNIA
  • US8903825B2 patent drawing
  • US8903825B2 patent drawing
  • US8903825B2 patent drawing

AI summary

A method of classifying a plurality of documents. The method includes steps of providing a first set of classification terms and a second set of classification terms, the second set of classification terms being different from the first set of classification terms; generating a first frequency array of a number of occurrences of each term from the first set of classification terms in each document; generating a second frequency array of a number of occurrences of each term from the second set of classification terms in each document; generating a first similarity matrix from the first frequency array; generating a second similarity matrix from the second frequency array; determining an entrywise combination of the first similarity matrix and the second similarity matrix; and clustering the plurality of documents based on the result of the entrywise combination.