Document Metadata Scoring via Term Co-occurrence Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document retrieval systems face challenges in efficiently organizing and storing documents for subsequent retrieval, particularly in effectively utilizing metadata to associate terms and determine their relevance across documents.

Innovation Solution

A data handling device that analyzes existing metadata to generate statistical data on term co-occurrence, assigns scores based on association strength and frequency, and selects a subset of terms with the highest scores for new documents, using a system that includes a CPU, storage devices, and network interface to facilitate controlled indexing and searching.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional metadata assignment methods are used, then the process is simple, but the search efficiency and relevance are insufficient

Engineering Contradiction:
Improvesearch relevanceVSAvoidmetadata analysis complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of existing metadata to generate statistical data on term co-occurrence patterns before new documents are indexed. This pre-computed statistical information is then used to automatically assign high-quality metadata to new documents, improving search relevance without requiring complex real-time analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The metadata assignment process serves itself by using previously analyzed metadata patterns to automatically generate appropriate metadata for new documents. The system leverages its own accumulated knowledge base of term co-occurrences to improve its metadata assignment capability iteratively.

Inventive Principle:
Principle #25Self-service

2Reliability

If comprehensive metadata is assigned to all documents, then search coverage is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvesearch coverageVSAvoidmetadata processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Instead of assigning all possible metadata terms to every document, the system selectively assigns only the most relevant terms based on statistical analysis of co-occurrence patterns. This partial action approach ensures sufficient search coverage while avoiding the time and computational overhead of processing excessive metadata for each document.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If manual metadata assignment is performed, then accuracy is high, but productivity is low

Engineering Contradiction:
Improvemetadata accuracyVSAvoiddocument processing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system replaces the manual mechanical process of metadata assignment with an automated computational system that uses statistical analysis of term co-occurrences. This substitution maintains high accuracy by leveraging learned patterns from existing metadata while dramatically increasing productivity through automated processing of new documents.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9165063B2Organising and storing documents
Publication Date: 2015.10.20 BRITISH TELECOM PLC
  • US9165063B2 patent drawing
  • US9165063B2 patent drawing
  • US9165063B2 patent drawing

AI summary

A data handling device has access to a store of existing metadata pertaining to existing documents having associated metadata terms. It analyses the metadata to generate statistical data as to the co-occurrence of pairs of terms in the metadata of one and the same document. When a fresh document is received, it is analysed to assign to it a set of terms and determine for each a measure of their strength of association with the document. Then, for each term of the set, a score is generated that is a monotonically increasing function of (a) the strength of association with the document and of (b) the relative frequency of co-occurrence of that term and another term that occurs in the set; metadata for the fresh document are then selected as the subset of the terms in the set having the highest scores.