Document Fusion Score via Hierarchical Structure Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge of efficiently retrieving relevant information from voluminous electronic documents is exacerbated by the difficulty in comparing and matching the structural and semantic features across documents, leading to ineffective search results.

Innovation Solution

A document management system generates a fusion score by extracting and comparing sets of features, including hierarchical structures and semantic data, using machine learning models to compute similarity scores and weighted features, thereby determining the relevance between electronic documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple feature types including hierarchical structure are extracted and compared, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvedocument comparison accuracyVSAvoidfeature extraction and comparison system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments document features into distinct types including hierarchical structure, semantic content, and metadata. Each feature type is extracted and processed separately through dedicated processing steps, allowing the system to handle complex multi-dimensional document comparison by breaking it down into manageable feature segments that can be independently analyzed and then integrated.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If machine learning models are used to generate weighted features, then manufacturing precision is improved, but ease of manufacture deteriorates

Engineering Contradiction:
Improvefeature weighting accuracyVSAvoidsystem implementation complexity
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent employs machine learning models to pre-compute optimal weights for different feature types during a training phase. This preliminary action allows the system to learn the relative importance of various document features beforehand, so that during actual document comparison operations, the pre-determined weights can be directly applied without requiring complex real-time calculations, thus improving precision while reducing operational complexity.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If fusion scores are generated by comparing multiple feature sets, then reliability is improved, but loss of time increases

Engineering Contradiction:
Improvedocument retrieval accuracyVSAvoidcomputation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements a tiered fusion score computation approach where not all feature types are compared with equal depth for every document pair. The system computes fusion scores using a combination of feature types, applying more rigorous comparison to critical features while using simpler comparison methods for less important features. This partial action approach maintains reliable retrieval accuracy by focusing computational effort on the most discriminative features rather than exhaustively processing all features at maximum detail.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11625423B2Linear late-fusion semantic structural retrieval
Publication Date: 2023.04.11 JPMORGAN CHASE BANK NA
  • US11625423B2 patent drawing
  • US11625423B2 patent drawing
  • US11625423B2 patent drawing

AI summary

Systems and methods for generating a fusion score between electronic documents. The method includes receiving a first electronic document by a document management system. The method further includes extracting a first set of features from the first electronic document including at least one feature type indicating the hierarchical structure of the first electronic document. The method also includes receiving a second electronic document by the document management server. The method further includes extracting a second set of features from the second electronic document including at least one feature type indicating the hierarchical structure of the second electronic document. The method further includes generating a fusion score based on a comparison of the first set of features and the second set of features.