File Relationship Mapping via Metadata Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face challenges in efficiently processing and analyzing unstructured data, particularly in determining relationships between files, which is time-consuming and prone to human error, especially in the context of rapidly growing information and complex regulatory documents.

Innovation Solution

A method and system that extracts metadata from files to generate a data structure indicating relationships between them, applying weighting factors to determine the degree of similarity, allowing for automated mapping and updating of file relationships, and providing a report on potential impacts such as risk assessments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis is used to determine relationships between unstructured files, then accuracy can be maintained through human judgment, but processing time and labor costs increase significantly

Engineering Contradiction:
Improveaccuracy of relationship determinationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces metadata as an intermediary layer between unstructured files and relationship analysis. Metadata extracts key attributes (author, date, subject, keywords) that serve as mediators for automated comparison, enabling machines to efficiently determine file relationships without manually analyzing entire document contents

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual human analysis (mechanical system) with automated computational processing. By using algorithms to compare metadata attributes and calculate similarity scores, the system substitutes human cognitive work with machine-based automated relationship determination

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated processing is applied to unstructured data, then processing speed increases, but accuracy decreases due to lack of contextual understanding

Engineering Contradiction:
Improveprocessing speedVSAvoidaccuracy of relationship determination
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent extracts essential contextual information from unstructured files in the form of metadata attributes (author, creation date, subject, keywords). By taking out only the most relevant identifying characteristics, the system enables automated processing to achieve both speed and accuracy by focusing on key discriminative features rather than analyzing entire documents

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If comprehensive metadata analysis is performed to map file relationships, then relationship accuracy improves, but system complexity increases

Engineering Contradiction:
Improverelationship mapping accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the relationship determination process into distinct modular components: metadata extraction, attribute comparison, similarity scoring, and relationship classification. This segmentation allows each component to be independently optimized and maintained, reducing overall system complexity while achieving comprehensive analysis through coordinated simple operations

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10997181B2Generating a data structure that maps two files
Publication Date: 2021.05.04 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10997181B2 patent drawing
  • US10997181B2 patent drawing
  • US10997181B2 patent drawing

AI summary

A first file and a second file are retrieved from a database, in which the first and second files include an unstructured text stream. Metadata of the first and second files are extracted. The extracted metadata include a description category, entity source, geographic region, and a set of sub-files linked to the file. A data structure indicative of relationship between the first and second files is generated. Weighting factor is applied to the generated data structure, which indicates a degree of relationship between the first file and the second file. The relationship and the degree of the relationship are determined based on the extracted metadata of the first and second files. In response to a user requesting the first file, it is determined whether the second file should be provided in conjunction with the first file based on the weighting factor as applied to the data structure.