File Management System with Self-Learning Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document categorization systems are limited by their reliance on mathematical comparisons and require costly and ambiguous pairwise comparisons, often necessitating explicit training and prior knowledge of cluster numbers, which can be cumbersome and inefficient.

Innovation Solution

A hybrid method that uses recognition samples with defined key characteristics to identify clusters of similar files, eliminating the need for explicit training and allowing for arbitrary cluster formation based on file characteristics, and employing an AI engine for semi-autonomous categorization and characterization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If mathematical comparison methods are used for document categorization, then categorization can be performed with existing knowledge samples, but the system is limited to existing knowledge and requires explicit training

Engineering Contradiction:
Improveease of system setupVSAvoidflexibility of categorization
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system performs self-learning by automatically analyzing document characteristics and forming clusters without requiring explicit training data or predefined categories. The categorization model autonomously identifies patterns and creates category structures based on the documents it processes, eliminating the need for manual training setup while maintaining adaptability to new document types.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If pairwise comparison methods are used to form similarity groups, then documents can be grouped by similarity, but the approach is costly and ambiguous

Engineering Contradiction:
Improveaccuracy of similarity groupingVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts key characteristics and features from documents to create condensed representations that capture essential similarity information. By working with these extracted features rather than performing full pairwise comparisons of entire documents, the system achieves accurate similarity grouping while dramatically reducing computational complexity and ambiguity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms documents into standardized feature vectors with consistent parameter representations. This parameter transformation enables efficient comparison and clustering by converting complex document data into a uniform format that highlights key similarities and differences, reducing both computational cost and ambiguity in grouping decisions.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If explicit training and prior knowledge of cluster numbers are required, then categorization can be performed with structured guidance, but the process becomes cumbersome and inefficient

Engineering Contradiction:
Improvereliability of categorizationVSAvoidoperational simplicity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically determines the number of clusters and category structures through self-organization of document features. It performs self-learning by analyzing document characteristics and autonomously forming categories without requiring users to specify cluster numbers or provide structured training guidance, thereby maintaining reliability while dramatically improving operational simplicity.

Inventive Principle:
Principle #25Self-service

4Ease of manufacture

If traditional categorization systems are used, then file organization can be achieved, but the systems lack dynamic cluster formation and require expert training

Engineering Contradiction:
Improveease of implementationVSAvoidlevel of autonomous operation
Core Design Contradiction:
Ease of manufactureVSExtent of automation

Solution Approach 1:

The categorization model performs self-learning and automatic cluster formation by analyzing document characteristics and autonomously determining category structures. This self-organizing capability eliminates the need for expert training or manual configuration while maintaining ease of implementation, as the system adapts automatically to the specific document set it processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements dynamic cluster formation where category structures evolve automatically based on the documents being processed. Rather than relying on static predefined categories, the system dynamically adjusts and creates category groupings that reflect the actual characteristics and relationships within the document set, enabling high automation while remaining easy to implement.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12117964B2File management systems and methods
Publication Date: 2024.10.15 DOKKIO INC
  • US12117964B2 patent drawing
  • US12117964B2 patent drawing
  • US12117964B2 patent drawing

AI summary

Example file management systems and methods are described. In one implementation, a system detects a user entry in a document. The system then retrieves knowledge relevant to the user entry. The system also presents the knowledge to a user.