Social Community Identification for Document Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data classification methods fail to effectively identify data files with common characteristics, particularly in large datasets, due to the lack of consideration for social community associations, which can lead to missed relevant information and inefficient retrieval of data.

Innovation Solution

A method and system that generate and update lists of key terms by analyzing data files within social communities, using hierarchical structures and decision trees to classify data files based on social community associations, physical, and semantic connections, enabling the identification of data files with common characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional data classification methods are used, then the classification process is simple, but the ability to identify data files with common characteristics is poor

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the classification process into multiple hierarchical levels (upper nodes for general similarities, lower nodes for specific similarities). This segmentation allows the system to progressively refine classification accuracy without overwhelming complexity at any single level, resolving the contradiction between measurement precision and device complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces social community associations as an additional dimension for classification beyond traditional content-based methods. By incorporating social context (who created the data, their relationships, community memberships), the system achieves higher classification accuracy without proportionally increasing system complexity, as this is a new dimension rather than a multiplication of existing complex processes.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If social community associations are incorporated into classification, then data retrieval efficiency improves, but processing complexity increases

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-establishing social community associations and hierarchical structures before actual data classification and retrieval operations. Social networks, community memberships, and hierarchical node structures are built in advance, allowing rapid classification during retrieval without real-time computation of complex social relationships, thus improving productivity while controlling processing complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces hierarchical structures and key term lists as intermediaries between raw social community data and final classification decisions. These intermediaries simplify the processing by pre-processing social associations into structured formats (hierarchical nodes, key terms), reducing the complexity of direct social network analysis while maintaining improved retrieval efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If hierarchical structures with multiple nodes are used, then classification precision improves, but computational requirements increase

Engineering Contradiction:
Improveclassification precisionVSAvoidcomputational energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The hierarchical structure segments classification into upper nodes (general similarities) and lower nodes (specific similarities), allowing the system to stop at appropriate levels based on needs. This segmentation enables precise classification when necessary while avoiding unnecessary computational energy expenditure by not always traversing to the most granular levels, resolving the contradiction between precision and energy consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial action by selectively applying detailed lower-node classification only when higher-node classification is insufficient or when precision requirements demand it. For many routine retrievals, upper-node classification provides sufficient precision with lower computational energy, while the option for more precise lower-node classification remains available when needed, balancing precision and energy usage.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9317594B2Social community identification for automatic document classification
Publication Date: 2016.04.19 SAS INSTITUTE INC
  • US9317594B2 patent drawing
  • US9317594B2 patent drawing
  • US9317594B2 patent drawing

AI summary

Systems and methods for identifying data files that have a common characteristic are provided. A plurality of data files are received. The plurality of data files include one or more data files having the common characteristic. A list of key terms is generated from the plurality of data files. Data files from the plurality of data files that have an association with a social community are identified, where the social community is defined by one or more features. The list of key terms is updated based on an analysis of the identified features. The updated list of key terms is used to identify other data files that have the common characteristic.