Graph Grammar Induction for ML Classifier Data Pre-processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional machine learning methods struggle with classifying minor ontological differences in complex networks due to the complexity of extracting naively represented graph structures, requiring extensive data and computational resources, especially in applications like medical scans and design analysis.

Innovation Solution

A data processing system that employs a probabilistic chunking method paired with a multi-scale random walk based graph exploration approach to efficiently induce grammars for design representations, reducing computational complexity and enabling classification with minimal training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional machine learning methods are used for classifying complex networks, then classification can be performed, but the method requires extensive training data and computational resources

Engineering Contradiction:
Improveclassification accuracyVSAvoidtraining data quantity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the graph structure into multiple levels of abstraction by extracting grammatical rules at different scales. The system identifies local patterns (nodes, edges, cycles) and combines them into higher-order grammatical structures, allowing classification to operate on compressed representations rather than raw graph data, thus reducing the effective data quantity needed while maintaining classification accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces grammatical rules as an intermediary representation between raw graph data and classification decisions. These rules serve as a compressed language that captures essential graph structures without requiring the full original data, enabling accurate classification with minimal training data by learning the grammatical patterns rather than memorizing specific graph instances.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If traditional machine learning methods are used for classifying complex networks, then classification can be performed, but computational resources and time are excessively consumed

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the graph analysis into hierarchical levels, extracting grammatical rules at different scales from local to global structures. This segmentation allows the system to process graphs by combining results from smaller units rather than analyzing the entire graph at once, significantly reducing computational time while maintaining accuracy through the hierarchical aggregation of grammatical information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts grammatical rules from the graph structure, separating the essential structural information from the redundant data. By taking out and representing graphs in terms of their grammatical components (nodes, edges, cycles, and their relationships), the system reduces the computational complexity of classification while preserving the essential characteristics needed for accurate classification.

Inventive Principle:
Principle #2Taking out (Extraction)

3Manufacturing precision

If grammars are manually created for design representation, then accurate representation can be achieved, but the process is time-consuming and computationally intensive

Engineering Contradiction:
Improvegrammar representation accuracyVSAvoidgrammar induction time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the system to automatically induce grammars from graph data without manual intervention. The algorithm autonomously extracts grammatical rules, identifies patterns, and generates the complete grammar representation, eliminating the time-consuming manual process while maintaining the accuracy that would otherwise require expert knowledge and extensive time to achieve.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary extraction of grammatical rules from the graph structure before classification is performed. By pre-computing the grammatical representation through automated rule induction, the system prepares the data in an optimized format that accelerates subsequent classification operations, reducing overall processing time while ensuring accurate representation through systematic extraction of structural patterns.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11899669B2Searching of data structures in pre-processing data for a machine learning classifier
Publication Date: 2024.02.13 CARNEGIE MELLON UNIV
  • US11899669B2 patent drawing
  • US11899669B2 patent drawing
  • US11899669B2 patent drawing

AI summary

A data processing system is configured to pre-process data for a machine learning classifier. The data processing system includes an input port that receives one or more data items, an extraction engine that extracts a plurality of data signatures and structure data, a logical rule set generation engine configured to generate a data structure, select a particular data signature of the data structure, identify each instance of the particular data signature in the data structure, segment the data structure around instances of the particular data signature, identify one or more sequences of data signatures connected to the particular data signature, and generate a logical ruleset. A classification engine executes one or more classifiers against the logical ruleset to classify the one or more data items received by the input port.