Automated Structuring of Free Form Heterogeneous Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing approaches to structuring free form heterogeneous data in technical helpdesk systems are inefficient due to the unstructured and noisy nature of ticketing data, leading to difficulties in searching for relevant problem resolutions.

Innovation Solution

The method involves segmenting and labeling free form data using machine learning techniques to structure it in a format that facilitates IT operations, employing supervised and semi-supervised learning algorithms to identify key information structures and generate annotation models for automatic labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If free form heterogeneous data is stored in unstructured format, then storage flexibility is maintained, but search efficiency and information retrieval accuracy deteriorate

Engineering Contradiction:
Improvesearch efficiencyVSAvoiddata structuring complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments free form heterogeneous data into distinct structured fields including problem description, resolution steps, product information, and environment details. This segmentation transforms unstructured text into organized data elements that can be efficiently queried and retrieved, directly improving search efficiency while maintaining manageable complexity through systematic categorization.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If manual labeling and structuring of ticket data is performed, then data quality and search accuracy improve, but processing time and labor costs increase

Engineering Contradiction:
Improveinformation retrieval accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements automatic labeling and structuring systems that enable the helpdesk data to self-organize into structured formats without extensive manual intervention. Machine learning algorithms and automated text processing techniques allow the system to independently categorize and structure data, maintaining high information retrieval accuracy while dramatically reducing the time and labor required compared to manual processing.

Inventive Principle:
Principle #25Self-service

3Ease of manufacture

If existing word-based heuristic rules are used for data classification, then implementation simplicity is maintained, but classification accuracy and relevance deteriorate due to noisy heterogeneous data

Engineering Contradiction:
Improvesystem implementation easeVSAvoiddata classification accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transitions from simple word-based heuristic rules to more sophisticated classification parameters including n-gram analysis, contextual understanding, and multiple feature extraction. This parameter enhancement allows the system to accurately classify noisy heterogeneous data by considering broader contextual patterns rather than relying solely on individual keyword matches, thereby improving classification accuracy while maintaining reasonable implementation complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8996587B2Method and apparatus for automatically structuring free form hetergeneous data
Publication Date: 2015.03.31 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8996587B2 patent drawing
  • US8996587B2 patent drawing
  • US8996587B2 patent drawing

AI summary

Techniques are provided for automatically structuring free form heterogeneous data. In one aspect of the invention, the techniques include obtaining free form heterogeneous data, segmenting the free form heterogeneous data into one or more units, automatically labeling the one or more units based on one or more machine learning techniques, wherein each unit is associated with a label indicating an information type, and structuring the one or more labeled units in a format to facilitate one or more operations that use at least a portion of the labeled units, e.g., information technology (IT) operations.