Data Field Extraction Model Training for Log Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data systems face challenges in efficiently analyzing and searching massive quantities of diverse machine data due to the lack of user-friendly tools for visual identification of data subsets, particularly in large-scale IT environments.

Innovation Solution

A data intake and query system that utilizes a flexible schema for extracting information from events, allowing for late-binding schema application during search time, and employs a pipelined search language to process and index data for efficient querying and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data systems store massive quantities of raw data for later retrieval and analysis, then data analysis flexibility and insight capability are improved, but data search and analysis efficiency deteriorate due to the lack of user-friendly tools for visual identification of data subsets

Engineering Contradiction:
Improvedata analysis flexibilityVSAvoiddata search efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the data analysis process into multiple stages: data ingestion, schema application during search, and visual identification of data subsets. This allows raw data to be stored in its entirety while enabling efficient retrieval through staged processing and visual tools that break down complex search tasks into manageable steps.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that sits between raw data storage and user analysis. This layer includes schema application mechanisms and visual identification tools that translate user-friendly search criteria into efficient data retrieval operations, mediating between the need for data flexibility and search efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If pre-processing is applied to extract specified data items for efficient retrieval, then data retrieval efficiency is improved, but data analysis flexibility deteriorates because only a fraction of generated data can be analyzed

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoiddata analysis flexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic schema application approach where the data retrieval strategy adapts based on user needs and search context. Rather than static pre-processing, the system applies schemas dynamically during search operations, allowing efficient retrieval to be adjusted and optimized for different analysis requirements while maintaining access to all raw data.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent performs preliminary data ingestion and storage of complete raw data sets before analysis occurs. This preliminary action ensures all data is available for any analysis need, while subsequent schema application and visual identification tools provide the efficiency needed for specific retrieval operations without sacrificing flexibility.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If tools are provided for searching data systems separately and collecting results over a network, then data collection capability is improved, but ease of operation deteriorates because analysts cannot quickly search and analyze large sets of raw machine data to visually identify data subsets of interest

Engineering Contradiction:
Improvedata collection capabilityVSAvoiduser-friendliness of search tools
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent implements a nested structure where visual identification tools are embedded within the search interface, which itself is nested within the data retrieval system. This nested arrangement allows users to progressively drill down from high-level data subsets to specific data items, with each layer providing user-friendly controls while maintaining access to the full data collection capability.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent enables self-service data subset identification through visual tools that automatically analyze and present data patterns without requiring complex query formulation. The system serves itself by automatically applying schemas, identifying relevant data subsets, and presenting results in visually intuitive formats, making the system easy to operate while maintaining comprehensive data collection capabilities.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11663176B2Data field extraction model training for a data intake and query system
Publication Date: 2023.05.30 CISCO TECHNOLOGY INC
  • US11663176B2 patent drawing
  • US11663176B2 patent drawing
  • US11663176B2 patent drawing

AI summary

Systems and methods are described for training an artificial intelligence model to extract one or more data fields from a log. For example, the artificial intelligence model may be a neural network. The neural network may be trained using training data obtained by iterating through a plurality of logs using active learning, and selecting a subset of the logs in the plurality to be labeled by a user. For example, the selected subset of logs may be logs that are not similar to other logs already labeled by a user. The user may be prompted to label the selected subset of logs to identify one or more data fields to extract. Once the selected subset of logs are labeled, these labeled logs can be used as the training data to train the neural network.