Data Field Extraction Model Training for Log Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data systems face challenges in efficiently analyzing and searching massive quantities of diverse machine data due to the lack of user-friendly tools for visual identification of data subsets, particularly in large-scale IT environments.
Innovation Solution
A data intake and query system that utilizes a flexible schema for extracting information from events, allowing for late-binding schema application during search time, and employs a pipelined search language to process and index data for efficient querying and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data systems store massive quantities of raw data for later retrieval and analysis, then data analysis flexibility and insight capability are improved, but data search and analysis efficiency deteriorate due to the lack of user-friendly tools for visual identification of data subsets
Solution Approach 1:
The patent segments the data analysis process into multiple stages: data ingestion, schema application during search, and visual identification of data subsets. This allows raw data to be stored in its entirety while enabling efficient retrieval through staged processing and visual tools that break down complex search tasks into manageable steps.
Solution Approach 2:
The patent introduces an intermediary processing layer that sits between raw data storage and user analysis. This layer includes schema application mechanisms and visual identification tools that translate user-friendly search criteria into efficient data retrieval operations, mediating between the need for data flexibility and search efficiency.
2Productivity
If pre-processing is applied to extract specified data items for efficient retrieval, then data retrieval efficiency is improved, but data analysis flexibility deteriorates because only a fraction of generated data can be analyzed
Solution Approach 1:
The patent implements a dynamic schema application approach where the data retrieval strategy adapts based on user needs and search context. Rather than static pre-processing, the system applies schemas dynamically during search operations, allowing efficient retrieval to be adjusted and optimized for different analysis requirements while maintaining access to all raw data.
Solution Approach 2:
The patent performs preliminary data ingestion and storage of complete raw data sets before analysis occurs. This preliminary action ensures all data is available for any analysis need, while subsequent schema application and visual identification tools provide the efficiency needed for specific retrieval operations without sacrificing flexibility.
3Quantity of substance
If tools are provided for searching data systems separately and collecting results over a network, then data collection capability is improved, but ease of operation deteriorates because analysts cannot quickly search and analyze large sets of raw machine data to visually identify data subsets of interest
Solution Approach 1:
The patent implements a nested structure where visual identification tools are embedded within the search interface, which itself is nested within the data retrieval system. This nested arrangement allows users to progressively drill down from high-level data subsets to specific data items, with each layer providing user-friendly controls while maintaining access to the full data collection capability.
Solution Approach 2:
The patent enables self-service data subset identification through visual tools that automatically analyze and present data patterns without requiring complex query formulation. The system serves itself by automatically applying schemas, identifying relevant data subsets, and presenting results in visually intuitive formats, making the system easy to operate while maintaining comprehensive data collection capabilities.
Data Source
AI summary
Systems and methods are described for training an artificial intelligence model to extract one or more data fields from a log. For example, the artificial intelligence model may be a neural network. The neural network may be trained using training data obtained by iterating through a plurality of logs using active learning, and selecting a subset of the logs in the plurality to be labeled by a user. For example, the selected subset of logs may be logs that are not similar to other logs already labeled by a user. The user may be prompted to label the selected subset of logs to identify one or more data fields to extract. Once the selected subset of logs are labeled, these labeled logs can be used as the training data to train the neural network.


