Security Ontology Evolution via LSTM Crawling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current generic search engines lack specific domain-focused security information, suffer from ambiguity, irrelevance, and credibility issues, and existing solutions are either expensive or limited in scope, with manual seed URL identification failing to ensure comprehensive subdomain coverage.

Innovation Solution

A method and system that extracts seed URLs from social media using keywords, crawls security-related content, classifies it into subdomains, extracts text with acronyms, evolves a security ontology using LSTM deep learning, and ranks results by credibility to provide comprehensive security information on vulnerabilities, threats, incidents, and controls to security experts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a generic search engine is used to access security information, then a single site for information retrieval is provided, but the information suffers from ambiguity, irrelevancy, bias and credibility issues

Engineering Contradiction:
Improvesingle site accessVSAvoidinformation credibility
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the search system into specialized modules: a crawler module that extracts seed URLs from social media, a classification module that categorizes content into security subdomains, and a ranking module that evaluates credibility based on domain relevance. This segmentation allows the system to maintain single-site accessibility while improving information reliability through specialized processing at each stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by creating security-specific search functionality within the general search engine framework. The credibility ranking is calculated specifically for security-related URLs using domain relevance metrics, while other parts of the search engine remain generic. This allows credibility improvement for security information without requiring complete system redesign.

Inventive Principle:
Principle #3Local quality

2Reliability

If expensive security information providers such as AlientVault, Symantec DeepSight, and IBM x-Force are used, then comprehensive threat information is obtained, but the cost is very high (∼150K USD per annum) and coverage is limited only to threat information

Engineering Contradiction:
Improvethreat information qualityVSAvoidinformation coverage scope
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal security search system that can handle multiple types of security information including vulnerabilities, threats, incidents, and controls through a single platform. The classification module categorizes content into various security subdomains, making the system versatile for different security information needs without requiring multiple specialized paid services.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs automated crawling, extraction, and classification algorithms that self-serve the information gathering function. Instead of relying on expensive proprietary databases, the system automatically extracts seed URLs from social media, crawls relevant content, and classifies it into security categories, providing comprehensive coverage at lower cost.

Inventive Principle:
Principle #25Self-service

3Ease of manufacture

If manual seed URL identification is used, then some security information sources are identified, but subdomain coverage is not assured and the process is time-consuming

Engineering Contradiction:
Improveseed URL identificationVSAvoidsubdomain coverage
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent replaces the manual mechanical process of seed URL identification with an automated computational system. The crawler module uses algorithms to automatically extract seed URLs from social media platforms based on security-related keywords, eliminating manual effort while ensuring comprehensive subdomain coverage through systematic web crawling and classification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Quantity of substance

If automated crawling and classification is implemented, then comprehensive security information is gathered across multiple subdomains, but system complexity increases

Engineering Contradiction:
Improvesecurity information coverageVSAvoidsystem architecture
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent divides the complex information processing system into distinct modular components: a crawler module for URL extraction, a classification module for content categorization into security subdomains, and a ranking module for credibility assessment. This segmentation manages system complexity by making each component independently manageable while achieving comprehensive security information coverage collectively.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11489859B2System and method for retrieving and extracting security information
Publication Date: 2022.11.01 INT INST OF INFORMATION THCHNOLOGY HYDERABAD
  • US11489859B2 patent drawing
  • US11489859B2 patent drawing
  • US11489859B2 patent drawing

AI summary

A system and method for retrieving and extracting security information is provided. The method includes (i) extracting seed Uniform Resource Locators (URLs) from social media based on keywords that are identified for each sub-domain, (ii) crawling a security related content in the extracted seed URLs to determine relevant URLs that are related to a security domain from the extracted seed URLs, (iii) classifying the security related content into sub-domains of security to obtain domain coverage, (iv) extracting text that include acronyms from the relevant URLs, (v) automatically evolving a security ontology based on extracted text using a Long Short-Term Memory (LSTM) deep Learning model, (vi) ranking search results by accessing credibility of the URLs that include the security related content based on domain relevance and (vii) providing the ranked search results that includes trends.