Security Ontology Evolution via LSTM Crawling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current generic search engines lack specific domain-focused security information, suffer from ambiguity, irrelevance, and credibility issues, and existing solutions are either expensive or limited in scope, with manual seed URL identification failing to ensure comprehensive subdomain coverage.
Innovation Solution
A method and system that extracts seed URLs from social media using keywords, crawls security-related content, classifies it into subdomains, extracts text with acronyms, evolves a security ontology using LSTM deep learning, and ranks results by credibility to provide comprehensive security information on vulnerabilities, threats, incidents, and controls to security experts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a generic search engine is used to access security information, then a single site for information retrieval is provided, but the information suffers from ambiguity, irrelevancy, bias and credibility issues
Solution Approach 1:
The patent segments the search system into specialized modules: a crawler module that extracts seed URLs from social media, a classification module that categorizes content into security subdomains, and a ranking module that evaluates credibility based on domain relevance. This segmentation allows the system to maintain single-site accessibility while improving information reliability through specialized processing at each stage.
Solution Approach 2:
The patent applies local quality by creating security-specific search functionality within the general search engine framework. The credibility ranking is calculated specifically for security-related URLs using domain relevance metrics, while other parts of the search engine remain generic. This allows credibility improvement for security information without requiring complete system redesign.
2Reliability
If expensive security information providers such as AlientVault, Symantec DeepSight, and IBM x-Force are used, then comprehensive threat information is obtained, but the cost is very high (∼150K USD per annum) and coverage is limited only to threat information
Solution Approach 1:
The patent creates a universal security search system that can handle multiple types of security information including vulnerabilities, threats, incidents, and controls through a single platform. The classification module categorizes content into various security subdomains, making the system versatile for different security information needs without requiring multiple specialized paid services.
Solution Approach 2:
The system employs automated crawling, extraction, and classification algorithms that self-serve the information gathering function. Instead of relying on expensive proprietary databases, the system automatically extracts seed URLs from social media, crawls relevant content, and classifies it into security categories, providing comprehensive coverage at lower cost.
3Ease of manufacture
If manual seed URL identification is used, then some security information sources are identified, but subdomain coverage is not assured and the process is time-consuming
Solution Approach 1:
The patent replaces the manual mechanical process of seed URL identification with an automated computational system. The crawler module uses algorithms to automatically extract seed URLs from social media platforms based on security-related keywords, eliminating manual effort while ensuring comprehensive subdomain coverage through systematic web crawling and classification.
4Quantity of substance
If automated crawling and classification is implemented, then comprehensive security information is gathered across multiple subdomains, but system complexity increases
Solution Approach 1:
The patent divides the complex information processing system into distinct modular components: a crawler module for URL extraction, a classification module for content categorization into security subdomains, and a ranking module for credibility assessment. This segmentation manages system complexity by making each component independently manageable while achieving comprehensive security information coverage collectively.
Data Source
AI summary
A system and method for retrieving and extracting security information is provided. The method includes (i) extracting seed Uniform Resource Locators (URLs) from social media based on keywords that are identified for each sub-domain, (ii) crawling a security related content in the extracted seed URLs to determine relevant URLs that are related to a security domain from the extracted seed URLs, (iii) classifying the security related content into sub-domains of security to obtain domain coverage, (iv) extracting text that include acronyms from the relevant URLs, (v) automatically evolving a security ontology based on extracted text using a Long Short-Term Memory (LSTM) deep Learning model, (vi) ranking search results by accessing credibility of the URLs that include the security related content based on domain relevance and (vii) providing the ranked search results that includes trends.


