AI-Indexed File Search for Non-Aggregated Cyber Event Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data collection systems for cyber events lack scalability, accuracy, and integration of non-aggregated data sources, leading to incomplete and unsearchable historical records of cyber incidents, which are crucial for corporate risk assessment and supply chain security.

Innovation Solution

A method and system using AI-driven Named Entity Recognition (NER) and machine learning to process and index non-aggregated data from various sources, including darknets, corporate sources, and official channels, enriching data with metadata and correlating entities for comprehensive search and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is collected from multiple non-aggregated sources manually, then data coverage is improved, but productivity deteriorates due to manual collection requirements

Engineering Contradiction:
Improvedata coverageVSAvoiddata collection efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent replaces manual mechanical data collection with automated web scraping systems that use software agents to systematically extract data from multiple non-aggregated sources including darknets, corporate sources, and official channels, thereby maintaining comprehensive data coverage while dramatically improving productivity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system implements self-service automation where the scraping infrastructure automatically collects, processes, and indexes data from diverse sources without requiring manual intervention for each data collection task, enabling continuous operation across thousands of organizations

Inventive Principle:
Principle #25Self-service

2Measurement precision

If AI-based NER processing is applied to extract entities, then measurement precision is improved, but device complexity worsens due to AI processing requirements

Engineering Contradiction:
Improveentity extraction accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the data processing pipeline into distinct stages: data collection, AI-based NER processing for entity extraction, metadata enrichment, and indexing. This segmentation allows the complex AI processing to be isolated to specific stages while maintaining overall system manageability and enabling parallel processing of different data streams

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If data is aggregated and indexed for search, then search precision is improved, but loss of information worsens due to aggregation requirements

Engineering Contradiction:
Improvesearch precisionVSAvoiddata detail loss
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent implements a nested data structure where aggregated index files contain references to and excerpts from original non-aggregated source files. This allows the system to provide efficient search capabilities through aggregation while preserving access to the complete detailed information in the source files, effectively nesting the aggregated view within the context of the original data

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentEP4579476A1Method and system for file search and method and system for creating index files
Publication Date: 2025.07.02 DIGITAL INTELLIGENCE LAB SRL
  • EP4579476A1 patent drawingFigure 1
  • EP4579476A1 patent drawingFigure 2
  • EP4579476A1 patent drawingFigure 3~4

AI summary

The invention pertains to a method for creating a computer-executable index file for indexing and searching non-aggregated files, wherein: during the phase of integrating and processing first contents to transform them into second contents, extractions are performed, and the extracted entities are used for at least one of the following phases: - correlating a first extracted entity to the second contents; - using the first extracted entity as a search key in external informational databases and, if information is found, enriching the second contents; - correlating the second contents with the official website of the first extracted entity; - correlating the TLD portion of an extracted domain name associated with the first extracted entity with the second contents through metadata related to the nation of the first extracted entity.