Automated Web Content Analysis System for Prohibited Material Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual content analysis of web sites for prohibited content is labor-intensive, slow, and ineffective due to cognitive limitations and inability to process multiple linguistic connections and long chains of connected texts, with keyword filtering being insufficient.
Innovation Solution
An automated system for web site content analysis that searches for terms using a dictionary, performs multi-factor genre content analysis, and uses a thematic rubricator to determine the activity and purpose of web resources, replicating human thought processes by combining thematic, pragmatic, and lexical properties to differentiate between propaganda and information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual content analysis is used, then human judgment and context understanding are applied, but the analysis speed is slow and cannot keep up with site update rates
Solution Approach 1:
The patent replaces manual human analysis with an automated computer system that uses algorithms to detect prohibited content. The system automatically crawls web pages, extracts text, and applies detection rules without human intervention, thereby maintaining accuracy while dramatically increasing processing speed and keeping pace with rapid site updates.
Solution Approach 2:
The system performs self-service analysis by automatically monitoring and detecting prohibited content on websites without requiring human administrators to manually review each site. The automated detection system continuously scans for banned words and patterns, enabling continuous monitoring at high speed.
2Productivity
If keyword filtering is used, then simple detection is achieved, but it cannot distinguish between proper and improper usage of words
Solution Approach 1:
The patent applies local quality analysis by examining the contextual environment of keywords rather than treating them in isolation. The system analyzes the surrounding text, sentence structure, and semantic relationships to determine whether a keyword is used properly or improperly, allowing distinction between legitimate and prohibited content usage.
Solution Approach 2:
The system changes the detection parameters from simple keyword matching to multi-dimensional analysis including contextual parameters, semantic parameters, and syntactic parameters. This enables the system to evaluate word usage in different contexts and distinguish between acceptable and prohibited meanings.
3Measurement precision
If manual analysis is performed, then contextual understanding is achieved, but the administrator cannot process multiple linguistic connections and long chains of connected texts
Solution Approach 1:
The patent segments the content analysis process into independent modules: text extraction, keyword detection, contextual analysis, and decision-making. Each module processes specific aspects of the text independently, allowing the system to handle complex linguistic connections and long text chains by breaking them down into manageable analysis units.
Solution Approach 2:
The automated system provides multi-functional capabilities including text crawling, keyword detection, contextual analysis, semantic understanding, and automated decision-making. This universal approach allows a single system to handle diverse content types and complex linguistic relationships that would overwhelm manual analysis.
Data Source
AI summary
A system and method for an automated web source content analysis. The system of automated content analysis performs the following: a search of terms, i.e. key words and phrases, presented in the special dictionary, in the text content; executes a multi-factor genre content analysis based on structural, pragmatic and stylistics properties; executes thematic content analysis using a rubricator built based on illegal subjects and topics and their antagonists; and the system makes a decision based on a combination of thematic and genre properties of the text. The proposed method allows for providing a final decision in terms that are easily understood by a user.


