Mail Server Keyword Learning for Email Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing email filtering systems face challenges in maintaining accuracy and keeping up with changing keywords, relying on subjective administrator input and requiring significant time for monitoring, which can lead to missed harmful emails and increased administrative burden.
Innovation Solution
A mail server system that learns and updates keyword sets based on co-occurrence probability and clustering, automatically adding or deleting keywords to improve the accuracy of identifying and filtering out harmful emails, thereby adapting to changes in keyword usage over time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If administrators manually monitor and update keywords in email filtering systems, then extraction accuracy can be maintained, but administrative burden and time consumption increase significantly
Solution Approach 1:
The system performs self-learning by automatically analyzing email content and identifying harmful keywords without administrator intervention. The learning unit autonomously updates the keyword set based on analyzed results, eliminating the need for manual monitoring and maintenance while maintaining high extraction accuracy.
Solution Approach 2:
The system implements a feedback mechanism where analysis results are continuously fed back to update the keyword set. The learning unit uses the analyzed harmful emails and their characteristics to refine and expand the keyword database, creating a self-improving cycle that maintains accuracy without additional administrative time.
2Adaptability or versatility
If a fixed keyword set is used for email extraction, then system complexity is reduced, but the system cannot adapt to changing keyword usage over time
Solution Approach 1:
The keyword set transitions from a static, fixed list to a dynamic, self-updating database. The learning unit continuously analyzes email content and automatically adjusts the keyword set to reflect changing usage patterns and emerging harmful terms, enabling the system to adapt without manual intervention.
Solution Approach 2:
The system autonomously manages its own keyword database through self-learning mechanisms. The learning unit independently identifies new harmful keywords, updates the keyword set, and maintains the extraction conditions, reducing system complexity from the administrator's perspective while enhancing adaptability.
3Reliability
If comprehensive keyword monitoring is performed to catch all harmful emails, then extraction accuracy improves, but the administrative burden increases
Solution Approach 1:
The learning unit autonomously performs comprehensive analysis of email content to identify harmful keywords and patterns. This self-service approach maintains high detection reliability by continuously learning from new data while eliminating the administrative burden of manual keyword monitoring and updates.
Solution Approach 2:
The system implements continuous feedback loops where analysis results automatically update the keyword set. This feedback mechanism ensures comprehensive harmful email detection by constantly refining the extraction conditions based on newly identified patterns, without requiring additional administrative effort.
Data Source
AI summary
A mail server identifies first keyword set including a keyword that is not included in a second keyword set, the key word being a keyword that appear in mail data with a frequency higher than a predetermined frequency, the mail data being extracted based on the second keyword set including a keyword used in extraction conditions of the mail data. Then, the mail server adds the first keyword set to the extraction conditions of the mail data.


