Feature String Extraction Using Markov Transition Entropy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting feature strings in network traffic identification require significant manual intervention and resources, failing to provide an automated solution that adapts to the rapid updating cycle of applications.
Innovation Solution
A method utilizing a first-order Markov transition probability matrix and genetic variation coefficient to automatically extract feature strings from data packets, determining transition entropy and identifying usable feature strings without manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data packet analysis method is used to extract feature strings, then extraction accuracy can be maintained through expert analysis, but a large amount of time and manual resources are required
Solution Approach 1:
The system performs self-service by automatically extracting feature strings through probabilistic models and entropy calculations without requiring manual expert intervention. The apparatus autonomously completes session filtration, candidate generation, and feature string extraction that previously required manual analysis
Solution Approach 2:
Manual mechanical analysis is replaced with automated computational systems using Markov transition probability matrices and entropy-based algorithms. The system substitutes human expert analysis with mathematical models that automatically identify and extract feature strings from network traffic
2Reliability
If traditional analysis techniques are used to maintain feature strings, then identification reliability can be ensured, but frequent manual updates are required due to application updates
Solution Approach 1:
The system implements dynamic adaptation by continuously monitoring network traffic and automatically updating feature strings in response to application changes. Instead of static manual updates, the system dynamically adjusts feature strings based on real-time traffic analysis and entropy-based validation
Solution Approach 2:
The system uses feedback mechanisms where extracted feature strings are validated against traffic patterns and identification results. This feedback loop ensures reliability by confirming that extracted features actually work for identification while reducing manual update frequency through automated validation
3Extent of automation
If semi-automatic scripting method is used for feature string extraction, then some automation is achieved, but manual intervention is still required for traffic triggering and feature screening
Solution Approach 1:
The system achieves complete self-service by automating all previously manual steps including traffic triggering, session filtration, candidate generation, and feature string validation. No manual intervention is required as the system autonomously performs the entire extraction pipeline
4Loss of information
If application disassembling analysis method is used to extract feature strings, then deep understanding of application protocol can be achieved, but a lot of manpower and material resources are required
Solution Approach 1:
Manual protocol analysis is replaced with automated disassembling and analysis systems that use probabilistic models to understand application protocols. The system substitutes human analysts with computational algorithms that perform disassembling, candidate generation, and validation without requiring expert manual intervention
Data Source
AI summary
Disclosed are a method for extracting a feature string, a device, a network apparatus, and a storage medium. The method comprises: determining, for each candidate feature string, a transition probability for each pair of adjacent characters in the candidate feature string according to a first-order Markov transition probability matrix; detennining a transition entropy value of the candidate feature string according to the transition probability of each pair of adjacent characters and the logarithm of the transition probability; and labeling a candidate feature string having a transition entropy value greater than a pre-determined threshold as a first usable feature string, and using a valid first usable feature string as an extracted target feature string.


