N-gram String Search for Non-Delimited Language Content Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional content filtering mechanisms struggle with keyword search in non-delimited languages like Chinese, Japanese, and Thai, as they lack spaces between words, making real-time filtering inefficient and time-consuming due to the reliance on word delimiters.
Innovation Solution
The implementation of an efficient string search method using N-grams, where a finite state machine is constructed to search for pre-selected keyword sequences (N-grams) in a string of bytes, allowing simultaneous search without relying on word delimiters, thus enabling efficient content classification and filtering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional keyword search methods are used in non-delimited languages, then the search can be performed using standard tokenization approaches, but the processing time increases significantly and real-time filtering becomes difficult
Solution Approach 1:
The patent segments the keyword search problem into N-gram units (sequences of N characters or bytes) that can be independently identified and matched. By breaking down the search task into discrete N-gram segments rather than attempting to tokenize entire words, the system can efficiently search for keywords in non-delimited languages without requiring word boundary delimiters, thus reducing processing time while maintaining search reliability
Solution Approach 2:
The patent performs preliminary action by pre-processing and indexing N-grams from the text content before the actual keyword search is needed. N-grams are extracted, normalized, and stored in an inverted index structure in advance, allowing the search engine to quickly retrieve and count matching N-grams without performing complex tokenization during the search phase, thereby reducing real-time processing time
2Productivity
If word delimiter-based tokenization is used, then the search process is simple and efficient for delimited languages, but it fails to correctly identify words in non-delimited languages like Chinese, Japanese, and Thai
Solution Approach 1:
Instead of relying on space delimiters to segment words, the patent segments text into N-grams (sequences of N characters or bytes). This segmentation approach works universally for both delimited and non-delimited languages, maintaining tokenization efficiency while improving word boundary recognition accuracy by not depending on the absence of delimiters
Solution Approach 2:
The patent changes the fundamental parameter of tokenization from space-delimiter-based word segmentation to fixed-length or variable-length N-gram segmentation. By changing the segmentation parameter from relying on delimiter presence to fixed character/byte sequences, the system achieves both high productivity (simple fixed-length segmentation) and high reliability (accurate keyword representation in non-delimited languages)
3Measurement precision
If real-time content filtering is implemented with comprehensive keyword search, then content classification accuracy improves, but the filtering speed decreases causing noticeable delays
Solution Approach 1:
The patent performs preliminary action by pre-processing text into N-grams and building inverted indexes before filtering is needed. This pre-computation stores the mapping between N-grams and their positions in the text, allowing the filtering system to quickly count keyword occurrences and classify content without performing complex analysis in real-time, thus maintaining both accuracy and speed
Solution Approach 2:
The patent substitutes the mechanical tokenization process (which requires sequential word boundary detection) with a direct N-gram matching mechanism using inverted indexes. This substitution replaces the complex mechanical process of word segmentation with a simpler lookup-based system that counts N-gram occurrences directly, improving filtering speed while maintaining classification accuracy through comprehensive keyword matching
Data Source
AI summary
Some embodiments of an efficient string search have been presented. In one embodiment, a string of bytes representing content written in a non-delimited language is received, wherein the content has been classified into a predetermined category. In a single pass through the string of bytes, a set of N-grams is searched for simultaneously. Statistical information on occurrences of the N-grams, if any, in the string of bytes is collected. In some embodiments, a model is generated based on the statistical information, where the model is usable by a content filter to classify content.


