N-gram String Search for Non-Delimited Language Content Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional content filtering mechanisms struggle with keyword search in non-delimited languages like Chinese, Japanese, and Thai, as they lack spaces between words, making real-time filtering inefficient and time-consuming due to the reliance on word delimiters.

Innovation Solution

The implementation of an efficient string search method using N-grams, where a finite state machine is constructed to search for pre-selected keyword sequences (N-grams) in a string of bytes, allowing simultaneous search without relying on word delimiters, thus enabling efficient content classification and filtering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional keyword search methods are used in non-delimited languages, then the search can be performed using standard tokenization approaches, but the processing time increases significantly and real-time filtering becomes difficult

Engineering Contradiction:
Improvekeyword search reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the keyword search problem into N-gram units (sequences of N characters or bytes) that can be independently identified and matched. By breaking down the search task into discrete N-gram segments rather than attempting to tokenize entire words, the system can efficiently search for keywords in non-delimited languages without requiring word boundary delimiters, thus reducing processing time while maintaining search reliability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by pre-processing and indexing N-grams from the text content before the actual keyword search is needed. N-grams are extracted, normalized, and stored in an inverted index structure in advance, allowing the search engine to quickly retrieve and count matching N-grams without performing complex tokenization during the search phase, thereby reducing real-time processing time

Inventive Principle:
Principle #10Preliminary action

2Productivity

If word delimiter-based tokenization is used, then the search process is simple and efficient for delimited languages, but it fails to correctly identify words in non-delimited languages like Chinese, Japanese, and Thai

Engineering Contradiction:
Improvetokenization efficiencyVSAvoidword boundary recognition accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

Instead of relying on space delimiters to segment words, the patent segments text into N-grams (sequences of N characters or bytes). This segmentation approach works universally for both delimited and non-delimited languages, maintaining tokenization efficiency while improving word boundary recognition accuracy by not depending on the absence of delimiters

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the fundamental parameter of tokenization from space-delimiter-based word segmentation to fixed-length or variable-length N-gram segmentation. By changing the segmentation parameter from relying on delimiter presence to fixed character/byte sequences, the system achieves both high productivity (simple fixed-length segmentation) and high reliability (accurate keyword representation in non-delimited languages)

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If real-time content filtering is implemented with comprehensive keyword search, then content classification accuracy improves, but the filtering speed decreases causing noticeable delays

Engineering Contradiction:
Improvecontent classification accuracyVSAvoidfiltering speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent performs preliminary action by pre-processing text into N-grams and building inverted indexes before filtering is needed. This pre-computation stores the mapping between N-grams and their positions in the text, allowing the filtering system to quickly count keyword occurrences and classify content without performing complex analysis in real-time, thus maintaining both accuracy and speed

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent substitutes the mechanical tokenization process (which requires sequential word boundary detection) with a direct N-gram matching mechanism using inverted indexes. This substitution replaces the complex mechanical process of word segmentation with a simpler lookup-based system that counts N-gram occurrences directly, improving filtering speed while maintaining classification accuracy through comprehensive keyword matching

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10460041B2Efficient string search
Publication Date: 2019.10.29 SONICWALL US HOLDINGS INC
  • US10460041B2 patent drawing
  • US10460041B2 patent drawing
  • US10460041B2 patent drawing

AI summary

Some embodiments of an efficient string search have been presented. In one embodiment, a string of bytes representing content written in a non-delimited language is received, wherein the content has been classified into a predetermined category. In a single pass through the string of bytes, a set of N-grams is searched for simultaneously. Statistical information on occurrences of the N-grams, if any, in the string of bytes is collected. In some embodiments, a model is generated based on the statistical information, where the model is usable by a content filter to classify content.