Automated Diaculture Identification via Gram Tokenization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for identifying diaculture in text data rely heavily on manual keyword selection and require skilled analysts, making it time-consuming and inefficient, as they focus on topic-specific words rather than cultural indicators.

Innovation Solution

A method utilizing tokenization and gram construction to compare text against training data sets, assigning scores based on grammatical types and context, allowing for automated identification of diaculture by distinguishing between topic-centric and non-topic-centric words.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual keyword selection and hand selected communications channels are used to identify diaculture data streams, then the analyst can identify specific cultural backgrounds, but it requires skilled analysts, takes time, and relies on innate abilities and experience

Engineering Contradiction:
Improvediaculture identification accuracyVSAvoidtime required for analysis
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical analysis with an automated computational system. The processor automatically tokenizes text, constructs grams, compares them against training data sets, and assigns scores to identify diaculture, eliminating the need for manual keyword selection and skilled analyst intervention while maintaining identification accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically learning from training data sets and applying the learned patterns to new text without requiring human expertise. The automated process independently completes the entire analysis workflow from tokenization to diaculture identification, freeing analysts from time-consuming manual tasks

Inventive Principle:
Principle #25Self-service

2Measurement precision

If hand selected keywords are used for searching diaculture data, then specific cultural backgrounds can be targeted, but it requires skilled analysts to identify critical combinations of keywords or phrases

Engineering Contradiction:
Improvediaculture identification accuracyVSAvoidcomplexity of keyword combination identification
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the text processing into distinct automated steps: tokenization, gram construction, comparison against training data, and scoring. This segmentation eliminates the complex manual task of identifying critical keyword combinations by breaking down the analysis into systematic computational operations that automatically handle all combinations

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system substitutes manual expert judgment with automated computational algorithms that systematically generate and evaluate all possible gram combinations against trained models, eliminating the need for skilled analysts to manually identify critical keyword phrases while maintaining or improving identification precision

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If topic-centric words are focused on for analysis, then the topic of the text can be identified, but cultural background identification is hindered because topic words are replaced with tokens

Engineering Contradiction:
Improvecultural background identification accuracyVSAvoidtopic information loss
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments word analysis into two distinct categories: topic-centric words (verbs, nouns, adverbs, adjectives) that are replaced with grammatical tokens to remove topic bias, and non-topic-centric words (pronouns, articles, prepositions) that are retained to preserve cultural linguistic patterns. This segmentation allows simultaneous achievement of topic independence and cultural signal preservation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities to different word types: topic words receive transformation (replacement with tokens) to eliminate topic contamination, while non-topic words retain their original form to preserve cultural diacultural markers. This local differentiation ensures that cultural identification is not hindered by topic-specific vocabulary while maintaining topic awareness through the tokenized grammatical structure

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9158761B2Identifying cultural background from text
Publication Date: 2015.10.13 LEIDOS INNOVATIONS TECHNOLOGY INC
  • US9158761B2 patent drawing
  • US9158761B2 patent drawing
  • US9158761B2 patent drawing

AI summary

Diaculture of text can be determined or analyzed by tokenizing words of the text according to a rule set to generate tokenized text, the rule set defining: a first set of grammatical types of words, which are words that are replaced with tokens that respectively indicate a grammatical type of a respective word, and a second set of grammatical types of words, which are words that are passed as tokens without changing. Grams can be constructed from the tokenized text, each gram including one or more of consecutive tokens from the tokenized text. The grams can be compared to a training data set that corresponds to a known diaculture to obtain a comparison result that indicates how well the text matches the training data set for the known diaculture.