Dynamic Linguistic Assessment for Text Mining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cognitive systems, such as those used in natural language processing, are inherently non-deterministic, leading to inconsistent results due to their susceptibility to input data, and conventional text mining techniques require rebuilding indices for new words, which is inefficient.

Innovation Solution

A system with a knowledge engine that applies linguistic algorithms to form cluster representations and identifies associative relationships using a nearness factor, allowing for dynamic linguistic assessment and measurement without the need for re-indexing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional text mining techniques are used to add new words to the dictionary, then the text mining system can incorporate new linguistic terms, but the associated index must be rebuilt which reduces efficiency and increases processing time

Engineering Contradiction:
Improveability to incorporate new linguistic termsVSAvoidtime required for re-indexing
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system pre-computes and stores linguistic relationships, proximity scores, and cluster associations in advance. When new words are added to the dictionary, the system leverages pre-established linguistic patterns and relationships rather than performing complete re-indexing, significantly reducing the time required to incorporate new terminology while maintaining comprehensive text mining capabilities

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The index structure is divided into modular components that can be independently updated. Instead of rebuilding the entire index when adding new words, only the specific segments related to the new linguistic terms are updated, allowing the system to incorporate new vocabulary efficiently while preserving the performance of existing indexed content

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If cognitive systems process natural language based on acquired knowledge, then the system can provide intelligent language understanding, but the results are non-deterministic and inconsistent due to susceptibility to input data variations

Engineering Contradiction:
Improvelanguage understanding capabilityVSAvoidconsistency of processing results
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system incorporates feedback mechanisms that monitor and adjust processing based on established linguistic patterns and proximity relationships. By continuously referencing pre-computed linguistic relationships and cluster associations, the system maintains consistent behavior across varying inputs while preserving adaptive language understanding capabilities

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system transforms unstructured natural language input into structured representations with defined parameters such as proximity scores, cluster memberships, and linguistic relationship metrics. This parameterization creates deterministic processing pathways that yield consistent results while maintaining the ability to understand diverse language inputs

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If a linguistic algorithm is applied to form cluster representations, then the system can identify linguistically related elements, but the process requires significant computational resources and time without optimization

Engineering Contradiction:
Improvelinguistic relationship identification accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system pre-computes linguistic relationships, proximity scores, and cluster associations between terms and stores them in optimized data structures. When text mining is performed, the system queries these pre-computed relationships rather than calculating them in real-time, maintaining high linguistic relationship identification accuracy while dramatically improving processing speed and reducing computational resource requirements

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates and maintains copies of linguistic relationship data in optimized formats suitable for rapid querying. By storing pre-computed proximity scores and cluster associations in accessible memory structures, the system avoids repeated computational operations while preserving accurate linguistic relationship identification

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11361031B2Dynamic linguistic assessment and measurement
Publication Date: 2022.06.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11361031B2 patent drawing
  • US11361031B2 patent drawing
  • US11361031B2 patent drawing

AI summary

Embodiments are directed to a system, a computer program product, and a method for identification of linguistically related elements, and more specifically to prediction of a linguistically related element. A linguistic algorithm forms a cluster representation of corpus entries. A linguistic term is identified and applied to the cluster representation to identify proximally related linguistic terms. Associative relationships between the proximally related terms and category metadata are iteratively investigated. One or more linguistic terms related across the two more metadata categories is identified and designated as the linguistically related element.