Linguistic Analysis System Using Hierarchical Phrase Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer methods for linguistic analysis struggle to effectively understand and translate human languages, particularly failing to handle idiomatic words and phrases, and require lengthy and complex rule-based programming.
Innovation Solution
A method that splits input text into words and sentences, compares phrases with stored phrases using a combination of hierarchical pattern storage and bidirectional pattern matching, allowing for the conversion of text into its constituent grammatical parts, and includes error correction and word sense disambiguation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If rule-based systems are used for linguistic analysis, then the system can follow structured grammar rules, but the programming becomes lengthy and complex
Solution Approach 1:
The patent replaces the mechanical rule-based programming system with a neural network-based cognitive system. Instead of manually coding grammar rules, the system uses trained neural networks that automatically learn and apply linguistic patterns, thereby reducing programming complexity while maintaining analysis reliability
Solution Approach 2:
The neural network system performs self-learning from training data without requiring explicit programming of linguistic rules. The system automatically acquires grammar and semantic knowledge through exposure to language samples, eliminating the need for lengthy manual rule creation
2Productivity
If statistical methods based on word sequence likelihood are used, then translation can be performed, but idiomatic words and phrases are not handled effectively
Solution Approach 1:
The patent combines multiple neural network components with different functions: one network handles statistical word sequence analysis for general translation, while another network specifically processes idiomatic expressions and phrases. This composite approach allows the system to leverage both statistical methods and specialized pattern recognition for comprehensive language understanding
Solution Approach 2:
The system applies different processing strategies to different parts of the text based on their characteristics. Standard words are processed using statistical methods, while identified idiomatic expressions are routed to specialized processing pathways that handle figurative language, ensuring each type of linguistic element receives appropriate treatment
3Reliability
If the longest stored phrase is matched first, then idiomatic phrases and names are matched ahead of grammatically-based phrases, but shorter phrases may be missed
Solution Approach 1:
The matching process is made dynamic and iterative rather than static and single-pass. The system performs multiple comparison passes, adjusting the matching strategy based on what was found in previous passes. This allows the system to capture both long idiomatic phrases and shorter grammatical phrases that may have been overlooked in earlier iterations
4Measurement precision
If bidirectional pattern matching is used to convert text to hierarchical patterns, then linguistic analysis accuracy is improved, but processing time increases
Solution Approach 1:
The system performs preliminary processing steps before the main bidirectional matching: text is pre-tokenized, basic grammatical structures are identified in advance, and common patterns are pre-loaded. This preliminary preparation reduces the computational burden during the intensive bidirectional matching phase, thereby reducing overall processing time while maintaining accuracy
Data Source
AI summary
A method of operating a computer to perform linguistic analysis includes the steps of splitting an input text into words and sentences; for each sentence, comparing phrases in the sentence with known phrases stored in a database, as follows: for each word in the sentence, comparing its value and values of words following it with values of words of stored phrases, starting with the longest stored phrase that starts with that word, and working from longest to shortest; in the event a match is found for two or more consecutive words, and considering the words around the phrase, labelling the matched phrase with an overphrase that describes the grammar use of the matched phrase; after the penultimate word has been compared, recasting the sentence by replacing the matched phrases by their respective overphrases; and then repeating the comparison process with the recast sentence until there is no further recasting.


