Text Disabbreviation via Vector Space Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text data processing techniques in communications systems face challenges in accurately disabbreviating text units, such as abbreviations, due to high rates of false positives and negatives, especially when dealing with non-canonical text units.
Innovation Solution
A method involving a communications server apparatus that compares text data elements with a database to determine similarity measures, selects candidate text units with ordered relationships, and uses these measures to nominate a disabbreviated text unit, incorporating vector space models and frequency analysis to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If rudimentary techniques are used for comparing text data elements, then device complexity is reduced, but measurement precision deteriorates due to high false positive and negative rates
Solution Approach 1:
The patent transforms the text comparison problem from simple string matching to vector space distance measurement. By converting text data elements and database entries into vectors and comparing them using distance metrics, the system achieves higher measurement precision while maintaining reasonable computational complexity. This parameter transformation from discrete symbols to continuous vector representations resolves the contradiction between simplicity and accuracy.
2Measurement precision
If complex techniques are used for comparing text data elements, then measurement precision improves, but device complexity increases
Solution Approach 1:
The patent replaces complex mechanical or algorithmic text comparison systems with a vector space model approach. Instead of using intricate rule-based systems or multiple processing stages, the invention substitutes a unified mathematical framework based on vector representations and distance measurements, achieving high precision with more elegant and manageable complexity.
3Measurement precision
If frequency analysis is incorporated into the comparison process, then measurement precision improves through better candidate selection, but use of energy increases due to additional processing
Solution Approach 1:
The patent performs frequency analysis and vector space model construction as preliminary actions during database preparation, rather than during each text comparison operation. By pre-computing frequency statistics and storing them in the database, the system reduces the computational energy required during actual text processing while maintaining high measurement precision through informed candidate selection.
Data Source
AI summary
A communications server apparatus (100) is configured to receive (202) text data comprising at least one text data element associated with an abbreviated text unit. The text data element is compared (204) with a plurality of candidate text data elements from a representation of a given text database, each candidate text data element associated with a respective candidate text unit in the database. Values for a similarity measure between the at least one text data element and the candidate text data elements are determined (206), and candidate text data elements are processed (208) to select candidate text data elements with associated candidate text units having an ordered relationship with the abbreviated text unit. The similarity measure values and the candidate text data element selections are used (210) to nominate an associated candidate text unit as a disabbreviated text unit for the abbreviated text unit.


