Domain Dictionary Construction via Inter-Character Correlation Index
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing domain dictionary constructing methods are inefficient and have limited usage conditions due to the need for manual annotation and high recognition accuracy, requiring significant accumulation of domain words.
Innovation Solution
A method and apparatus that segment a domain corpus sample to calculate inter-character correlation indices, determining character segments with high correlation to construct a domain dictionary without manual annotation, improving efficiency and accuracy by filtering and expanding domain words.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If statistical-based domain dictionary constructing method is used with manual annotation and model training, then domain word recognition accuracy can be improved, but construction efficiency deteriorates and usage conditions become limited
Solution Approach 1:
The patent enables the domain dictionary construction system to automatically identify and extract domain words from corpus samples without requiring manual annotation. The system uses unsupervised learning techniques where the model autonomously learns domain word patterns and performs self-correction, eliminating the need for human annotators while maintaining high recognition accuracy.
Solution Approach 2:
The patent replaces the mechanical manual annotation process with an automated computational system. Instead of relying on human experts to manually label domain words in corpus samples, the system uses natural language processing algorithms and machine learning models to automatically identify domain words, thereby dramatically improving construction efficiency while preserving accuracy.
2Measurement precision
If manual annotation and model training are required, then domain word extraction accuracy can be improved, but the complexity of the construction process increases
Solution Approach 1:
The patent extracts only the essential computational components needed for domain word identification, separating the automated extraction process from complex manual annotation workflows. By focusing on key NLP techniques and removing unnecessary manual steps, the system maintains high extraction accuracy while simplifying the overall construction process.
3Reliability
If accumulated domain words and manual annotation are used, then domain dictionary quality can be improved, but the time required for construction increases
Solution Approach 1:
The patent performs preliminary processing of corpus samples by pre-segmenting text into character units and pre-calculating correlation indices before domain word identification. This preliminary action prepares the data in advance, allowing the main extraction process to run faster while maintaining quality, thereby reducing overall construction time without sacrificing dictionary reliability.
Data Source
AI summary
The present disclosure provides a domain dictionary constructing method and apparatus, the method including: segmenting a domain corpus sample to obtain a first character segment set, where the first character segment set includes at least one first character segment; calculating an inter-character correlation index of the at least one first character segment; determining a second character segment set according to the correlation index, where the second character segment set includes a first character segment of which the correlation index is greater than or equal to a preset threshold; determining a third character segment set according to the second character segment set, where the third character segment set includes the second character segment set and second character segments of the domain corpus; constructing a domain dictionary according to the third character segment set.

