Document Quality Evaluation Feature Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document quality evaluation systems fail to aggregate similar feature values extracted from document data, making it difficult for users to understand the nature of the documents.
Innovation Solution
An information processing apparatus that includes an extractor, an obtainer, and an aggregator, where the extractor extracts feature values from input data, the obtainer converts these values into distributed representations, and the aggregator clusters them based on these representations to aggregate similar values.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple feature values are extracted from document data to improve evaluation accuracy, then measurement precision is improved, but device complexity increases due to the large number of similar feature values
Solution Approach 1:
The patent merges similar feature values by calculating similarity between feature value vectors and grouping them into clusters. Feature values with high similarity are combined into representative feature values, reducing the total number of features while preserving the essential information needed for accurate document quality evaluation.
Solution Approach 2:
The patent extracts and removes redundant similar feature values from the feature value set. By identifying and eliminating duplicate or highly similar features, the system reduces complexity while maintaining the most discriminative features for evaluation accuracy.
2Measurement precision
If multiple similar feature values are extracted from document data, then measurement precision is improved, but ease of operation deteriorates as users cannot grasp document nature
Solution Approach 1:
The patent combines multiple similar feature values into clustered groups with representative names that are meaningful to users. This aggregation allows users to understand document characteristics through a manageable number of grouped features rather than being overwhelmed by individual similar features.
Solution Approach 2:
The patent transforms the representation of feature values by changing from individual raw features to clustered groups with aggregated meanings. This parameter transformation makes the feature set more interpretable for users while maintaining the underlying evaluation precision.
3Device complexity
If similar feature values are aggregated into classifications, then device complexity is reduced, but measurement precision may deteriorate due to information loss
Solution Approach 1:
The patent merges similar feature values while preserving their combined information content. By using vector-based similarity calculation and appropriate aggregation methods, the system maintains the essential information from individual features while reducing redundancy, thus preventing measurement precision deterioration.
Solution Approach 2:
The patent applies different aggregation strategies to different clusters of feature values based on their specific characteristics. Each cluster is processed according to its local properties, ensuring that information preservation is optimized for each group while maintaining overall system simplicity.
Data Source
AI summary
A plurality of feature values are extracted from input data as document data, distributed representations of words that correspond to the respective extracted plurality of feature values is obtained, and the extracted plurality of feature values are aggregated into a plurality of classifications based on the obtained distributed representation.


