Semantic Section Identification via Concept Density and Affinity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing question and answer systems face challenges in accurately identifying sections within electronic documents due to reliance on structural delimiters, which can lead to misrepresentation of subject matter and difficulties in cases without explicit section headers.
Innovation Solution
The method involves correlating concepts within textual content to identify concept groups and generate section metadata, using concept density and affinity to partition documents into distinct sections, regardless of structural clues, and associating sections with corresponding headings or inferred headings based on semantic links.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If structural delimiters are used to identify sections, then section identification is simplified, but accuracy deteriorates due to misrepresentation of subject matter and failure to handle documents without explicit headers
Solution Approach 1:
The patent replaces the mechanical approach of using structural delimiters (headers, formatting) with a semantic approach based on concept extraction and analysis. The system identifies concepts within text passages and uses concept density and affinity measurements to determine section boundaries, substituting structural analysis with semantic understanding to achieve both ease of operation and high accuracy.
Solution Approach 2:
The patent introduces new parameters for section identification: concept density (number of concepts per passage) and concept affinity (relatedness between concepts). By changing from structural parameters to semantic parameters, the system can accurately identify sections based on subject matter content rather than formatting, resolving the contradiction between operational simplicity and identification accuracy.
2Measurement precision
If concept correlation and analysis are performed to identify sections, then section identification accuracy is improved, but system complexity increases
Solution Approach 1:
The patent segments the complex task of section identification into distinct manageable steps: concept extraction from passages, concept correlation analysis, concept density calculation, concept affinity measurement, and threshold-based section boundary determination. This segmentation reduces system complexity by breaking down the overall complex process into simpler, modular operations.
Solution Approach 2:
The patent introduces concept vectors as an intermediary representation between raw text and section identification decisions. By converting text passages into concept vectors and using these as intermediaries for comparison and analysis, the system simplifies the complexity of direct text analysis while maintaining high identification accuracy.
Data Source
AI summary
Mechanisms are provided for generating section metadata for an electronic document. These mechanisms receive a document and analyze the document to identify concepts present within textual content of the document. The mechanisms correlate concepts within the textual content with one another to identify concept groups based on the application of one or more rules defining related concepts or concept patterns. The mechanisms determine sections of text within the textual content based on the correlation of concepts within the textual content. Based on results of the determining, the mechanisms generate section metadata for the document and store the section metadata in association with the document for use by a document processing system.


