Text Coreference Analysis via Subject Grouping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current natural language processing technologies face challenges in accurately analyzing and processing text data to identify coreference relationships between different names and entities, making it difficult to extract relevant information and insert targeted advertisements without affecting the integrity of the text.
Innovation Solution
A method and system for text processing that involves acquiring text, extracting subject and entity words, grouping subject words to determine entity words referencing a common object, and generating processing policies to extract and analyze comments or insert advertisements effectively, using techniques such as subject relationship networks, named entity recognition, and emotion analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional natural language processing technology is used to analyze character coreference, then character-level coreference can be determined, but object-level coreference analysis accuracy deteriorates
Solution Approach 1:
The patent segments the coreference analysis process into character-level and object-level stages. First, character coreference is resolved to identify referring expressions, then these expressions are grouped to form objects, and finally object coreference is established. This segmentation allows the system to achieve both character-level precision and object-level versatility.
Solution Approach 2:
The patent introduces a new dimension of analysis by moving from character-level coreference to object-level coreference. It creates an additional layer of abstraction where referring expressions are grouped into objects before establishing coreference relationships, thereby expanding the capability from simple character matching to comprehensive object identification.
2Loss of information
If comprehensive text analysis is performed to identify all coreference relationships, then information extraction completeness improves, but processing time increases
Solution Approach 1:
The patent performs preliminary actions by first resolving character coreference relationships and identifying all referring expressions before proceeding to object formation and object-level coreference. This preliminary characterization of the text enables more efficient subsequent processing and reduces the need for repeated analysis passes.
Solution Approach 2:
The system uses the results of character coreference resolution to automatically inform object formation, and then uses object structure to guide object coreference analysis. Each stage serves itself by utilizing the output of previous stages, reducing redundant computation and improving overall efficiency.
3Productivity
If targeted advertisement insertion is performed based on entity coreference, then advertising effectiveness improves, but text integrity may deteriorate
Solution Approach 1:
The patent introduces processing policies as an intermediary layer between the coreference analysis results and the actual advertisement insertion. These policies act as mediators that evaluate whether and how to insert advertisements based on the identified object coreference relationships, ensuring that advertising effectiveness is achieved while text integrity is preserved.
4Measurement precision
If subject word grouping is performed to form subject word groups, then object coreference determination accuracy improves, but system complexity increases
Solution Approach 1:
The patent segments the complex task of object coreference determination into manageable steps: extracting subject words, grouping them into subject word groups based on grammatical and semantic relationships, and then determining coreference within these groups. This segmentation reduces system complexity by breaking down the problem into smaller, more tractable sub-tasks.
Solution Approach 2:
The subject word grouping mechanism is dynamic and adaptive, adjusting the grouping criteria based on the specific text context and the types of entities being analyzed. This dynamic approach allows the system to maintain high accuracy without requiring a fixed, overly complex structure.
Data Source
AI summary
System, method and program product for text processing using object coreference technology. In particular, the invention provides a text processing method which includes, acquiring text to be processed; extracting subject words and entity words corresponding to the subject words from the text; grouping the subject words; determining entity words that reference a same concerned object according to the grouped subject words; and generating processing policy for entity words that reference a same concerned object. The invention also includes a system with means for carrying out the method. The invention generally realizes automatic, more comprehensive, accurate, efficient analysis and processing on text data. The invention can be used to dig a large amount of comment data about some entity, and the invention can also be used to suggest insertion place in an article where embedded advertisement is inserted.


