Annotation Assessment and Adjudication for Ground Truth Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cognitive systems, such as AI platforms, are inherently non-deterministic, leading to inconsistencies and inaccuracies in natural language processing due to the susceptibility of data output to input variations, which results in incorrect data extraction and annotation errors.
Innovation Solution
A system comprising a natural language manager, annotation manager, assessment manager, and ground truth manager that identifies grammatical elements, assigns and assesses annotations, dynamically adjusts scores, and selectively re-assigns annotations to construct accurate ground truth data, ensuring deterministic behavior and improving annotation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If cognitive systems process natural language based on knowledge from databases or corpora, then the system can provide relevant recommendations and solve problems, but the resulting outcome can be incorrect or inaccurate due to peculiarities of language constructs and human reasoning
Solution Approach 1:
The patent implements a feedback mechanism where multiple annotators annotate the same data, and their annotations are compared and assessed. The system provides feedback by identifying discrepancies between annotations and using assessment scores to improve future annotation accuracy, thereby resolving the contradiction between adaptability and reliability
Solution Approach 2:
The system performs self-assessment of annotation quality by automatically comparing multiple annotations against each other and generating assessment scores without requiring external validation. This self-service approach improves reliability while maintaining the system's ability to process diverse natural language inputs
2Adaptability or versatility
If new machine learning models are deployed to improve system capabilities, then the system can extract entities and provide insights, but there is no guarantee that the system will extract the same entities as done previously, leading to non-deterministic behavior
Solution Approach 1:
The patent applies preliminary action by establishing a baseline set of annotations before deploying new models. These baseline annotations serve as a reference point to measure changes and ensure consistency, allowing the system to adapt to new models while maintaining stable entity extraction behavior
Solution Approach 2:
The system dynamically adjusts to new machine learning models by re-assessing annotations when models change. The assessment manager dynamically updates annotation scores based on comparisons with baseline annotations, enabling the system to maintain consistency despite model updates and version changes
3Reliability
If manual annotation processes are used to create ground truth data, then annotations can be assigned to grammatical elements, but the process is time-consuming and prone to human error
Solution Approach 1:
The patent merges multiple annotators' work on the same data elements, combining their annotations into a single assessed result. This merging approach improves reliability through consensus while reducing the total time required compared to sequential manual annotation by multiple individuals
Solution Approach 2:
The assessment manager acts as an intermediary that automatically evaluates and compares annotations from multiple annotators. This intermediary process eliminates the need for time-consuming manual review and resolution, automatically resolving conflicts and producing final ground truth annotations efficiently
Data Source
AI summary
Embodiments relate to an intelligent computer platform to construct ground truth data. A document is subjected to NLP and grammatical elements of the document are identified. The identified grammatical elements are annotated and one or more annotations are assigned to the elements, with each annotation having corresponding meta-data. The assigned annotations are assessed for accuracy and a score associated with the meta-data is adjusted. The annotations are selectively re-assigned with respect to the identified grammatical elements responsive to the adjusted score. Ground truth data is constructed and the selectively re-assigned annotations are assigned to corresponding and identified grammatical elements.


