Linguistic Annotation Projection via Parallel Corpus Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual annotation of linguistic resources is time and cost intensive, and direct projection of linguistic labels from resource-rich languages to resource-poor languages in parallel corpora often results in errors due to mismatches and translation shifts, limiting the availability of linguistic resources for many languages.
Innovation Solution
A two-stage approach using a parallel corpus, where linguistic annotations are projected from a source language to a target language, filtered to remove errors, and then used to train a machine learning model that iteratively adds annotations to incomplete target language texts, improving both precision and recall of linguistic labels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If linguistic annotations are manually annotated, then annotation accuracy is improved, but time consumption and cost increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-annotating a source language corpus with linguistic annotations before projection. This pre-prepared annotated corpus serves as the foundation for automated annotation of target languages, eliminating the need for manual annotation in each target language while maintaining reasonable accuracy through subsequent filtering and machine learning refinement.
Solution Approach 2:
The patent uses copying by projecting linguistic annotations from the source language corpus to parallel target language corpora. Instead of creating annotations from scratch for each language, the system copies annotations across languages using parallel corpus alignment, significantly reducing time and cost while preserving annotation quality through filtering and iterative improvement.
2Productivity
If linguistic annotations are projected from source language to target language, then productivity is improved, but annotation precision deteriorates due to translation shifts and mismatches
Solution Approach 1:
The patent implements feedback through iterative refinement where projected annotations are filtered, evaluated, and used to retrain machine learning models. The system continuously improves annotation precision by feeding back corrected annotations into the training process, allowing the model to learn from errors and progressively reduce translation shift mismatches while maintaining high productivity.
Solution Approach 2:
The patent replaces the mechanical system of direct annotation projection with a machine learning-based system. Instead of简单地 copying annotations, the system uses trained models to predict and refine annotations, substituting rigid mechanical projection with adaptive intelligent processing that can handle translation shifts and linguistic variations while maintaining both precision and productivity.
3Measurement precision
If filters are applied to remove erroneous projected annotations, then annotation precision is improved, but the quantity of annotated data decreases
Solution Approach 1:
The patent applies discarding and recovering by filtering out erroneous projected annotations and then recovering lost annotation coverage through machine learning prediction. The system discards low-confidence or erroneous annotations identified by filters, then recovers the annotation coverage by using the ML model to predict annotations for filtered-out cases, maintaining both precision and adequate data quantity for training.
4Quantity of substance
If a machine learning model is trained iteratively to add annotations, then completeness of linguistic resources is improved, but device complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the annotation completion process into distinct stages: initial projection, filtering, ML model training, and iterative refinement. Each stage handles a specific aspect of the annotation process, making the overall complex system more manageable and enabling systematic improvement of linguistic resource completeness through structured incremental development.
Data Source
AI summary
One embodiment provides a method for generating a natural language resource using a parallel corpus, the method including: utilizing at least one processor to execute computer code that performs the steps of: receiving, from a parallel corpus, natural language text in a source language and a corresponding translation of the natural language text in a target language, wherein the natural language text in the source language comprises linguistic annotations; projecting the linguistic annotations from the source language natural language text to the target language natural language text; applying one or more filters to remove at least one projected linguistic annotation from the target language natural language text that results in at least one error; selecting at least one target language natural language text having substantially complete linguistic annotations; training a machine learning model using the selected at least one target language natural language text and annotations; and adding, using the trained machine learning model, linguistic annotations to at least one target language natural language text having incomplete linguistic annotations. Other aspects are described and claimed.


