Character Detection Using Word-Level Annotated Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning models for character detection are not fully trained due to the lack of character-level annotation samples, and only samples meeting strict annotation requirements can participate in training, limiting the number of available annotation samples.
Innovation Solution
A method and apparatus that utilize word-level annotated images as input to a machine learning model, selecting characters based on predicted results and annotation information within an annotation region to train the model, allowing for full training and character detection without strict annotation requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If word-level annotated images are used for training, then the number of available annotation samples increases, but the machine learning model cannot be fully trained for character detection
Solution Approach 1:
The patent segments the annotation process by automatically extracting individual character-level annotations from word-level annotated images using OCR technology and manual correction interfaces. This segmentation transforms coarse word-level annotations into fine-grained character-level training data, enabling the model to learn character detection while utilizing the larger pool of word-level annotated images.
Solution Approach 2:
The patent introduces an intermediary processing system that bridges word-level annotations and character-level training requirements. This system includes automated OCR-based character extraction, positioning algorithms, and manual correction interfaces that convert word-level annotation data into character-level training samples, allowing the model to be fully trained on otherwise insufficient data.
2Measurement precision
If strict annotation requirements are imposed, then the quality of training samples improves, but the number of samples that can participate in training decreases
Solution Approach 1:
The patent implements a self-service annotation system where the machine learning model actively participates in its own training data creation. The model performs OCR on word-level annotated images to extract character-level annotations, then uses manual correction interfaces to refine these annotations. This self-service approach maintains high annotation quality while dramatically increasing the number of usable training samples.
Solution Approach 2:
The patent applies preliminary automated processing to generate character-level annotations from word-level images before manual review. OCR technology pre-extracts character positions and content, creating a draft annotation that is then refined through manual correction. This preliminary action ensures both high quality (through manual verification) and quantity (through automated processing of large datasets).
Data Source
AI summary
Disclosed embodiments relate to a character detection method and apparatus. In some embodiments, the method includes: using an image including an annotated word as an input to a machine learning model; selecting, based on a predicted result of characters inside an annotation region of the annotated word predicted and annotation information of the annotated word, characters for training the machine learning model from the characters inside the annotation region of the annotated word predicted; and training the machine learning model based on features of the selected characters. This implementation manner implements the full training of a machine learning model by using existing word level annotated images, to obtain a machine learning model capable of detecting characters in images, thereby reducing the costs for the training of a machine learning model capable of detecting characters in images.


