Resource ID String Generation to Prevent OCR Misreads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in accurately identifying technology resource strings due to electronic misreads during optical character recognition, leading to errors and incorrect actions, as some characters are commonly misread, resulting in inefficiencies and manual review requirements.
Innovation Solution
A system that generates new technology resource strings by comparing them to existing strings, discarding those with a high likelihood of misread characters, and flags existing strings prone to misinterpretation, ensuring a low threshold for electronic misreads by using a combination of matching and commonly misread character pairs analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If optical character recognition is used to extract character strings from electronic documents, then automated processing efficiency is improved, but electronic misreads occur leading to identification errors
Solution Approach 1:
The system performs preliminary actions by generating multiple candidate character strings before final identification, evaluating each candidate against stored strings to predict potential misreads. This advance evaluation prevents incorrect identifications from occurring in the first place, rather than correcting errors after they happen.
Solution Approach 2:
The system incorporates feedback mechanisms by using stored character strings and their known misread patterns to evaluate new candidate strings. The evaluation process feeds back into the selection process, allowing the system to choose candidate strings that are less likely to be misread based on historical misread data from existing strings.
2Speed
If character strings are generated without considering misread patterns, then generation speed is improved, but misidentification errors increase requiring manual review
Solution Approach 1:
The system performs preliminary evaluation of generated character strings by comparing them against stored strings and assessing misread risk before finalizing the selection. This preliminary check ensures that only low-risk strings are selected, maintaining high precision without significantly impacting generation speed.
Solution Approach 2:
The system changes the parameters of character string generation by incorporating misread pattern analysis and similarity thresholds into the generation process. By adjusting these parameters, the system can balance between generation speed and identification accuracy, selecting strings that meet both speed and precision requirements.
3Device complexity
If existing character strings are not evaluated for misread potential, then system complexity is reduced, but errors propagate requiring corrective actions
Solution Approach 1:
The system extracts and analyzes specific misread patterns from stored character strings, separating the evaluation of misread risk from the overall identification process. By focusing on specific problematic patterns rather than evaluating all possible errors, the system maintains lower complexity while improving reliability.
Solution Approach 2:
The system performs self-service by automatically evaluating existing character strings for misread potential and identifying strings that require correction. This automated self-evaluation reduces the need for external manual review while maintaining high identification reliability, as the system independently identifies and flags problematic strings.
Data Source
AI summary
Embodiments of the invention are directed to a system, method, or computer program product structured for generating resource identification strings to avoid electronic misreads. In some embodiments, the system is structured for generating a new technology resource string of characters, comparing the new string to existing technology resource strings, and determining whether the new string is the same as an existing string. The system is also structured for, in response to determining it is not, for each existing string, pairing characters of the strings and determining whether the strings have at least a threshold number of matching character pairs; if there are, for at least one of the existing strings, determining whether characters of the non-matching pairs are commonly misread characters and determining whether there are a threshold combination of matching/commonly misread pairs; and if there are, discard the new string and generate a second new technology resource string.


