AI-POWERED DOCUMENT PROCESSING SYSTEM AND METHOD FOR VISUALLY IMPAIRED USERS
Patent Information
- Application Number
- TR202612315
- Authority / Receiving Office
- TR · TR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-21
Abstract
Description
AI-POWERED DOCUMENTATION FOR USERS WITH VISUAL IMPAIRMENT. PROCESSING SYSTEM AND METHOD Technical Field to Which the Invention Relates This invention will enable computer-based document processing, optical character recognition, and the reproduction of historical alphabets. transcription, natural language processing, accessible user interfaces, screen reader compatibility, even tolerant distributed data processing and text-to-speech conversion technologies It is related to the invention, particularly printed or handwritten documents written in Ottoman Turkish. Digital text that is independently accessible to users with visual impairments. It is a system and method for converting audio into sound output. State of the Art In known technological systems, printed or image documents are scanned and converted to optical format. It can be subjected to character recognition processing, and the resulting texts are converted from text to speech. It can be voiced using conversion engines. However, with historical alphabets, letter forms, combined characters, and spelling, especially in documents written in Ottoman Turkish. differences include right-to-left reading direction, page layout, old number and symbol representations, and handwriting. Due to the diversity, standard optical character recognition systems offer insufficient accuracy and It is unable to provide usability. Additionally, existing solutions include remote processing services for multiple and voluminous documents. excessive requests, connection interruptions, or service unavailability may occur during the sending process. When this happens, the process may have to start over, the page order may be disrupted, or data may be lost. This can occur. However, transferring the OCR output directly to the interface can result in formatting residues. The screen displays inappropriate symbols, confusing labels, and an incorrect reading sequence. This makes it difficult for the guards to follow the text continuously. The sound is incompatible with the target language. The choice of engine also includes Ottoman Turkish or Arabic script content versus today's Latin script content. This can prevent the correct pronunciation of the Turkish word. Our invention precisely addresses this issue with the technology. It aims to remedy the deficiency in the specified condition and for those with visual impairment. It helps users access printed or handwritten works. 1 Purpose of the Invention The purpose of this invention is to produce accessible audio output from the capture of a historical document image. by combining the entire process chain, from production to manufacturing, into a single system for visually impaired individuals allowing users to upload documents without requiring external settings or auxiliary software, Determine the direction of transformation, access the processed text, and vocalize the text. The aim is to provide the following information regarding the transaction status in the invention: file ID, page number, conversion direction, and Since the completed output information is stored along with the data, it occurs in the primary processing engine. After the error occurs, a gradual delay can be applied and a retry can be made, or It is possible to switch to the secondary engine and continue the process from the last completed section. This reduces the impact of network and service disruptions on multi-page document processing. The accessibility normalization module normalizes the OCR / transcription result to the screen reader. converting it into a linear reading sequence that it can follow; a specified number, character, and applying punctuation rules; a style specific to visual presentation and disrupting auditory reading. It removes the remnants. The speech synthesis interface then produces sounds suitable for the target text and language. It enables quick switching by automatically identifying the problem and using user commands. Description of the Invention The invention's application involves a document processing system, a server over a network, or cloud processing. accessible user interface running on user terminal communicating with its infrastructure It includes the interface, full keyboard functionality, focus order preservation, and screen reader recognition. It is configured with recognizable field names and single-command output copying functions. The user can access one or more files in PDF, word processing document, or image file formats. It loads the source document into the interface. The input and format conversion module processes the file. It determines the type of document; divides the document into pages or logical segments; and arranges the segments sequentially. placing each segment in a job queue and processing it as machine-processable coded data. It converts image data into a package. In an application, image data is encoded using Base64. The data is transferred to the application programming interface; however, different binary data transfer methods are used. Other protocols can also be used. The transformation direction and language analysis module allows the user to analyze Ottoman language. Transcription from a Turkish / Arabic script source into modern Turkish using Latin script. Latinization or the transformation from modern Turkish to Ottoman Turkish orthography. It allows the user to select one of the options. The selected orientation information is associated with the document segment. The same process log is kept and the technical data is sent to the OCR / transcription processing engine. 2 It defines the instruction set. The OCR / transcription processing engine processes the source image. Identifying character spaces, extracting line and reading order, character recognition and language It implements model-based transcription processes and performs intratextual morphological analysis. It does this. In an application, the primary processing model aims to reduce variability. It is operated at a production temperature of 0.2 or lower. Accessibility normalization. The module processes the resulting text according to a predefined set of accessibility rules. It is in operation. The rule set makes it difficult for the screen reader to pronounce the unique and accurate numbers. converting the representations to their equivalent Arabic numerals, removing unnecessary form marks removal, merging line breaks according to paragraph logic, right-to-left and left-to-right Arrange text fragments on the right in a linear reading order, heading and paragraph. preserving the distinctions and including interpretation or model explanation in the output This includes blocking Roman numerals 1–20 in one application. It is converted to Arabic numerals. Error management and process status module, File ID, segment number, completed segments, pending segments for each document. The status log includes segments, the model used, the type of error, and the number of attempts. This creates an error code or error code indicating an excessive request situation from the primary processing model. If a network / connection error occurs, the module will enter a staggered waiting period depending on the number of attempts. It applies the time limit and resends the request. The waiting time for an application is 4000. These are multiples of milliseconds. When the defined trial threshold is exceeded, the process switches to a secondary processing model. being transferred; thanks to the status record, previously completed segments can be re-recorded. The process continues from the next segment without processing. Accessible output module. Normalized text is plain text that the screen reader can focus on directly, or It presents the text within a pre-formatted code block. The output module outputs the text to the clipboard with a single command. functions include copying, exporting as a file, and saving to the process / archive data store. It provides at least one of the following: data repository, document ID and source file, conversion direction. By establishing a relationship between the segment sequence and the final text, it enables access to the digital library. It provides. The voice synthesis interface uses the available voice on the user terminal or server. They are querying the engines; according to language tags and writing system information, Turkish and Ottoman. It lists the sounds that match the Turkish / Arabic content. The sound engine selector displays the selected transformation. It determines the default sound according to the direction and provides the user with keyboard shortcuts for sound engines. It switches between them. In an application, the CTRL+1 and CTRL+2 shortcuts are two different shortcuts. It is used for switching between language / speech engines. Speech rendering is an external output. It is executed within the interface without needing to be transferred to the program. 3 Our invention, as described, has the potential to be applied to all historical and contemporary world languages in the future. and we are carrying out development work in this direction. How the invention can be applied to industry. The invention is relevant to university and research libraries, archives, museums, and the digital humanities. centers, public institutions, distance learning platforms, accessible publishing organizations and by all institutions serving individuals with visual impairments, server, cloud, can be produced and repeated on desktop or mobile computing platforms. applicable. 4
Claims
1. Resource documents written in Ottoman Turkish for users with visual impairments. It is a document processing system that converts document information into accessible digital text and audio output; Enables the user to upload at least one source document and select the conversion direction. Keyboard-operated accessible user interface; divides the source document into pages or segments. input and separates and converts each segment into a machine-processable coded data packet. Format conversion module; converts from Ottoman Turkish to modern Turkish according to the selected conversion direction. Instructions for processing the Turkish or the conversion from modern Turkish to Ottoman Turkish spelling. The transformation direction and language analysis module that is generated; optically analyzes the data packets in question. at least one primary processing unit that performs character recognition, transcription, and morphological analysis The OCR / transcription processing engine, which includes the model, processes characters, numbers, and other elements in the processed text. punctuation, formatting, and reading order elements are rules for screen reader compatibility. The accessibility normalization module transforms according to the set; processing each segment. a primary processing model that stores its state and has a progressive delay when an error occurs. retrying and completing the process in the secondary processing model under the specified condition Error management and process status module that transfers information from the last segment to the next segment; Normalized text can be directly focused on by the screen reader (plain text). an accessible output module that provides the appropriate voice engine for the target language; and by selecting the appropriate voice engine for the target language. It is characterized by containing a speech synthesis interface that vocalizes normalized text. It is done.
2. Document processing system according to claim 1, and its feature is the input and format conversion module. It must accept at least two of the following file formats: PDF, word processing document, and image file format. It is characterized by its ability to queue multiple files in a single process queue.
3. Document processing system according to claim 1 or 2, with input and format conversion module. network by converting the document image to Base64 or equivalent binary data encoding format It is characterized by its ability to transfer applications to the programming interface (API) on it.
4. The document processing system is the primary processing system according to any of the previous requirements. To reduce the output variability of the model, a production temperature of 0.2 or lower is used. It is characterized by being run with the specified parameter.
5. Document processing system and accessibility according to any of the previous requirements. The normalization modulus uses Roman numerals from I to XX, respectively, between 1 and 20. It is characterized by its conversion to Arabic numerals.
6. Document processing system and accessibility according to any of the previous requirements. normalization module sorts text from right to left and left to right for screen reader placing it in linear reading order, removing unnecessary formatting tags and headings It is characterized by maintaining paragraph boundaries.
7. Document processing system according to any of the previous requirements, including error management and processing. The status module attempts to detect an excessive request error code or connection error. Characterized by the application of waiting times in multiples of 4000 milliseconds, depending on the situation. It is done.
8. The document processing system is based on any of the previous requirements, and the transaction status record... file ID, segment number, transformation direction, processing model used, completion It is characterized by having at least three of the required knowledge and trial and error information.
9. A document processing system and voice synthesis system according to any of the previous requirements. the interface lists the available sound engines according to language tags and the sound engine selector Turkish and Ottoman Turkish / Arabic voice engines via at least two keyboard shortcuts. It is characterized by its ability to facilitate transitions between the two.
10. Resource documents written in Ottoman Turkish for visually impaired users. a computer-generated document that is converted into accessible digital text and audio output It is a processing method that involves accessing at least one source document through an accessible user interface. The process of obtaining the source document; dividing the source document into pages or segments and machine-sorting the segments. conversion of encoded data packets into processable data packets; from Ottoman Turkish to modern Turkish the direction of transformation from Turkish or modern Turkish to Ottoman Turkish orthography acquisition; optical character recognition, transcription and on encoded data packets. Morphological analysis is applied to assess the screen reader compatibility of the resulting text. according to the rule set in terms of character, number, punctuation, format, and reading order. normalization; storing the processing status for each segment; error in the processing engine. When this occurs, a phased delayed retry is performed and a secondary attempt is made under the specified condition. 6 The process involves switching to the next segment after the last completed segment in the processing engine; normalized text in a plain text area that can be focused by the screen reader presentation; and normalized text by selecting a speech engine suitable for the target language. It is characterized by the fact that it includes the steps of voice acting.
11. This method is based on Claim 10, and if the source document contains multiple files, the file... and creating a single job queue according to page order and assigning a unique segment to each job item. It is characterized by including the steps involved in identity assignment.
12. The method is according to claim 10 or 11, where the transformation direction and segment ID are the same operation. the instruction set to be kept in the record and sent to the OCR / transcription engine It is characterized by its automatic generation based on the direction of transformation.
13. The method according to either claim 10 or 12, after error detection. the last successful segment in a way that will prevent reprocessing of completed segments by storing the information and starting the secondary processing engine from the following segment It is the characterization of the situation.
14. The method is according to either claim 10 or 13, and during normalization, I to XX converting Roman numerals between 1 and 20 to Arabic numerals between 1 and 20. Characterized by the removal of unnecessary tags specific to the visual presentation within the text. It is done.
15. The method is according to either claim 10 or 14, and the normalized text is prioritized. presented in a formatted text block and clipped with a single user command. It is characterized by being copied or exported as a file.
16. The method is according to either 10 or 15, and is available on the user terminal. Querying voice engines based on language tags, defaulting to the appropriate voice for the conversion direction. by selecting the engine and switching between at least two sound engines using a keyboard shortcut It is the characterization of the situation.
17. Method according to either claim 10 or 16, source document, conversion direction, an archive record that establishes a relationship between segment sequence and final accessible text 7 by creating the record and making it accessible through the digital library It is the characterization of the situation. 8