METHOD FOR FILTERING ERRORS IN THE USE OF NUMBERS AND NUMBER IN INDONESIAN LANGUAGE TEXTS BASED ON SEQUENTIAL DETECTION AND GENERATIVE CORRECTION
IDS00202606864APending Publication Date: 2026-07-15UNIVERSITAS MULTIMEDIA NUSANTARA
Patent Information
- Authority / Receiving Office
- ID · ID
- Patent Type
- Utility models
- Current Assignee / Owner
- UNIVERSITAS MULTIMEDIA NUSANTARA
- Filing Date
- 2026-06-25
- Publication Date
- 2026-07-15
Abstract
This invention concerns the field of natural language processing (NLP) and the application of hybrid imitation intelligence to language, specifically a method for filtering errors in the use of numbers and numerals in Indonesian. This method consists of two main integrated stages, namely the error detection stage and the automatic correction stage for the use of numbers and numerals. The detection stage is carried out through collecting an Indonesian language text corpus, filtering sentences based on Regular Expressions (Regex), data pre-processing, sequential labeling based on the Inside, Outside, Beginning (IOB) scheme, linguistic feature engineering, and Conditional Random Fields (CRF) based modeling to recognize errors in the use of numbers and numerals contextually.The detection results serve as input for the automatic correction stage using the Transformer model, specifically the Indonesian language variant of the Text-to-Text Transfer Transformer (T5) (IndoT5), to generate correction recommendations according to the sentence context and Indonesian language rules. The output of this method is a filtering result containing errors in the use of numbers and figures along with automatic correction recommendations displayed through a web-based text checking system.
Need to check novelty before this filing date? Find Prior Art