METHOD FOR FILTERING ERRORS IN THE USE OF NUMBERS AND NUMBER IN INDONESIAN LANGUAGE TEXTS BASED ON SEQUENTIAL DETECTION AND GENERATIVE CORRECTION

IDS00202606864APending Publication Date: 2026-07-15UNIVERSITAS MULTIMEDIA NUSANTARA

Patent Information

Authority / Receiving Office
ID · ID
Patent Type
Utility models
Current Assignee / Owner
UNIVERSITAS MULTIMEDIA NUSANTARA
Filing Date
2026-06-25
Publication Date
2026-07-15
Patent Text Reader

Abstract

This invention concerns the field of natural language processing (NLP) and the application of hybrid imitation intelligence to language, specifically a method for filtering errors in the use of numbers and numerals in Indonesian. This method consists of two main integrated stages, namely the error detection stage and the automatic correction stage for the use of numbers and numerals. The detection stage is carried out through collecting an Indonesian language text corpus, filtering sentences based on Regular Expressions (Regex), data pre-processing, sequential labeling based on the Inside, Outside, Beginning (IOB) scheme, linguistic feature engineering, and Conditional Random Fields (CRF) based modeling to recognize errors in the use of numbers and numerals contextually.The detection results serve as input for the automatic correction stage using the Transformer model, specifically the Indonesian language variant of the Text-to-Text Transfer Transformer (T5) (IndoT5), to generate correction recommendations according to the sentence context and Indonesian language rules. The output of this method is a filtering result containing errors in the use of numbers and figures along with automatic correction recommendations displayed through a web-based text checking system.
Need to check novelty before this filing date? Find Prior Art