A
machine-
learning based (ML-based)
system and method for automatically extracting one or more data fields from one or more documents, are disclosed. The ML-based
system includes a document obtaining subsystem to obtain documents, a document pre-
processing subsystem to generate pre-processed data, a field identifying subsystem to identify data fields using a trained ML model, and a field extracting subsystem to extract financial information. The ML-based
system also comprises an output subsystem to deliver the extracted data to end users via user interfaces. The ML model is trained using historical documents, labelled data fields, and features such as distance-based features, direction-based features, dimension-based features, positional features, and value-based features. The M-based system employs
hyperparameter optimization,
noise removal, and accuracy assessment mechanisms to enhance performance. This ML-based system provides a scalable, accurate, and automated solution for financial
information extraction, ensuring efficiency, adaptability, and seamless integration with enterprise systems.