Systems and methods for extracting information from documents using language models are disclosed. In an embodiment, a method includes receiving an input document containing unstructured or semi-structured data, performing
document classification by providing a first prompt to a
first language model, the first prompt including a section defining
document classification parameters, example documents, and the input document, performing format
pattern detection by providing a second prompt to a second
language model, the second prompt including a section defining format analysis parameters, a
list of pre-configured document patterns, and the input document, performing
information extraction by providing a third prompt to a third
language model, the third prompt including a section defining
information extraction parameters, pattern extraction
metadata associated with the matched document pattern, a schema specification, and the input document, wherein the third
language model outputs a normalized data
record according to the schema with field values and confidence scores.