A computer-implemented method for identifying a
problem list section from an
electronic document includes receiving, by one or more processors, the
electronic document, generating, by the one or more processors and based on applying an
optical character recognition algorithm to the
electronic document, unstructured text, and identifying, by the one or more processors, one or more
problem list words in the unstructured text, the one or more
problem list words belonging in a dataset for identifying a presence of a problem
list section. The method also includes associating, by the one or more processors, a portion of the unstructured text that corresponds to the one or more problem
list words in the unstructured text with the problem
list section and outputting, by the one or more processors, at least a portion of the problem list section.