Robotic process automation for document understanding that integrates generative artificial intelligence models
The document understanding engine with a generative AI model addresses the inefficiencies of traditional document processing by rapidly and accurately classifying and validating large volumes of scanned documents, reducing time and resources.
Patent Information
- Application Number
- JP2024070393
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-04-24
- Publication Date
- 2025-07-24
AI Technical Summary
Existing document processing methods are time-consuming and resource-intensive due to the lack of automatic solutions for identifying types, items, fields, and information in scanned documents, leading to inefficient scanning, review, and editing of large volumes of paper documents.
Implementing a document understanding engine that integrates a generative AI model for pre-annotating, classifying, and verifying files, utilizing models like generative pre-trained transformers (GPTs) to accurately and efficiently classify and understand document content.
Reduces processing time from weeks to minutes for large document sets, minimizing manual labor and resource consumption while enhancing accuracy and efficiency in document classification and validation.
Smart Images

Figure 2025109160000009 
Figure 2025109160000010 
Figure 2025109160000011
Abstract
Description
Technical Field
[0001] The present invention generally relates to robotic process automation (RPA) for document understanding, and more specifically to RPA for document understanding that integrates a generative artificial intelligence (AI) model.
Background Art
[0002] Companies have thousands of paper documents. To store these paper documents electronically, each paper document needs to be scanned. When a paper document is scanned, a repository of thousands of scanned files that are not marked or organized is created. Each paper document is scanned by a scanner or printer, resulting in a flat image without markings (e.g., a scanned file). Additionally, it should be noted that the scanned files from the scanner or printer are assigned random alphanumeric names that do not indicate the type or items, fields, and information within the flat image. As a result, companies need to spend significant resources and time on the conventional scanning, review, identification, and editing of all scanned files.
[0003] Conventional scanning, review, identification, and editing include scanning paper documents to create scanned files in a database, opening and viewing the scanned files individually, identifying the content of the scanned files during viewing, editing the content of the scanned files to reflect categories and other qualification information, and editing individual annotations of items, fields, and information within the scanned files during viewing of the scanned files.
[0004] Each step of traditional scanning, review, identification, and editing (traditional document processing) is time-consuming, boring, and very time-consuming (for example, when performed manually, it may take 4 minutes to process individual paper documents in the best scenario). Furthermore, when scanning a large number of thousands of paper files, the cost of traditional document processing further increases. Therefore, there is a problem that companies can only perform traditional document processing on scanned files because there is no automatic document understanding solution that can identify the types, items, fields, and information of scanned documents.
[0005] Therefore, an automatic document understanding solution is needed.
Summary of the Invention
[0006] According to one or more embodiments, a method is provided. The method is performed by a document understanding engine implemented as a computer program. The method includes pre-annotating one or more files with one or more annotation proposals by a special model of the document understanding engine. The method includes presenting, by the document understanding engine, one or more files including one or more annotation proposals for verification to a user interface. The user interface includes scoring of the one or more files. The method includes training, by the document understanding engine, the special model based on one or more inputs to improve scoring and pre-annotation, the inputs being received in response to the annotation proposals.
[0007] According to any of one or more embodiments herein or method embodiments, the document understanding engine may be implemented as an apparatus, a system, and a computer program product.
Brief Description of the Drawings
[0008] To facilitate an understanding of the advantages of particular embodiments of the present invention, a more particular description of the invention briefly described above is presented with reference to the specific embodiments illustrated in the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are not to be considered limiting of its scope, but that the invention will be described and explained with additional particularity and detail by use of the following accompanying drawings.
[0009]
Figure 1
[0010]
Figure 2
[0011]
Figure 3
[0012]
Figure 4
[0013]
Figure 5
[0014]
Figure 6
[0015]
Figure 7
[0016]
Figure 8
[0017]
Figure 9
[0018]
Figure 10
[0019]
Figure 11
[0020]
Figure 12
[0021]
Figure 13
[0022]
Figure 14
[0023]
Figure 15
[0024]
Figure 16
[0025] Unless otherwise specified, like reference characters consistently refer to corresponding features throughout the accompanying drawings.
Best Mode for Carrying Out the Invention
[0026] (Detailed Description of Embodiments) According to one or more embodiments, robotic process automation (RPA) for document understanding is provided. Document understanding RPA integrates a generative artificial intelligence (AI) model for document classification operations such as, for example, digitizing, classifying, pre-annotating, and verifying files. For example, document understanding RPA uses special models and / or language models to pre-annotate the content of files to improve file understanding, processing, and storage.
[0027] According to one or more embodiments, document understanding RPA is integrated with a generative AI model within a document understanding engine. The document understanding engine includes software implemented as processor-executable code stored in memory. The document understanding engine may be implemented by a combination of processor-executable code and hardware described herein. For example, the processor-executable code of the document understanding engine is executed by at least one processor to perform document classification operations including one or more of digitizing, classifying, pre-annotating, and verifying files. Thus, the document understanding engine needs to be based on the operation of at least one processor to improve or replace the conventional document processing of scanned files.
[0028] Accordingly, as a technical effect, merit, and advantage, the document understanding engine provides a mechanism for adapting an AI model, or a mechanism for creating an AI model that classifies files more accurately and efficiently, reducing the time and resources required for conventional document processing (e.g., reducing the extraction time and labor of documents, thereby enabling processed files to be distributed more accurately and quickly) to a specific domain (e.g., type / classification). As an example, the document understanding engine uses one or more generative pre-trained transformers (GPTs) to discover the type of document (e.g., domain or classification) and to discover fields within the document to understand the document. By comparison, in conventional document processing, it can take up to 66 business hours, or 1.5 weeks of labor time, to process 1,000 paper documents, whereas the document understanding engine can process the same 1,000 paper documents in 1 to 10 minutes after receiving the scanned files. Further, as a technical effect, benefit, and advantage, the document understanding engine provides a mechanism for validating files at runtime, reducing the time and resources required for the document understanding engine itself (e.g., minimizing or eliminating reviews, thereby enabling processed files to be distributed more accurately and quickly).
[0029] FIG. 1 is an architectural diagram showing a hyper-automation system 100 according to one or more embodiments. As used herein, "hyper-automation" refers to an automation system that combines process automation components, integration tools, and technologies that amplify the ability to automate work. For example, in some embodiments, RPA is used at the core of the hyper-automation system, and in certain embodiments, the automation capabilities can be extended by AI / ML, process mining, analytics, and / or other advanced tools. When the hyper-automation system learns processes, trains AI / ML models, and employs analytics, for example, more knowledge work can be automated, and both computing systems within an organization, such as those used by individuals and those that operate autonomously, can all participate as actors in the hyper-automation process. The hyper-automation systems of some embodiments enable users and organizations to discover, understand, and extend automation efficiently and effectively.
[0030] The hyper-automation system 100 includes user computing systems such as desktop computer 102, tablet 104, and smartphone 106. However, any desired computing system, including but not limited to smartwatches, laptop computers, servers, Internet of Things (IoT) devices, etc., may be used without departing from the scope of one or more embodiments described herein. Also, although FIG. 1 shows three user computing systems, any suitable number of computing systems may be used without departing from the scope of one or more embodiments described herein. For example, in some embodiments, dozens, hundreds, thousands, or millions of computing systems may be used. The user computing systems may be actively used by a user or may be automatically executed with little or no user input.
[0031] Each computing system 102, 104, 106 has respective automation process(es) 110, 112, 114 executing thereon. The automation process(es) 102, 104, 106 can include, without limitation and without departing from the scope of one or more embodiments described herein, RPA robots, a part of an operating system, downloadable application(s) for each computing system, any other suitable software and / or hardware, or any combination thereof. In some embodiments, one or more process(es) 110, 112, 114 can be listeners. The listener(s) can be, without departing from the scope of one or more embodiments described herein, RPA robots, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware. Indeed, in some embodiments, the logic of the listener(s) is implemented partially or fully via physical hardware.
[0032] The listener monitors and records data related to user interactions with each computing system and / or the operation of unattended computing systems, and transmits the data to the core hyper-automation system 120 via a network (e.g., a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, any combination thereof, etc.). The data can include, but is not limited to, which buttons were clicked, where the mouse moved, the text entered into fields, that one window was minimized and another window was opened, the applications associated with the windows, and the like. In certain embodiments, data from the listener can be transmitted periodically as part of a heartbeat message. In some embodiments, the data can be transmitted to the core hyper-automation system 120 when a predetermined amount of data has been collected, after a predetermined period of time has elapsed, or both. One or more servers, such as server 130, receive the data from the listener and store it in a database, such as database 140.
[0033] The automation process can perform logic developed in the workflow during design time. In the case of RPA, the workflow can include a set of steps performed in a sequence or some other logical flow, defined herein as an "activity". Each activity can include actions such as clicking a button, reading a file, writing to a log panel, and the like. In some embodiments, the workflow can be nested or embedded.
[0034] The long-running workflows for RPA in some embodiments are master projects that support service orchestration, human intervention, and long-running transactions in an unattended environment. See U.S. Patent No. 10,860,905. This content is hereby incorporated by reference in its entirety. Human intervention occurs when a particular process requires human input for exception handling, approval, or verification before proceeding to the next step of the activity. In this case, the execution of the process is paused and the RPA robot is released until the human task is completed.
[0035] The long-running workflow may support fragmentation of the workflow via persistence activities, combined with call processes and non-user interaction activities, and orchestrate human tasks with RPA robot tasks. In some embodiments, multiple or a large number of computing systems may participate in the execution of the logic of the long-running workflow. The long-running workflow may be executed in a session to facilitate rapid execution. In some embodiments, the long-running workflow may orchestrate a background process that executes application programming interface (API) calls and may include activities that execute in the long-running workflow session. These activities may, in some embodiments, be called by a call process activity. A process having a user interaction activity that executes in a user session may be called by starting a job from a conductor activity (the conductor is described in more detail later in this specification). The user may, in some embodiments, interact through a task that requires the user to complete a form in the conductor. Activities may be included that cause the RPA robot to wait for the form task to be completed and then resume the long-running workflow.
[0036] One or more automation processes 110, 112, 114 communicate with a core hyper-automation system 120. In some embodiments, the core hyper-automation system 120 may execute a conductor application on one or more servers such as server 130. Although one server 130 is shown for illustration purposes, multiple or numerous servers in close proximity to each other or in a distributed architecture may be employed without departing from the scope of the one or more embodiments described herein. For example, one or more servers may be provided for conductor functionality, AI / ML model provisioning, authentication, governance, and / or any other suitable functionality without departing from the scope of the one or more embodiments described herein. In some embodiments, the core hyper-automation system 120 may incorporate or be part of a public cloud architecture, a private cloud architecture, a hybrid cloud architecture, etc. In certain embodiments, the core hyper-automation system 120 may host multiple software-based servers on one or more computing systems such as server 130. In some embodiments, one or more servers of the core hyper-automation system 120, such as server 130, may be implemented via one or more virtual machines (VMs).
[0037] In some embodiments, one or more automation processes 110, 112, 114 may call one or more AI / ML models 132 deployed on or accessible by a core hyper-automation system 120. The AI / ML models 132 can be trained for any suitable purpose without departing from the scope of one or more embodiments described herein, as discussed in more detail herein. In some embodiments, two or more AI / ML models 132 may be chained (e.g., in series, in parallel, or a combination thereof) such that they collectively provide collaborative output(s). The AI / ML models 132 may perform or assist with computer vision (CV), optical character recognition (OCR), document processing and / or understanding, semantic learning and / or analysis, analytical prediction, process discovery, task mining, testing, automated RPA workflow generation, sequence extraction, clustering detection, speech-to-text translation, combinations of any of these, etc. However, any desired number and / or type(s) of AI / ML models may be used without departing from the scope of one or more embodiments described herein. By using multiple AI / ML models, for example, a system can develop an overall picture of what is happening on a given computing system. For example, one AI / ML model can perform OCR, another can detect buttons, another can compare sequences, etc. Patterns may be determined individually by an AI / ML model or collectively by multiple AI / ML models. In certain embodiments, one or more AI / ML models are deployed locally on at least one of the computing systems 102, 104, 106.
[0038] In some embodiments, multiple AI / ML models 132 may be used. Each AI / ML model 132 is an algorithm (or model) that executes on data, and the AI / ML model itself may be, for example, a deep learning neural network (DLNN) of artificial "neurons" trained on training data. In some embodiments, the AI / ML model 132 may have multiple layers that perform various functions such as statistical modeling (e.g., hidden Markov model (HMM)), and may utilize deep learning techniques (e.g., long short-term memory (LSTM) deep learning, encoding of previous hidden states, etc.) to perform desired functions.
[0039] The hyper-automation system 100 may provide four main functional groups in some embodiments: (1) discovery, (2) automation construction, (3) management, and (4) engagement. Automation (e.g., executed on a user computing system, server, etc.) may be performed by software robots such as RPA robots in some embodiments. For example, attended robots, unattended robots, and / or test robots may be used. Attended robots collaborate with the user to assist the user in tasks (e.g., via UiPath Assistant™). Unattended robots operate independently of the user and may potentially execute in the background without the user's knowledge. Test robots are unattended robots that execute test cases against an application or RPA workflow. Test robots may be executed in parallel on multiple computing systems in some embodiments.
[0040] The discovery function can discover various opportunities for business process automation and provide automated recommendations therefor. Such a function can be implemented by one or more servers such as server 130. The discovery function may include, in some embodiments, providing an automation hub, process mining, task mining, and / or task capture. An automation hub (e.g., UiPath Automation Hub (trademark)) can provide a mechanism for managing an automation rollout with visibility and controllability. Automation ideas can be crowdsourced from employees, for example, via a submission form. A feasibility and return on investment (ROI) calculation for automating these ideas is provided, documentation for future automation is collected, and collaboration for quickly going from discovery to construction of the automation can be provided.
[0041] (e.g., via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) Process mining refers to the process of collecting and analyzing data from applications (such as enterprise resource planning (ERP) applications, customer relationship management (CRM) applications, email applications, call center applications, etc.) to identify what end-to-end processes exist in an organization, how to effectively automate them, and the impact of the automation. This data can be obtained, for example, from user computing systems 102, 104, 106 by a listener and processed by a server such as server 130. In some embodiments, one or more AI / ML models 132 can be employed for this purpose. This information can be exported to an automation hub to speed up implementation and avoid manual information transfer. The goal of process mining can be to increase business value by automating processes within an organization. Some examples of the goals of process mining include, but are not limited to, increased profits, improved customer satisfaction, regulatory and / or compliance, and improved employee efficiency.
[0042] (For example, via UiPath Automation Cloud™ and / or UiPath AI Center™) Task mining identifies and aggregates workflows (e.g., employee workflows), then applies AI to reveal patterns and variations in daily tasks, and scores such tasks for ease of automation and potential savings (e.g., time and / or cost savings). One or more AI / ML models 132 may be employed to reveal repetitive task patterns in the data. Repetitive tasks ripe for automation can then be identified. This information can first be provided by a listener and, in some embodiments, analyzed on a server of a core hyperautomation system 120 such as server 130. Discoveries from task mining (e.g., extensive application markup language (XAML) process data) are exported to a process document or a designer application such as UiPath Studio™ to enable faster creation and deployment of automation.
[0043] Task mining in some embodiments can include taking screenshots with user actions (e.g., mouse click location, keyboard input, application windows and graphical elements the user interacted with, timestamps for the interaction, etc.), collecting statistical data (e.g., execution time, number of actions, text input, etc.), editing and annotating the screenshots, specifying the types of actions being recorded, etc.
[0044] (Via UiPath Automation Cloud (trademark) and / or UiPath AI Center (trademark)) Task capture automatically documents attended processes as the user works or provides a framework for unattended processes. Such documentation may include process definition documents (PDDs), skeleton workflows, capture of actions for each part of the process, recording of user actions and automatic generation of comprehensive workflow diagrams including details about each step, tasks that are desirably automated in formats such as Microsoft Word (registered trademark) documents, XAML files, etc. Constructible workflows may, in some embodiments, be directly exported to designer applications such as UiPath Studio (trademark). Task capture can simplify the requirements gathering process for both subject matter experts who describe the process and Center of Excellence (CoE) members who provide production-grade automation.
[0045] Automation can be achieved through designer applications (such as UiPath Studio™, UiPath StudioX™, UiPath Web™, etc.). For example, RPA developers at the PA development facility 150 can use the RPA designer application 154 on the computing system 152 to build and test automation for various applications and environments such as web, mobile, SAP®, and virtual desktop. API integration can be provided for various applications, technologies, and platforms. Pre-defined activities, drag-and-drop modeling, and workflow recorders can facilitate automation with minimal coding. The document understanding feature can be provided through drag-and-drop AI skills for data extraction and interpretation that call one or more AI / ML models 132. Such automation can process virtually any document type and format, including tables, checkboxes, signatures, and handwritten. When data is validated or exceptions are handled, this information may be used to retrain the respective AI / ML models, improving their accuracy over time.
[0046] With the integrated service, developers can seamlessly combine, for example, the automation of user interfaces (UIs) and APIs. Automations that require APIs or that span both API and non-API applications and systems can be built. A repository (e.g., UiPath Object Repository (trademark)) or marketplace (e.g., UiPath Marketplace (trademark)) of pre-built RPA and AI templates and solutions can be provided so that developers can automate a wide variety of processes more quickly. Thus, when building an automation, the hyper-automation system 100 can provide user interfaces, development environments, API integrations, pre-built and / or custom-built AI / ML models, development templates, integrated development environments (IDEs), and advanced AI capabilities. The hyper-automation system 100, in some embodiments, enables the development, deployment, management, configuration, monitoring, debugging, and maintenance of RPA robots, which can provide automation for the hyper-automation system 100.
[0047] In some embodiments, components of the hyperautomation system 100, such as designer applications and / or external rule engines, provide support for managing and enforcing governance policies for controlling the various functions provided by the hyperautomation system 100. Governance is the ability of an organization to introduce policies to prevent users from developing automations (such as RPA robots) that can harm the organization, such as violating the EU General Data Protection Regulation (GDPR), the U.S. Health Insurance Portability and Accountability Act (HIPAA), the terms of use of third-party applications, etc. Otherwise, developers could create automations that violate privacy laws, terms of use, etc. during the execution of their automations. Thus, some embodiments implement access control and governance restrictions at the robot and / or robot design application level. This can provide an additional level of security and compliance in the automation process development pipeline in some embodiments by preventing developers from introducing security risks or taking dependencies on unapproved software libraries that could operate in ways that violate policies, regulations, privacy laws, and / or privacy policies. See U.S. Non-Provisional Patent Application No. 16 / 924,499. This content is hereby incorporated by reference in its entirety.
[0048] The management function can provide organization-wide automation management, deployment, and optimization. The management function may include, in some embodiments, orchestration, test management, AI capabilities, and / or insights. The management function of the hyper-automation system 100 can also act as an integration point with third-party solutions and applications for automation applications and / or RPA robots. The management function of the hyper-automation system 100 can include, among other things, but not limited to, facilitating the provisioning, deployment, configuration, queuing, monitoring, logging, and interconnectivity of RPA robots.
[0049] Conductor applications such as UiPath Orchestrator (trademark) (which may be provided as part of UiPath Automation Cloud (trademark) in some embodiments, or on-premises, VM, private or public cloud, on a Linux (trademark) VM, or as a cloud-native single-container suite via UiPath Automation Suite (trademark)) provide orchestration capabilities to deploy, monitor, optimize, scale, and secure RPA robot deployments. A test suite (e.g., UiPath Test Suite (trademark)) can provide test management for monitoring the quality of deployed automation. The test suite can facilitate test planning and execution, requirement fulfillment, and defect traceability. The test suite may include comprehensive test reports.
[0050] Analytics software (e.g., UiPath Insights (trademark)) can track, measure, and manage the performance of deployed automation. The analytics software can align automation operations with specific key performance indicators (KPIs) and strategic outcomes of the organization. The analytics software can present results in dashboard form for easier understanding by human users.
[0051] A data service (e.g., UiPath Data Service (trademark)) can, for example, be stored in a database 140 and bring data into a single, scalable, and secure place using a drag-and-drop storage interface. Some embodiments may provide low-code or no-code data modeling and storage for automation while ensuring seamless access to data, enterprise-grade security, and scalability. AI capabilities may be provided by an AI Center (e.g., UiPath AI Center (trademark)), which facilitates the incorporation of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options may enable non-data scientists to access such capabilities. Deployed automation (e.g., an RPA robot) can call an AI / ML model from an AI Center such as AI / ML model 132. The performance of the AI / ML model can be monitored and trained and improved using human-verified data such as that provided by a data review center 160. A human reviewer may provide labeled data to the core hyperautomation system 120 via a review application 152 on a computing system 154. For example, a human reviewer may verify that the predictions by the AI / ML model 132 are accurate or, if not, provide corrections. This dynamic input may then be saved as training data for retraining the AI / ML model 132 and may be stored, for example, in a database such as database 140. The AI Center can then schedule and perform a training job to train a new version of the AI / ML model using the training data. Both positive and negative examples can be stored and used for retraining the AI / ML model 132.
[0052] The engagement function involves humans and automation as one team for seamless collaboration regarding a desired process. Low-code applications can be built (e.g., via UiPath Apps™) even if they lack APIs in some embodiments to connect browser tabs and legacy software. Applications can be quickly created using a web browser, for example, through a rich library of drag-and-drop controls. An application can be connected to one automation or multiple automations.
[0053] The Action Center (e.g., UiPath Action Center™) provides an easy and efficient mechanism for passing a process from automation to humans or vice versa. Humans can provide approvals or escalations and perform exception handling, etc. Then, the automation can execute the automated functions of a given workflow.
[0054] The local assistant can be provided as a launch pad for the user to launch automation (e.g., UiPath Assistant (trademark)). This feature can be provided, for example, in the tray provided by the operating system, and can enable the user to interact with RPA robots and RPA robot - enabled applications on their computing system. The interface can list the automations approved for a given user and enable the user to execute them. These can include off - the - shelf automations from an automation marketplace, an internal automation store in an automation hub, etc. When an automation is running, they can execute as a local instance in parallel with other processes on the computing system so that the user can use the computing system while the automation is performing its actions. In certain embodiments, the assistant is integrated with a task capture function so that the user can document the processes that will soon be automated from the assistant's launch pad.
[0055] Chatbots (e.g., UiPath Chatbots (trademark)), social messaging applications, and / or voice commands can enable the user to execute automations. This can simplify access to the information, tools, and resources necessary to conduct customer interactions or other activities. Human - to - human conversations can be easily automated in the same way as other processes. Trigger RPA robots launched in this way may be able to perform actions such as order status checks and data posting to CRM using plain - language commands.
[0056] End-to-end measurement of automation programs at any scale and governance can be provided by the hyper-automation system 100 in some embodiments. As such, analytics (e.g., via UiPath Insights™) may be employed to understand the performance of the automation. Data modeling and analytics that use any combination of available business metrics and operational insights can be used for various automation processes. Custom-designed and pre-built dashboards visualize data across desired metrics, discover new analytical insights, track performance indicators, discover ROI for automation, perform remote instrumentation monitoring on the user's computing system, detect errors and anomalies, and debug the automation. An automation management console (e.g., UiPath Automation Ops™) can be provided to manage the automation throughout its lifecycle. An organization can govern how the automations are built, what users can do with them, and which automations users can access.
[0057] The hyper-automation system 100 provides an iterative platform in some embodiments. Processes can be discovered, automations can be built, tested, and deployed, performance can be measured, use of the automation can be easily provided to users, feedback can be obtained, AI / ML models can be trained and retrained, and the process itself can be repeated. This facilitates a more robust and effective set of automations.
[0058] FIG. 2 is an architectural diagram showing an RPA system 200 according to one or more embodiments. In some embodiments, the RPA system 200 is part of the hyper-automation system 100 of FIG. 1. The RPA system 200 includes a designer 210 that enables developers to design and implement workflows. The designer 210 provides solutions for application integration and automates third-party applications, management information technology (IT) tasks, and business IT processes. The designer 210 may facilitate the development of automation projects that are graphical representations of business processes. Put simply, the designer 210 facilitates the development and deployment of workflows and robots (indicated by arrow 211). In some embodiments, the designer 210 may be an application that runs on a user's desktop, an application that runs remotely on a VM, a web application, and the like.
[0059] Automation projects enable the automation of rule-based processes by giving developers control over the execution order and relationships between a custom set of steps developed in a workflow as defined herein as an "activity" above. A commercial example of an embodiment of the designer 210 is UiPath Studio (trademark). Each activity may include actions such as clicking a button, reading a file, writing to a log panel, and the like. In some embodiments, workflows may be nested or embedded.
[0060] Some types of workflows may include, but are not limited to, sequences, flowcharts, finite state machines (FSMs), and / or global exception handlers. A sequence may be particularly suitable for a linear process that enables the flow from one activity to another without cluttering the workflow. A flowchart may be particularly suitable for more complex business logic and enables the integration of decision-making and the connection of activities in more diverse ways through multiple branching logic operators. An FSM may be particularly suitable for large-scale workflows. An FSM may use a finite number of states triggered by conditions (i.e., transitions) or activities during their execution. A global exception handler may be particularly suitable for determining the behavior of a workflow when an execution error is encountered or for debugging the process.
[0061] When a workflow is developed within the designer 210, the execution of the business process is coordinated by the conductor 220, which coordinates one or more robots 230 that execute the workflow developed within the designer 210. A commercial example of an embodiment of the conductor 220 is UiPath Orchestrator (trademark). The conductor 220 facilitates the management of the generation, monitoring, and deployment of resources in an environment. The conductor 220 may operate as an integration point with third-party solutions and applications. As such, in some embodiments, the conductor 220 may be part of the core hyperautomation system 120 of FIG. 1.
[0062] The conductor 220 can manage all the robots 230 and connect and execute the robots 230 from a central point (indicated by arrow 231). The types of robots 230 that can be managed include, but are not limited to, attended robots 232, unattended robots 234, development robots (similar to unattended robots 234 but used for development and testing purposes), and non-production robots (similar to attended robots 232 but used for development and testing purposes). The attended robot 232 is triggered by user events and operates alongside humans on the same computing system. The attended robot 232 can be used with the conductor 220 for centralized process deployment and logging media. The attended robot 232 may assist human users in achieving various tasks and may be triggered by user events. In some embodiments, the process cannot start from the conductor 220 on this type of robot and / or they cannot execute under a locked screen. In certain embodiments, the attended robot 232 can only be launched from a robot tray or from a command prompt. The attended robot 232 preferably operates under human supervision in some embodiments.
[0063] The unattended robot 234 can operate unmanned in a virtual environment and automate many processes. The unattended robot 234 can be responsible for providing remote execution, monitoring, scheduling, and work queue support. Debugging for all robot types can be performed by the designer 210 in some embodiments. Both attended and unattended robots can automate various systems and applications (shown by the dashed box 290) including, but not limited to, mainframes, web applications, VMs, enterprise applications (e.g., those generated by SAP®, SalesForce®, Oracle®, etc.), and computing system applications (e.g., desktop and laptop applications, mobile device applications, wearable computer applications, etc.).
[0064] The conductor 220 may have various capabilities including, but not limited to, provisioning, deployment, configuration, queuing, monitoring, logging, and / or providing interconnectivity (as indicated by arrow 232). Provisioning may include creating and maintaining a connection between the robot 230 and the conductor 220 (e.g., a web application). Deployment may include ensuring the correct delivery of a package version to the robot 230 assigned for execution. Configuration may include maintaining and delivering the robot environment and process configuration. Queuing may include providing management of queues and queue items. Monitoring may include tracking specific data of the robot and maintaining user permissions. Logging may include saving and indexing logs to a database (e.g., a Structured Query Language (SQL) database or a NoSQL database) and / or another storage mechanism (e.g., ElasticSearch® which provides the ability to store large datasets and execute queries quickly). The conductor 220 may provide interconnectivity by operating as a central point of communication for third - party solutions and / or applications.
[0065] The robot 230 is an execution agent that implements the workflow constructed by the designer 210. One commercial example of some embodiments of the robot(s) 230 is UiPath Robots™. In some embodiments, the robot 230, by default, installs the Microsoft Windows® Service Control Manager (SCM) management service. As a result, such a robot 230 can open an interactive Windows® session under a local system account and may have the rights of a Windows® service.
[0066] In some embodiments, the robot 230 can be installed in user mode. For such a robot 230, it means having the same rights as the user in which a given robot 230 is installed. This feature may also be available for high-density (HD) robots that ensure maximum utilization of each machine. In some embodiments, any type of robot 230 can be configured in an HD environment.
[0067] The robot 230 in some embodiments is divided into a plurality of components, each specialized for a specific automation task. Robot components in some embodiments include, but are not limited to, SCM management robot services, user mode robot services, an executor, an agent, and a command line. The SCM management robot service manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host (i.e., the computing system on which the robot 230 is executed). These services are entrusted with managing the qualification information of the robot 230. The console application is launched by the SCM under the local system.
[0068] The user mode robot service in some embodiments manages and monitors Windows® sessions and operates as a proxy between the conductor 220 and the execution host. The user mode robot service may be entrusted with managing the qualification information of the robot 230. If the SCM management robot service is not installed, a Windows® application can be automatically launched.
[0069] The executor can perform a job given under a Windows® session (i.e., can perform a workflow). The executor can recognize the dots per inch (DPI) setting per monitor. The agent can be a Windows® Presentation Foundation (WPF) application that displays jobs available in the system tray window. The agent can be a client of the service. The agent can request the start or stop of a job and the change of settings. The command line is a client of the service. The command line is a console application that can request the start of a job and wait for its output.
[0070] As described above, the fact that the components of the robot 230 are split helps developers, support users, and computing systems to more easily execute, identify, and track what each component is doing. In this way, special behaviors can be configured for each component, such as setting different firewall rules for the executor and the service. The executor can always, in some embodiments, recognize the DPI setting per monitor. As a result, the workflow can be performed at any DPI, regardless of the configuration of the computing system on which the workflow was created. Also, in some embodiments, projects from the designer 210 can be made independent of the browser zoom level. In the case of applications marked as not recognizing or intentionally not recognizing DPI, DPI can be disabled in some embodiments.
[0071] The RPA system 200 in this embodiment is part of a hyper-automation system. Developers can use the designer 210 to build and test RPA robots that utilize AI / ML models deployed in the core hyper-automation system 240 (e.g., as part of its AI center). Such RPA robots can send inputs for the execution of the AI / ML model(s) and receive outputs therefrom via the core hyper-automation system 240.
[0072] One or more robots 230 may be listeners, as described above. These listeners can provide information to the core hyper-automation system 240 regarding what the user is doing when using their computing system. This information can then be used by the core hyper-automation system for process mining, task mining, task capture, etc.
[0073] An assistant / chatbot 250 can be provided on the user computing system to enable the user to launch an RPA local robot. The assistant can be placed, for example, in the system tray. The chatbot can have a user interface so that the user can view the text of the chatbot. Alternatively, the chatbot can run in the background without a user interface and can listen for the user's utterances using the microphone of the computing system.
[0074] In some embodiments, data labeling can be performed by a user of the computing system that the robot is executing on, or on another computing system that the robot provides information to. For example, if the robot calls an AI / ML model to perform CV on an image for a VM user, but the AI / ML model does not correctly identify a button on the screen, the user can draw a rectangle around the mis-identified or non-identified component and potentially provide text with the correct identification. This information can be provided to the core hyper-automation system 240 and can then be used later for training a new version of the AI / ML model.
[0075] Figure 3 is an architectural diagram showing a deployed RPA system 300 according to one or more embodiments. In some embodiments, the RPA system 300 can be part of the RPA system 200 of FIG. 2 and / or the hyper-automation system 100 of FIG. 1. The deployed RPA system 300 can be a cloud-based system, an on-premises system, a desktop-based system that provides enterprise-level, user-level, or device-level automation solutions for the automation of different computing processes.
[0076] Note that either the client side 301, the server side 302, or both may include any desired number of computing systems without departing from the scope of one or more embodiments described herein. On the client side 301, the robot application 310 includes an executor 312, an agent 314, and a designer 316. However, in some embodiments, the designer 316 may not be running on the same computing system as the executor 312 and the agent 314. The executor 312 is executing a process. As shown in FIG. 3, multiple business projects may be executed simultaneously. The agent 314 (e.g., Windows® service) is, in this embodiment, a single connection point for all executors 312. All messages in this embodiment are logged into the conductor 340, which further processes them via a database server 355, an AI / ML server 360, an indexer server 370, or any combination thereof. As described above with respect to FIG. 2, the executor 312 may be a robot component.
[0077] In some embodiments, the robot represents an association between a machine name and a user name. The robot may manage multiple executors simultaneously. In a computing system (such as Windows® Server 2012) that supports multiple interactive sessions running simultaneously, multiple robots may be executed simultaneously, each running in a separate Windows® session using a unique user name. This is referred to as the HD robot described above.
[0078] Agent 314 is also responsible for transmitting the state of the robot (e.g., periodically sending a "heartbeat" message indicating that the robot is still functioning) and downloading the required version of the package to be executed. In some embodiments, the communication between agent 314 and conductor 340 is always initiated by agent 314. In a notification scenario, agent 314 may open a WebSocket channel that is later used by conductor 330 to send commands (e.g., start, stop, etc.) to the robot.
[0079] Listener 330 monitors and records data related to user interactions with the operation of the attended computing system and / or unattended computing system in which listener 330 resides. Listener 330 can be an RPA robot, a part of an operating system, a downloadable application for each computing system, or any other software and / or hardware without departing from the scope of one or more embodiments described herein. In fact, in some embodiments, the listener logic is implemented partially or fully via physical hardware.
[0080] In addition to the conductor 340, the server side 302 includes a presentation layer 333, a service layer 334, and a persistence layer 336. The presentation layer 333 may include a web application 342, an Open Data Protocol (OData) Representative State Transfer (REST) application programming interface (API) endpoint 344, and notifications and monitoring 346. The service layer 334 may include an API implementation / business logic 348. The persistence layer 336 may include a database server 355, an AI / ML server 360, and an indexer server 370. For example, the conductor 340 includes the web application 342, the OData REST API endpoint 344, notifications and monitoring 346, and the API implementation / business logic 348. In some embodiments, most actions performed by a user at the interface of the conductor 340 (e.g., via the browser 320) are performed by calling various APIs. Such operations may include, but are not limited to, launching jobs on a robot, adding / removing data in a queue, scheduling jobs to be executed unattended, etc., without departing from the scope of one or more embodiments described herein. The web application 342 may be the visual layer of the server platform. In this embodiment, the web application 342 uses Hypertext Markup Language (HTML) and JavaScript (JS). However, any desired markup language, scripting language, or any other format may be used without departing from the scope of one or more embodiments described herein. The user interacts with a web page from the web application 342 via the browser 320 in this embodiment to perform various operations for controlling the conductor 340. For example, the user may create a robot group, assign packages to robots, analyze logs for each robot and / or process, start and stop robots, etc.
[0081] In addition to web application 342, conductor 340 also includes a service layer 334 that exposes an OData REST API endpoint 344. However, other endpoints may be included without departing from the scope of one or more embodiments described herein. The REST API is consumed by both web application 342 and agent 314. Agent 314 is, in this embodiment, a supervisor for one or more robots on a client computer.
[0082] The REST API of this embodiment includes configuration, logging, monitoring, and queuing functions (at least as indicated by arrow 349). Configuration endpoints may be used, in some embodiments, to define and configure the users, permissions, robots, assets, releases, and environments of an application. The logging REST endpoint may be used to log various information, such as errors, explicit messages sent by robots, and other environment-specific information. The deployment REST endpoint may be used by robots to query the version of a package to be executed when a job start command is used in conductor 340. The queuing REST endpoint may be responsible for managing queues and queue items, such as adding data to a queue, retrieving transactions from a queue, and setting the status of a transaction.
[0083] Monitoring of the REST endpoints may monitor web application 342 and agent 314. The notification and monitoring API 346 may be a REST endpoint used for registration of agent 314, distribution of configuration settings to agent 314, and sending and receiving notifications from the server and agent 314. The notification and monitoring API 346 may use WebSocket communication in some embodiments. As shown in Figure 3, one or more activities / actions described herein are represented by arrows 350 and 351.
[0084] In some embodiments, the API of the service layer 334 can be accessed through the configuration of an appropriate API access path, for example, based on whether the conductor 340 and the overall hyper-automation system have an on-premises deployment type or a cloud-based deployment type. The API for the conductor 340 can provide custom methods for querying statistics regarding various entities registered with the conductor 340. Each logical resource may be an OData entity in some embodiments. In such an entity, components such as robots, processes, queues, etc. may have properties, relationships, and operations. The API of the conductor 340 can be consumed by the web application 342 and / or the agent 314 in two ways in some embodiments: by obtaining API access information from the conductor 340 or by registering an external application for using the OAuth flow.
[0085] In this embodiment, the persistent layer 336 includes three servers: a database server 355 (e.g., an SQL server), an AI / ML server 360 (e.g., a server that provides AI / ML model providing services such as an AI center function), and an indexer server 370. The database server 355 in this embodiment stores configurations such as robots, robot groups, related processes, users, roles, and schedules. In some embodiments, this information is managed via a web application 342. The database server 355 may manage queues and queue items. In some embodiments, the database server 355 may store (in addition to or instead of the indexer server 370) messages recorded by robots. The database server 355 may also store, for example, process mining, task mining, and / or task capture related data received from a listener 330 installed on the client side 301. Although no arrow is shown between the listener 330 and the database 355, it should be understood that in some embodiments, the listener 330 can communicate with the database 355 and vice versa. This data can be stored in the form of PDD, images, XAML files, etc. The listener 330 may be configured to eavesdrop on user actions, processes, tasks, and performance metrics on each computing system where the listener 330 resides. For example, the listener 330 may record user actions (e.g., clicks, typed characters, locations, applications, active elements, time, etc.) on its respective computing system and then convert them into a form suitable for being provided to and stored in the database server 355.
[0086] The AI / ML server 360 facilitates the integration of AI / ML models into automation. Pre-built AI / ML models, model templates, and various deployment options can enable non-data scientists to access such capabilities. Deployed automation (e.g., RPA robots) can call AI / ML models from the AI / ML server 360. The performance of AI / ML models can be monitored and trained and improved using human-verified data. The AI / ML server 360 can schedule and execute training jobs to train new versions of AI / ML models.
[0087] The AI / ML server 360 can store data related to AI / ML models and ML packages for configuring various ML skills for users during development. The ML skills used herein are, for example, pre-built and trained ML models for processes that can be used by automation. The AI / ML server 460 can also store data related to document understanding techniques and frameworks, algorithms, and software packages for various AI / ML capabilities, including but not limited to intent analysis, natural language processing (NLP), speech analysis, different types of AI / ML models, etc.
[0088] Optionally in some embodiments, the indexer server 370 stores information recorded by robots and creates an index. In certain embodiments, the indexer server 370 may be disabled via configuration settings. In some embodiments, the indexer server 370 uses ElasticSearch®, an open-source project full-text search engine. Messages recorded by robots (e.g., using activities such as log messages or line writes) may be sent to the indexer server 370 via logging REST endpoint(s), where they are indexed for future use.
[0089] FIG. 4 is an architectural diagram illustrating the relationships between a designer 410, activities 420, 430, 440, 450, a driver 460, an API 470, and an AI / ML model 480, according to one or more embodiments. As shown herein, a developer uses the designer 410 to develop a workflow to be performed by a robot. Various types of activities may be presented to the developer in some embodiments. The designer 410 may be local to or remote from the user's computing system (e.g., accessed via a local web browser that interacts with a VM or a remote web server). The workflow may include user-defined activities 420, API-driven activities 430, AI / ML activities 440, and / or UI automation activities 450. By way of example (as shown by the dashed lines), the user-defined activities 420 and API-driven activities 440 interact with an application via their APIs. The user-defined activities 420 and / or AI / ML activities 440 may then call one or more AI / ML models 480, which may be located locally to and / or remotely from the computing system on which the robot operates, in some embodiments.
[0090] In some embodiments, non-text visual components in an image can be identified, which is referred to herein as CV. CV can be at least partially performed by an AI / ML model(s) 480. Some CV activities related to such components can include, but are not limited to, extraction of text from segmented label data using OCR, fuzzy text matching, cropping of segmented label data using ML, comparison of the extracted text in the label data with ground truth data, etc. In some embodiments, the number of activities that can be implemented in the user-defined activity 420 can be in the hundreds or thousands. However, any number and / or type of activities can be used without departing from the scope of one or more embodiments described herein.
[0091] The UI automation activity 450 is a subset of special low-level activities described in low-level code that facilitate interaction with the screen. The UI automation activity 450 facilitates these interactions via a driver 460 that enables the robot to interact with the desired software. For example, the driver 460 can include an operating system (OS) driver 462, a browser driver 464, a VM driver 466, an enterprise application driver 468, and the like. In some embodiments, one or more AI / ML models 480 can be used by the UI automation activity 450 to perform interactions with the computing system. In certain embodiments, the AI / ML model 480 can enhance or completely replace the drivers 460. In fact, in certain embodiments, the drivers 460 are not included.
[0092] Driver 460 can interact with the OS at a low level, such as searching for hooks and monitoring keys, via the OS driver 462. Driver 460 may facilitate integration with Chrome (registered trademark), IE (registered trademark), Citrix (registered trademark), SAP (registered trademark), etc. For example, a "click" activity serves the same role in these different applications via driver 460.
[0093] Figure 5 is an architectural diagram showing a computing system 500 configured to provide automatic form input and ticket extraction from email according to one or more embodiments. In some embodiments, computing system 500 may be one or more computing systems depicted and / or described herein. In certain embodiments, computing system 500 may be part of a hyperautomation system as shown in FIGS. 1 and 2. Computing system 500 includes a bus 505 or other communication mechanism for communicating information, and one or more processors 510 coupled to bus 505 for processing information. The one or more processors 510 can be any type of general or special-purpose processor, including a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), multiple instances thereof, and / or any combination thereof. The one or more processors 510 may also have multiple processing cores, and at least some of the cores may be configured to perform specific functions. In some embodiments, multiple parallel processing may be used. In certain embodiments, at least one of the one or more processors 510 can be a neuromorphic circuit including processing elements that mimic biological neurons. In some embodiments, the neuromorphic circuit may not require the typical components of a von Neumann computing architecture.
[0094] Computing system 500 further includes a memory 515 for storing information and instructions to be executed by processor(s) 510. Memory 515 may be composed of random access memory (RAM), read-only memory (ROM), flash memory, cache, a static storage device such as a magnetic disk or optical disk, or other types of non-transitory computer-readable media, or any combination thereof. The non-transitory computer-readable media may be any available media accessible by processor(s) 510, and may include volatile media, non-volatile media, or both. Also, the media may be removable, non-removable, or both.
[0095] Furthermore, computing system 500 includes a communication device 520, such as a transceiver, to provide access to a communication network via a wireless and / or wired connection. In some embodiments, the communication device 520 is Frequency Division Multiple Access (FDMA), Single Carrier FDMA (SC-FDMA), Time Division Multiple Access (TDMA), Code Division Multiple Access (CDMA), Orthogonal Frequency Division Multiplexing (OFDM), Orthogonal Frequency Division Multiple Access (OFDMA), Global System for Mobile (GSM) communication, General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), cdma2000, Wideband CDMA (W-CDMA), High-Speed Downlink Packet Access (HSDPA), High-Speed Uplink Packet Access (HSUPA), High-Speed Packet Access (HSPA), Long Term Evolution (LTE), LTE Advanced (LTE-A), 802.11x, Wi-Fi, Zigbee, Ultra-WideBand (UWB) radio, 802.16x, 802.15, Home Node-B (HnB), Bluetooth, Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Near-Field Communications (NFC), 5th Generation (5G), New Radio (NR), any combination thereof, and / or may be configured to use any other currently existing or future-implemented communication standards and / or protocols without departing from the scope of one or more of the embodiments described herein.In some embodiments, the communication device 520 may include one or more antennas that are a single antenna, an array antenna, a panel antenna, a phased antenna, a switched antenna, a beamforming antenna, a beam steering antenna, combinations thereof, and / or any other antenna configuration without departing from the scope of one or more embodiments described herein.
[0096] The processor(s) 510 is further coupled via the bus 505 to a display 525 such as a plasma display, a liquid crystal display (LCD), a light emitting diode (LED) display, a field emission display (FED), an organic light emitting diode (OLED) display, a flexible OLED display, a flexible substrate display, a projection display, a 4K display, a high definition display, a Retina (registered trademark) display, an IPS (In-Plane Switching) display, or any other suitable display for presenting information to a user. The display 525 may be configured as a touch (haptic) display, a three-dimensional (3D) touch display, a multi-input touch display, a multi-touch display, etc. using a resistive method, a capacitive method, a surface acoustic wave (SAW) capacitive method, an infrared method, an optical imaging method, a scatter signal method, an acoustic pulse recognition method, a frustrated total internal reflection method, etc. Any suitable display device and haptic I / O may be used without departing from the scope of one or more embodiments described herein.
[0097] Keyboard 530 and cursor control devices 535, such as computer mice, touchpads, etc., are further coupled to bus 505 to enable a user to interface with computing system 500. However, in certain embodiments, there may be no physical keyboard and mouse, and the user can interact with the device only via display 525 and / or a touchpad (not shown). Any type and combination of input devices can be used as a matter of design choice. In certain embodiments, there is no physical input device and / or display. For example, the user may interact remotely with computing system 500 via another computing system communicating with it, or computing system 500 may operate autonomously.
[0098] Memory 515 stores software modules that provide functionality when executed by processor(s) 510. The modules include operating system 540 for computing system 500. The modules further include modules 545 configured to execute all or part of the processes described herein or derivatives thereof. Thus, module 545 represents, in this specification, a document understanding engine / module that implements digitalization, classification, pre-annotation, and verification of files, and sub-modules or functions therein (e.g., classification module). According to one or more embodiments, module 545 implements document understanding operations including receiving one or more files, automatically determining the domain of each of the one or more files from a predefined set of domains, and presenting the one or more file documents according to the domain in the user interface of the document understanding engine (e.g., predicting the domain from the content of each file). As an example, module 545 uses one or more GPTs to discover the type of document (e.g., domain or classification) to understand the document and discover fields within the document.
[0099] According to one or more embodiments, module 545 adapts or creates a model. According to one or more embodiments, the model of module 545 can include, but is not limited to, an AI model, a generative AI model, a language model, other AI / ML models, and other models (e.g., autoregressive language models, GPT, large language models (LLMs), text-to-text transfer transformers (T5) language models, etc.). According to one or more embodiments, module 545 generates a dataset for pre-training by cataloging data, versioning data, and leveraging existing datasets (e.g., a predefined set of domains).
[0100] For example, module 545 is a mechanism for adapting an AI model or creating an AI model for a specific domain (i.e., type / classification) that more accurately and efficiently classifies files and reduces the time and resources required for conventional document processing of scanned files (e.g., reducing the extraction time and effort of documents, thereby enabling processed files to be delivered more accurately and quickly). According to one or more embodiments, module 545 trains an LLM by using any task (e.g., translation, question answering, and classification) as an input feed and generating some target text as an output across a diverse set of tasks.
[0101] According to one or more embodiments, module 545 is a mechanism for validating files at runtime to reduce the time and resources required by module 545 itself (e.g., the document understanding engine / module itself). In an embodiment, module 545 can continuously train the LLM by, for example, leveraging feedback from runtime validation.
[0102] According to one or more embodiments related to the discovery of types, when an existing model cannot identify the type of a document, module 545 may utilize GPT, LLM, or other models to identify the document and determine all relevant contexts from the extracted document. Further, module 545 can construct a taxonomy and label data fields for the taxonomy. According to one or more embodiments related to the discovery of fields, module 545 may discover new items / objects / things on a document and propose the new items / objects / things for user interface extension. Further, module 545 can predict fields from a document and add them to the taxonomy.
[0103] Computing system 500 may include one or more additional functional modules 550 that include additional functionality.
[0104] One of ordinary skill in the art will understand that the "system" can be embodied as a server, an embedded computing system, a personal computer, a console, a personal digital assistant (PDA), a mobile phone, a tablet computing device, a quantum computing system, or any other suitable computing device, or a combination of devices, without departing from the scope of one or more embodiments described herein. Presenting the functions described above as being performed by a "system" is not intended to limit the scope of the embodiments described herein in any way, but rather to provide an example of many embodiments. In fact, the methods, systems, and devices disclosed herein may be implemented in a localized and distributed form consistent with computing techniques including cloud computing systems. The computing system may be part of, or otherwise accessible by, a local area network (LAN), a mobile communication network, a satellite communication network, the Internet, a public cloud or a private cloud, a hybrid cloud, a server farm, or any combination thereof. Any local or distributed architecture may be used without departing from the scope of one or more embodiments described herein.
[0105] It should be noted that some of the system features described herein are presented as modules to emphasize implementation independence. For example, a module can be implemented as a hardware circuit including off-the-shelf semiconductors such as custom very large scale integration (VLSI) circuits or gate arrays, logic chips, transistors, or other discrete components. Additionally, a module can be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, graphics processing units, and the like.
[0106] The module may also be implemented at least in part as software for execution by various types of processors. For example, a specified unit of executable code may include one or more physical or logical blocks of computer instructions that may be organized, for example, as objects, procedures, or functions. Nevertheless, a specified module that is executable need not be physically located together and may include modules when logically combined and may include separate instructions stored in different locations to achieve the purpose stated for the module. Further, the module may be stored on a non-transitory computer-readable medium such as, for example, a hard disk drive, a flash device, RAM, a tape, and / or any other non-transitory computer-readable medium used to store data without departing from the scope of one or more embodiments described herein.
[0107] In fact, a module of executable code may be a single instruction, or a number of instructions, and may even be distributed across multiple different code segments, between different programs, and between multiple memory devices. Similarly, the operational data may be identified within the module, may be shown herein, may be embodied in any suitable form, and may be organized within any suitable type of data structure. The operational data may be collected as a single data set, or may be distributed in different locations across different storage devices and may exist, at least in part, simply as electronic signals on a system or network.
[0108] Without departing from the scope of one or more embodiments described herein, various types of AI / ML models can be trained and deployed. For example, FIG. 6 shows an example of a neural network 600 trained to recognize graphical elements within an image according to one or more embodiments. Here, the neural network 600 receives pixels (shown in column 610) of a 1920×1080 screen screenshot image as input to the input “neurons” 1 to I of the input layer (shown in column 620). In this case, I is 2,073,600, which is the total number of pixels in the screenshot image.
[0109] The neural network 600 also includes a number of hidden layers (represented in columns 630 and 640). Both deep learning neural networks (DLNNs) and shallow learning neural networks (SLNNs) typically have multiple layers, but an SLNN may sometimes have only one or two layers and is usually less than a DLNN. Typically, a neural network architecture includes an input layer, a plurality of intermediate layers (e.g., hidden layers), and an output layer (represented in column 650), as in the case of the neural network 600.
[0110] Often, a DLNN has many layers (such as 10, 50, 200), and subsequent layers typically reuse functions from previous layers to compute more complex and general functions. On the other hand, an SLNN has only a few layers and tends to be trained relatively quickly because expert functions are pre-created from raw data samples. However, feature extraction is cumbersome. On the other hand, a DLNN usually does not require expert functions but takes time to train and tends to have more layers.
[0111] In either approach, the layers are trained simultaneously on a training set and usually checked for overfitting on a separate cross-validation set. Excellent results are obtained with both techniques, and there is considerable enthusiasm for both approaches. The optimal size, shape, and number of individual layers depend on the problem being addressed by each neural network.
[0112] Returning to FIG. 6, the pixels provided as the input layer are supplied as input to the J neurons of hidden layer 1. In this example, all pixels are supplied to each neuron, but are not limited to, feedforward networks, radial basis networks, deep feedforward networks, deep convolutional inverse graphics networks, convolutional neural networks, recurrent neural networks, artificial neural networks, long-term / short-term memory networks, gated recurrent unit networks, generative adversarial networks, liquid state machines, autoencoders, variational autoencoders, denoising autoencoders, sparse autoencoders, extreme learning machines, echo state networks, Markov chains, Hopfield networks, Boltzmann machines, restricted Boltzmann machines, deep residual networks, Kohonen networks, deep belief networks, deep convolutional networks, support vector machines, neural Turing machines, or any other suitable type or combination of neural networks that do not depart from the scope of one or more embodiments described herein. Various architectures can be used, either individually or in combination.
[0113] Hidden layer 2 (630) receives input from hidden layer 1 (620), hidden layer 3 receives input from hidden layer 2 (630), and so on, until the last hidden layer (represented by ellipse 655) provides its output as input to the output layer. The same is done for all hidden layers. Note that the numbers of neurons I, J, K, and L are not necessarily equal, and thus any desired number of layers can be used for a given layer of neural network 600 without departing from the scope of one or more embodiments described herein. In fact, in certain embodiments, the types of neurons in a given layer may not all be the same.
[0114] The neural network 600 is trained to assign a confidence score to graphical elements that are thought to be found within an image. To reduce matchings with unacceptably low likelihoods, in some embodiments, only those results having a confidence score above a confidence threshold may be provided. For example, if the confidence threshold is 80%, outputs having a confidence score exceeding this amount may be used and the rest may be ignored. In this case, the output layer indicates that two text fields (represented by outputs 661 and 662), a text label (represented by output 663), and a send button (represented by output 665) were found. The neural network 600 may provide the positions, dimensions, images, and / or confidence scores of these elements without departing from the scope of one or more embodiments described herein, which may then be used by an RPA robot or another process that uses this output for a given purpose.
[0115] Note that a neural network is typically a probabilistic construct that has a confidence score. This can be a score learned by an AI / ML model based on the frequency with which similar inputs were correctly identified during training. For example, a text field often has a rectangular shape and a white background. The neural network can learn to identify graphical elements having these features with a high degree of confidence. Common types of confidence scores include decimal numbers between 0 and 1 (interpretable as a percentage of confidence), numbers between negative infinity and positive infinity, or a set of expressions (e.g., "low", "medium", and "high"). Also, various post-processing calibration techniques such as temperature scaling, batch normalization, weight decay, negative log likelihood (NLL), etc., may be employed as an attempt to obtain a more accurate confidence score.
[0116] The "neurons" of a neural network are usually mathematical functions based on the functions of biological neurons. Neurons receive weighted inputs and have a sum and activation function that governs whether they pass the output to the next layer. This activation function can be a non-linear thresholded activity function that does nothing if the value is below the threshold, and responds linearly when the function exceeds the threshold (i.e., rectified linear unit (ReLU) non-linearity). Since actual neurons can have a nearly identical activity function, the sum function and ReLU function are used in deep learning. Through linear transformation, information can be subtracted, added, etc. Essentially, neurons function as gating functions that pass the output to the next layer governed by their underlying mathematical functions. In some embodiments, different functions can be used for at least some of the neurons.
[0117] JPEG2025109160000001.jpg55162
[0118] JPEG2025109160000002.jpg41161
[0119] JPEG2025109160000003.jpg26132
[0120] In this case, neuron 700 is a single-layer perceptron. However, any suitable neuron type or combination of neuron types can be used without departing from the scope of one or more embodiments described herein. It should also be noted that the range of values of the weights and / or output value(s) of the activation function can be different in some embodiments without departing from the scope of one or more embodiments described herein.
[0121] For example, in the case where the identification of graphical elements within an image is successful, a goal, or "reward function", is often used. The reward function guides the search in the state space and attempts to achieve a goal (e.g., successful identification of graphical elements, successful identification of the next sequence of activities in an RPA workflow, etc.) by using both short-term and long-term rewards to explore intermediate transitions and steps.
[0122] During training, various labeled data (in this case, images) are fed through the neural network 600. When the identification is successful, the weights of the inputs to the neurons are strengthened, while when the identification fails, those weights are weakened. A cost function such as mean squared error (MSE) or gradient descent can be used to make slightly incorrect predictions cost much less than greatly incorrect predictions. If the performance of the AI / ML model does not improve after a certain number of training iterations, the data scientist can change the reward function, indicate where un-identified graphical elements are, provide corrections for mis-identified graphical elements, etc.
[0123] Backpropagation is a technique for optimizing the synaptic weights in a feedforward neural network. Backpropagation can be used to "pop up" the hidden layers of the neural network to see how much loss each node is bearing, and then assign low weights to nodes with a high error rate and vice versa, to update the weights to minimize the loss. That is, backpropagation enables the data scientist to repeatedly adjust the weights so as to minimize the difference between the actual output and the desired output.
[0124] The algorithm of backpropagation is mathematically based on optimization theory. In supervised learning, training data with known outputs are passed through the neural network, the error is calculated using a cost function from the known target outputs, and this gives the error for backpropagation. The error is calculated at the output, and this error is converted into a correction of the weights of the network that minimizes the error.
[0125] JPEG2025109160000004.jpg46162
[0126] JPEG2025109160000005.jpg18162
[0127] JPEG2025109160000006.jpg87162
[0128] JPEG2025109160000007.jpg90119
[0129] JPEG2025109160000008.jpg65164
[0130] The AI / ML model can be trained over multiple epochs until it reaches a good level of accuracy (e.g., above 97% using an F2 or F4 threshold for detection, about 2000 epochs). This level of accuracy can be determined in some embodiments using an F1 score, an F2 score, an F4 score, or any other suitable technique that does not deviate from the scope of one or more of the embodiments described herein. Once trained on training data, the AI / ML model can be tested on a set of evaluation data that the AI / ML model has not previously encountered. This helps ensure that the AI / ML model does not "overfit" such that it performs well at identifying graphical elements in the training data but does not generalize well to other images.
[0131] In some embodiments, it may not be known what accuracy levels an AI / ML model can achieve. Thus, when the accuracy of an AI / ML model begins to decline when analyzing evaluation data (i.e., the model performs well on training data but its performance is starting to degrade on evaluation data), the AI / ML model can undergo additional training epochs on the training data (and / or new training data). In some embodiments, the AI / ML model is deployed only when the accuracy reaches a certain level or when the accuracy of the trained AI / ML model is better than that of an existing deployed AI / ML model.
[0132] In certain embodiments, the collection of the trained AI / ML models can be used to implement tasks such as employing an AI / ML model for each type of target graphical element, performing OCR by employing an AI / ML model, deploying yet another AI / ML model to recognize proximity relationships between graphical elements, and employing yet another AI / ML model to generate an RPA workflow based on the output from other AI / ML models. For example, this can enable semantic automation collectively by the AI / ML models.
[0133] In some embodiments, a transformer network such as SentenceTransformers™, a Python™ framework for state-of-the-art sentence, text, and image embeddings, can be used. Such a transformer network learns associations of words and phrases with both high and low scores. This trains the AI / ML model to determine what is close to the input and what is not, respectively. Instead of using only word / phrase pairs, the transformer network may also use field length and field type.
[0134] Figure 8 is a flowchart showing a process 800 for training an AI / ML model(s) according to one or more embodiments. Note that process 800 can also be applied to other UI learning operations such as NLP and chatbots. The process begins by training, in block 810, on data that provides labeled data such as labeled screens (e.g., with specified graphical elements and text), words and phrases, a "thesaurus" of semantic relatedness between words and phrases such that similar words and phrases can be identified for a given word or phrase, etc., as shown in FIG. 8. The nature of the training data provided depends on the purpose the AI / ML model is to achieve. The AI / ML model is then trained in block 820 over a plurality of epochs, and the results are reviewed in block 830.
[0135] If the AI / ML model does not meet the desired confidence threshold at decision block 840 (process 800 proceeds according to the no arrow), the training data is supplemented and / or the reward function is modified in block 850 to help the AI / ML model better achieve its purpose, and the process returns to block 820. If the AI / ML model meets the confidence threshold at decision block 840 (process 800 proceeds according to the yes arrow), the AI / ML model is tested against evaluation data in block 860 to confirm that the AI / ML model generalizes well and does not overfit with respect to the training data. The evaluation data may include screens, source data, etc. that the AI / ML model has not previously processed. If the confidence threshold for the evaluation data is met at decision block 870 (process 800 proceeds according to the yes arrow), the AI / ML model is deployed in block 880. Otherwise (process 800 proceeds according to the no arrow), the process returns to block 880 and the AI / ML model is further trained.
[0136] FIG. 9 is a flowchart showing a process 900 of document classification operations for digitizing, classifying, pre-annotating, and verifying files according to one or more embodiments. The process 800 executed in FIG. 8 and the process 900 executed in FIG. 9 can be executed by a document understanding engine implemented in a computer program according to one or more embodiments. The computer program may be stored in a non-transitory computer-readable medium. The computer-readable medium may be, but is not limited to, a hard disk drive, a flash device, RAM, a tape, and / or any other such medium or combination of media used to store data. The computer program may include coded instructions for controlling a processor(s) of a computing system (e.g., module 545 of FIG. 5) to implement all or part of the processes described in FIGS. 8-9, which may also be stored in a computer-readable medium.
[0137] The process 900 begins at block 930, where the document understanding engine receives one or more files. According to one or more embodiments, the document understanding engine can receive one or more files from any source with which it is communicating. The document understanding engine can receive uploads from a database, scans of bulk scans from a device (e.g., from a scanner or other device), registration by drag and drop, and other transfer mechanisms.
[0138] According to one or more embodiments, a document understanding project of a document understanding engine may receive one or more files. By way of example, the document understanding engine may receive one or more files via a drag-and-drop operation in which the one or more files are placed in the first folder or space of a graphical user interface (GUI), selected while within the first folder or space, dragged onto the GUI of the document understanding project, and dropped onto the document understanding project of the document understanding engine. Thus, the document understanding engine provides, by way of a combination of both hardware and software, one or more GUIs presenting a document understanding project for receiving one or more files, and a standardized domain presentation for displaying or updating one or more files within a determined domain (e.g., a document understanding project including a predefined set of domains of the document understanding engine receives one or more files).
[0139] The one or more files include electronic files in any format (e.g., non-standard organization). The format of the one or more files includes, but is not limited to, image file formats (e.g., JPEG, PNG, GIF, etc.), document or word processor file formats (e.g., Microsoft® Word, portable document format (PDF), rich text format (RFT), open document format (ODF), etc.), text file formats, handwritten document formats (e.g., whether scanned or electronically generated), electronically generated form formats, other formats, and combinations thereof. Each of the one or more files includes content such as characteristics, format, fields, annotations, internal labels, descriptions, information, data points, and other content.
[0140] Thus, while individual files can be of a specific format type, it is rare, if ever, for that format type to pervade the entirety of one or more files received by a conventional document processing or document understanding engine. In this regard, these one or more files are often in a non-standard format selected by either the hardware or software platform used to store these one or more files on this storage, stored locally on a computer, or remotely stored in a server / cloud environment. In some cases, the one or more files are electronic copies of paper documents scanned while ignoring format, annotations, labeling, or explanations, and are electronic files created continuously while ignoring the structure of the electronic copies. Thus, for a business, it is difficult to process one or more files stored in a branched format using conventional document processing for the problems discussed herein. This can cause problems in the management of, for example, medical records, insurance certificates, bank accounts, educational records, accounting documents, and thousands of paper documents and electronic files such as other consumer / customer / company / patient information. Currently, businesses must continuously and manually monitor thousands of paper documents and one or more electronically stored files to maintain information, and the information is often incomplete because records in separate locations (e.g., different consumer / customer / company / patient information) are not timely, or cannot be easily shared, or cannot be integrated due to format inconsistencies (e.g., non-standard organization). To solve this problem and other problems, a document understanding engine can receive / collect (e.g., convert and integrate as further described herein) one or more files in an inconsistent format from various sources into a standardized domain presentation. Thus, the document understanding engine provides one or more GUIs that present a standardized domain presentation to a user to provide access for displaying or updating one or more files within a determined domain through a combination of both hardware and software.
[0141] In block 950, the document understanding engine automatically determines the domain of each of one or more files. The document understanding engine may perform the determination using a model. The model may be one or a combination of an AI model, a generative AI model, a language model, other AI / ML models, and other models. According to one or more embodiments, the generative AI model of the document understanding engine determines the domain. The generative AI model operates in the background or backend of the document understanding engine. As an example, the generative AI model operates to predict a domain from a predefined set of domains. The domain may be determined from the predefined set of domains using the content of each file. Thus, the generative AI model intelligently identifies the content of each file considering the predefined set of domains and determines the domain most applicable to that file. Further, the document understanding engine can display the determination of the domain in a standardized domain presentation of one or more GUIs via one or more interface elements. As an example, the document understanding engine can move one or more files from a general folder / bin to their respective domain folders.
[0142] A domain is a class or division of files based on shared content. The content can include, but is not limited to, characteristics, formats, fields, annotations, internal labelings, descriptions, information, and other content. Examples of domains include, but are not limited to, invoices, purchase orders, bills, utility bills, identification documents, driver's licenses, vehicle documents, bank statements, citation documents, receipts, applications, claims, contact forms, emergency information forms, and other classes or divisions. Thus, by determining the domain of a file, a document understanding engine labels or classifies that file according to the shared content of that domain. Each domain differs by at least one characteristic. For example, the domains of identification documents and driver's licenses have substantially the same content (e.g., name, date of birth, residential address, height, weight, eye color, photo, issuing authority, card number, etc.), but the domain of an identification document can differ from the domain of a driver's license by an indication that the card is a driver's license (e.g., a driver's license has a "Driver's License" label). In some cases, multiple contents can distinguish one domain from another. For example, both invoices and bills contain content for the recipient and the payer. However, an invoice has company information in the recipient field, and a bill has company information in the payer field. A predefined set of domains includes a plurality of domains selected for application to one or more files. For example, a predefined set of domains can include more than 70 different domains that classify files more accurately and efficiently than conventional document processing and reduce the time and resources required for conventional document processing.
[0143] Turning to FIG. 10, a user interface 1000 according to one or more embodiments is shown. The user interface 1000 is a GUI of a document understanding engine. The user interface 1000 includes a first frame 1010 and a second frame 1020 that include one or more domain type folders (such as represented by the automatic classification folder 1022 and the claim folder 1024). The user interface 1000 also includes a classification evaluation frame 1030 that includes an evaluation 1032, a recommended action frame 1040, and a documentation frame 1050.
[0144] The first frame 1010 provides an operation menu that includes selectable elements. The selectable elements of the operation menu can include a "Build" interface element, a "Measure" interface element, a "Publish" interface element, a "Monitor" interface element, a "Project Settings" interface element, and a "Feedback" interface element. The "Build" interface element enables the document understanding engine to start, build, and edit a document understanding project (e.g., the document understanding project 1 shown in the second frame 1020). Further, the "Build" interface element enables the document understanding engine to define the domain / type of the document and train a classification / extraction model. The "Measure" interface element enables the document understanding engine to measure the performance of the classification / extraction model (e.g., at design time to obtain model metrics). The "Publish" interface element enables the document understanding engine to provide the classification / extraction model for consumption through the document understanding engine or an API (e.g., when the classification / extraction model is created). The "Monitor" interface element enables the document understanding engine to track the performance after the model is deployed (e.g., at runtime to obtain performance metrics) and obtain an audit trail of what the model has predicted. The "Project Settings" interface element enables the document understanding engine to configure OCR settings, role-based access, and other settings. The "Feedback" interface element enables the document understanding engine to provide an interface for the user to provide one or more inputs (e.g., text input describing an experience).
[0145] The second frame 1020 provides one or more domain type folders. The one or more domain type folders receive and present one or more files. The domain type folders include an automatic classification folder 1022, a claim folder 1024, an order form folder, an invoice folder, an identity certificate folder, and a bank statement folder. Each folder can be expanded to display the files therein (as shown, for example, by the automatic classification folder 1022) or folded to display only the header (as shown, for example, by the claim folder 1024).
[0146] The classification evaluation frame 1030 provides an evaluation 1032 of the document project. The evaluation 1032 can be indicated by graphical elements and / or alphanumerics along a range. The evaluation 1032 is automatically generated by a document understanding engine to indicate the quality of a document understanding project (for example, the document understanding project 1 shown in the second frame 1020). For example, the evaluation can be set on a scale from 0 to 100, with higher values indicating higher quality. If 288 domains are successfully determined for 300 files, the document understanding project can obtain an evaluation of 96. According to one or more embodiments, an evaluation is given to a model (for example, a classification model and an extraction model) to identify the quality of the model constructed by the document understanding engine. The evaluation (for example, poor, good, average) can be a combination of scores determined by one or more factors (for example, the evaluation is mapped to a score). For example, the classification model is scored over two factors, and the extraction model is scored over four factors.
[0147] The recommended action frame 1040 provides information that a user can take to continue building a document understanding project. For example, the document understanding engine has successfully processed 2,667 files (e.g., 2 claims, 541 purchase orders, 583 receipts, 641 identity documents, and 900 bank transaction statements). The automatic classification folder 1022 shows two sample documents (e.g., sample 1 and sample 2), which are unrecognized, have no date or tags, and an unknown status. The recommended action frame 1040 can instruct the user to review the two sample documents to assist in domain determination. Note that when 2,667 domains are determined for 2,669 files, the document understanding project 1 can receive an evaluation of 99.9.
[0148] The documentation frame 1050 provides the user with all the documents that are shared as reference materials.
[0149] In operation, the document understanding engine moves one or more files from a general folder / bin (e.g., external to the user interface 1000) to each of the domain type folders of the second frame 1020.
[0150] Returning to FIG. 9, at block 960, the document understanding engine performs pre-annotating one or more files. The content of the one or more files can be pre-annotated, for example, with one or more annotation proposals. Pre-annotating (e.g., annotation proposals) can be represented by different user interface elements or interface extensions including, but not limited to, underlines, graphics, shading, color-coding, etc. According to one or more embodiments, the content of the file can be pre-annotated with color-coded underlines and / or color-coded shapes. For example, all annotation proposals can be underlines that are converted to shapes upon confirmation / approval.
[0151] Pre-annotation can be performed using a taxonomy (such as a common field set). According to one or more embodiments related to the discovery of fields, module 545 can discover new items / objects / things on a document and propose the new items / objects / things to the user interface extension. Further, module 545 can predict fields from the document and add them to the taxonomy. The document understanding engine can include a plurality of common field sets, each of which corresponds to a domain such that all documents in that domain are similarly pre-annotated. Thus, according to one or more embodiments, the document understanding engine can pre-annotate when the file is sorted.
[0152] Pre-annotation can be performed by using a combination of a special model (e.g., GPT) and an LLM. According to one or more embodiments, when the special model does not know or recognize the content, the LLM can predict pre-annotating the content. Thus, the combination of the special model and the LLM automatically detects the most relevant content in the file and uses pre-annotation to visually suggest that this content is relevant. In this way, the combination of the specialized model and the LLM provides the technical effect and benefit of speeding up the annotation of documents by automatically labeling the content, regardless of whether the content is known or unrecognized. Further, by pre-annotating, the GUI of the document understanding engine becomes more intelligent and conversational, and since the document understanding engine presents the next steps and suggestions to the user, the user does not need to know how the document understanding engine functions.
[0153] Turning to FIG. 11, a user interface 1100 according to one or more embodiments is shown. The user interface 1100 is a GUI of a document understanding engine. The user interface 1100 includes a first frame 1110 and a second frame 1120. The first frame 1110 presents a pre-annotated document 1111 (e.g., a file). The document 1111 shown in the user interface 1100 is, for example, a claim. The second frame 1020 presents a list for pre-annotating, such as annotation 1, annotation 2, annotation 3, annotation 4, and annotation 5. The list for pre-annotating corresponds to specific interface extensions on the document 1111. For example, annotation 1 corresponds to interface extension 1151, annotation 2 corresponds to interface extension 1152, annotation 3 corresponds to interface extension 1153, annotation 4 corresponds to interface extension 1154, and annotation 5 corresponds to interface extension 1155. As shown in FIG. 11, the document understanding engine recognizes specific content of the document 1111 via a special model and identifies this content using interface extensions (e.g., label data fields to a taxonomy). For example, as determined by GPT, the underline of interface extension 1151 may be pre-annotated to label the filing date, the solid box of interface extension 152 is to be pre-annotated to label the claim name, the solid box of interface extension 1153 is to be pre-annotated to label the claim address, and the solid box of interface extension 1155 is to be pre-annotated to label the claim total. Further, the LLM uses the dashed box of interface extension 1156 to predict content (e.g., account number, A / C name, and bank details from payment information). Each prediction can be added to the taxonomy of the document understanding engine. Each prediction can further be identified and verified for the user.
[0154] Returning to FIG. 9, at block 970, the document understanding engine displays one or more files within the GUI. The one or more files are presented according to their determined domain. One or more annotation proposals (e.g., pre-annotating) are displayed for the one or more files for verification. While the one or more files are displayed in the GUI, the document understanding engine can receive one or more inputs to verify or override the domain and / or pre-annotating. As an example, if an override is provided for the determined domain, the document understanding engine automatically moves the file to a folder of the corresponding domain type. While the one or more files are displayed in the GUI, the document understanding engine can receive one or more inputs to tag the one or more files. Tags are a user interface feature that can include sub-classifications used to further organize the one or more files. The GUI can include scoring of the one or more files. For example, the scoring can be done for the set of all files presented in the GUI. In some cases, each file presented can have a sub-score that contributes to the scoring of the one or more files. The scoring can indicate the processing success rate of the document understanding engine for the one or more files with respect to the domain and / or annotation proposals.
[0155] Next, turning to FIG. 12, the flowchart shows a process 1200 according to one or more embodiments. Process 1200 begins at block 1210, where the document understanding engine generates a document understanding project (e.g., the document understanding project is created based on one or more inputs). The document understanding project includes a predefined set of domains. The predefined set of domains can be assigned to the document understanding project at creation time, for example, when a name is entered in the interface prompt for the document understanding project for creating the project.
[0156] In block 1220, the document understanding engine provides one or more user interfaces configured to receive one or more files (e.g., an interface is opened). For example, a document understanding project is displayed across one or more GUIs of the document understanding engine.
[0157] Turning to FIG. 13, a user interface 1300 according to one or more embodiments is shown. The user interface 1300 is a GUI of the document understanding engine. The user interface 1300 includes a first frame 1010 and a second frame 1320. The user interface 1000 also includes a recommended action frame 1340 and a documentation frame 1050. The first frame 1010 provides an action menu having a selected "build" interface element 1313 (as shown, for example, by a gray shading). Note that the document understanding engine has created and assigned a predefined set of domains to document understanding project 2. Thus, the second frame 1320 is empty and is ready to receive one or more files either by a drag-and-drop operation or an upload (e.g., "drag and drop your document here (click to upload)"). The recommended action frame 1340 provides information that the user can execute to continue creating document understanding project 2. For example, the recommended action frame 1340 instructs the user to "upload a file - upload your document sample to determine the appropriate document type."
[0158] Returning to FIG. 12, in block 1230, the document understanding engine receives one or more files. According to one or more embodiments, the document understanding engine can receive one or more files from any source with which it is communicating. The document understanding engine can receive uploads from a database, scans of bulk scans from a device (e.g., from a scanner or other device), registration by drag-and-drop, and other transfer mechanisms.
[0159] In block 1240, the document understanding engine digitizes one or more files. Digitization includes converting the files from their format to a usable data structure that can be processed by the document understanding engine. The format of the one or more files includes, but is not limited to, image file formats (e.g., JPEG, PNG, GIF, etc.), document or word processor file formats (e.g., Microsoft® Word, Portable Document Format (PDF), Rich Text Format (RFT), Open Document Format (ODF), etc.), text file formats, handwritten document formats (e.g., whether scanned or electronically generated), electronically generated form formats, other formats, and combinations thereof. In block 1245, the document understanding engine discovers what its model knows about the content of the one or more files. In this regard, the document understanding engine may provide context with a prompt regarding what the document understanding engine is doing (e.g., when the LLM understands the request and makes a classification looking at the name and description, e.g., makes a determination of the domain). Further, in block 1245, the document understanding engine trains the model. For example, the document understanding engine may train a special model based on the input to score and pre-annotate the files, i.e., improve the input. Further, if the input responds to annotation proposals (e.g., pre-annotating), the model may learn which annotation proposals failed or were incorrect.
[0160] In block 1250, the document understanding engine automatically determines the domain of each of one or more files. The document understanding engine may perform the determination using a model. The model may be one or a combination of an AI model, a generative AI model, a language model, other AI / ML models, and other models. According to one or more embodiments, the document understanding engine can move one or more files from a general folder / bin to their respective domain folders. FIG. 14 shows a user interface 1400 according to one or more embodiments. The user interface 1400 is a GUI of the document understanding engine. The user interface 1400 includes a first frame 1010 and a second frame 1420 including an auto-classification folder 1422 and a claims folder 1424. The user interface 1400 also includes a pop-up frame 1470. The pop-up frame 1470 includes a toggle arrow 1471 for switching between files, a domain determination 1472 (e.g., classification), a cancel interface button 1474, and a note interface button 1476.
[0161] In block 1255, the document understanding engine generates a custom domain type outside of a predefined set of domains. The custom domain type can be generated by the document understanding engine when an unknown file is discovered. According to one or more embodiments, the document understanding engine requires a name and a description. A user input request can be made. If no user input is received, the document understanding engine can propose and / or generate a name and a description using GTP (e.g., by using data of established domains).
[0162] According to one or more embodiments, if the document understanding engine does not recognize a file (not a general one such as a claim or an order form, but a file specific to an enterprise, country, or industry), the document understanding engine automatically identifies / detects that file. In conventional document processing, the user manually checks the file and extracts the necessary information, which is a cumbersome process that requires the creation and configuration of fields. In contrast, the document understanding engine uses GPT to automatically detect the most relevant data points on the file and visually propose those data points in the GUI. The visual representation can be interacted with. For example, when clicked, a tooltip is displayed that checks whether to create a field of this domain type (e.g., whether to add the field to the taxonomy). After this automatic detection and proposal, the user can further perform manual annotation. For example, the user can annotate the value of that field. Turning to FIG. 15, a user interface 1500 according to one or more embodiments is shown. The user interface 1500 is the GUI of the document understanding engine. The user interface 1500 includes a first frame 1510 showing pages 1511, 1512, 1513, and 1514, and a second frame 1120 showing file 1531. The second frame 1550 presents a list of pre-annotations, e.g., annotation 11, annotation 12, annotation 13, annotation 14, annotation 15, and annotation 16. The list of pre-annotations corresponds to specific interface extensions on document 1521. For example, annotation 11 corresponds to interface extension 1551, annotation 12 corresponds to interface extension 1552, annotation 13 corresponds to interface extension 1553, annotation 14 corresponds to interface extension 1554, annotation 15 corresponds to interface extension 1555, and annotation 16 corresponds to interface extension 1556. Annotation 13 prompts the user to confirm that the value of interface extension 1553 is the correct value.
[0163] Next, as indicated by arrow 1256, process 1200 may proceed to block 1260. Alternatively or in parallel, process 1200 may return to block 1245 as indicated by arrow 1257 and retrain the model with a new custom domain type.
[0164] At block 1260, the document understanding engine performs pre-annotating one or more files. The content of the one or more files is pre-annotated with various user interface elements. By pre-annotating, the GUI of the document understanding engine becomes more intelligent and conversational, and the user does not need to know how the document understanding engine functions because the document understanding engine presents the next steps and suggestions to the user.
[0165] According to one or more embodiments, pre-annotating includes at least two sets of information including annotation proposals (e.g., pre-annotating) and direct application. An annotation proposal may be a user interface element where the annotation proposal is deleted from the file if no direct application (e.g., manual input to approve the annotation proposal) is provided. An annotation proposal may be a user interface element that implements the annotation proposal from underlined to full-color rectangle when received / applied. An annotation proposal may be a user interface element that remains in the file if no action is taken. Direct application may be manually confirmed and made to match the received and accepted annotation proposal.
[0166] At block 1270, the document understanding engine displays one or more files within the GUI. The one or more files are presented according to their determined domains. While the one or more files are displayed in the GUI, the document understanding engine may receive one or more inputs to verify or override the domain and / or pre-annotating.
[0167] In block 1280, the document understanding engine enables annotations. For example, when a file is classified into a domain or pre-annotated, the user can then add annotations by the document understanding engine. The document understanding engine provides prompts so that the user can send input during browsing of the file and pre-annotating. The input may include confirmation of updates or explanations. In this regard, the document understanding engine may provide engine-assisted fine-tuning of one or more files. Next, process 1200 proceeds to block 1245 as indicated by arrow 1281 and may further update the model based on the annotations. As an example, FIG. 16 shows table 1600 according to one or more embodiments. Table 1600 provides examples of engine-assisted fine-tuning, such as examples of when and where the document understanding engine can provide prompts and what content those prompts include.
[0168] According to one or more embodiments, the document understanding engine includes technical effects, advantages, and benefits such as providing a unique user experience via the GUI because domain and / or annotation suggestions are automatically provided and placed. Further, the document understanding engine includes technical effects, advantages, and benefits such as providing automatic training to the model (e.g., "behind the scenes") so that the model grows and learns as corrections are made. Further, the document understanding engine includes technical effects, advantages, and benefits such as providing domain assignment and pre-annotation scoring and improved scoring, for example, in response to a low score, the document understanding engine can learn the correction content and update / refresh / improve the annotation suggestions. The document understanding engine includes technical effects, advantages, and benefits of providing a mechanism to prompt for direct instructions, receive feedback, and filter which corrections should be used for training the model (e.g., determining which 50 out of 1000 files train the model).
[0169] According to one or more embodiments, engine-assisted fine-tuning improves the overall accuracy of domain determination and reduces human error and human resource consumption. For example, engine-assisted fine-tuning during execution reduces the number of times a user verifies a file. Using GPT, user operations can be eliminated. Further, the panel of LLMs can be used to vote on model predictions, and engine-assisted fine-tuning may include user intervention where the majority of the LLMs do not agree on the prediction (e.g., to reduce overall user input). According to one or more embodiments, engine-assisted fine-tuning can identify points where the LLMs do not agree, determine the agreeing LLMs, and the agreeing LLMs can be the trusted LLMs.
[0170] According to one or more embodiments, the document understanding engine may include configuration logic. The configuration logic includes setting threshold logic regarding whether to request user verification. For example, based on the decision of whether human intervention is required, the user can configure threshold logic regarding the number of disagreeing LLMs.
[0171] The computer program can be implemented in hardware, software, or a hybrid implementation. The computer program can be composed of modules that communicate operably with each other and is designed to send information or instructions to a display. The computer program can be configured to operate on a general-purpose computer, an ASIC, or any other suitable device.
[0172] It will be readily understood that the components of the various embodiments, as generally described and illustrated herein, may be arranged and designed in a variety of different configurations. Accordingly, the detailed description of the embodiments as represented in the accompanying figures is not intended to limit the scope as claimed, but is merely representative of the selected embodiments.
[0173] The features, structures, or characteristics described throughout this specification may be combined in any suitable manner in one or more embodiments. For example, references throughout this specification to "certain embodiments", "some embodiments", or similar language mean that the particular features, structures, or characteristics described in connection with the embodiments are included in at least one embodiment. Thus, the appearances of "in certain embodiments", "in some embodiments", "in other embodiments", or similar language throughout this specification are not necessarily referring to the same group of all embodiments, and the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0174] It should be noted that references throughout this specification to features, advantages, or similar language do not mean that all of the realizable features and advantages should be, or are, in any single embodiment. Rather, the language referring to features and advantages is understood to mean that a particular feature, advantage, or characteristic described in connection with an embodiment is included in one or more embodiments. Thus, discussions of features and advantages throughout this specification, and similar language, may refer to the same embodiment, but not necessarily.
[0175] Furthermore, the features, advantages, and characteristics of one or more embodiments described in this specification can be combined in any appropriate manner. Those skilled in the relevant art will recognize that the present disclosure can be practiced without the specific features or advantages of one or more of the particular embodiments. In other instances, additional features and advantages may be recognized in certain embodiments but may not be present in all embodiments.
[0176] Those having ordinary skill in the art will readily understand that the present disclosure may be practiced using steps in a different order and / or using hardware elements of a different configuration than those disclosed. Accordingly, while the present disclosure has been described based on these preferred embodiments, it will be apparent to those skilled in the art that certain changes, modifications, and alternative configurations will become apparent while remaining within the spirit and scope of the present disclosure. Therefore, reference should be made to the appended claims to determine the scope of the present disclosure.
Claims
1. A method performed by a document understanding engine implemented as a computer program within a computing environment, the method comprising: pre-annotating one or more files including one or more annotation proposals by a special model of the document understanding engine; presenting, by the document understanding engine, to a user interface, the one or more files including the one or more annotation proposals for verification, the user interface including scoring of the one or more files; training, by the document understanding engine, the special model based on one or more inputs to improve the scoring and the pre-annotating, the one or more inputs including being received in response to the one or more annotation proposals.
2. The method of claim 1, wherein the pre-annotating of the file is performed by the special model and a language model.
3. The method of claim 2, wherein the language model predicts a field from a set of common fields for features within the one or more files when the special model cannot provide at least one of the one or more annotation proposals.
4. The method of claim 1, wherein the pre-annotating of the one or more files includes one or more color-coded underlines or outlines of features of the one or more files.
5. The method of claim 1, wherein the method includes pre-annotating the one or more files after the domain of the file is determined.
6. The method of claim 5, wherein the pre-annotating of the one or more files includes receiving the one or more inputs during browsing of the one or more annotation proposals from the pre-annotating, the one or more inputs including one or more confirmations or explanations for fine-tuning the document understanding engine during training.
7. The method of claim 1, wherein the one or more files are stored in a non-standardized format and received via a drag-and-drop operation.
8. The method of claim 1, wherein the one or more files are received from a device that scans one or more corresponding paper documents into a non-standardized format.
9. The method according to claim 1, wherein the document understanding engine provides the user interface for receiving the one or more files.
10. The method according to claim 1, including digitizing the one or more files from a non-standardized format into a data structure usable by the document understanding engine.
11. A computer program product including computer program code for a document understanding engine, the computer program code being pre-annotating one or more files including one or more annotation proposals by a special model of the document understanding engine, presenting, by the document understanding engine, to a user interface the one or more files including the one or more annotation proposals for verification, the user interface including scoring of the one or more files, training, by the document understanding engine, the special model based on one or more inputs to improve the scoring and the pre-annotating, the one or more inputs including being received in response to the one or more annotation proposals, and being executable by at least one processor to perform operations, a computer program product.
12. The computer program product according to claim 11, wherein the pre-annotating of the file is performed by the special model and a language model.
13. The computer program product according to claim 12, wherein the language model predicts a field from a set of common fields for features in the one or more files when the special model cannot provide at least one of the one or more annotation proposals.
14. The computer program product according to claim 11, wherein the pre-annotating of the one or more files includes one or more color-coded underlines or outlines of features of the one or more files.
15. The computer program product according to claim 11, wherein the operations further include pre-annotating the one or more files after the domain of the file is determined.
16. Pre-annotating the one or more files includes receiving the one or more inputs during browsing of the one or more annotation proposals from the pre-annotating, the one or more inputs including one or more confirmations or explanations for fine-tuning the document understanding engine during training, the computer program product according to claim 15.
17. The one or more files are stored in a non-standardized format and received via a drag-and-drop operation, the computer program product according to claim 11.
18. The one or more files are received from a device that scans one or more corresponding paper documents into a non-standardized format, the computer program product according to claim 11.
19. The document understanding engine provides the user interface for receiving the one or more files, the computer program product according to claim 11.
20. The operation further includes digitizing the one or more files from a non-standardized format into a data structure usable by the document understanding engine, the computer program product according to claim 11.