Systems and methods for analyzing a document
An AI-driven document analysis system simplifies complex documents by generating audio-visual summaries in user-selected languages, addressing misinterpretation and enhancing clarity for non-experts.
Patent Information
- Application Number
- PCT/US2025/016191
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2025-02-16
- Publication Date
- 2025-09-04
AI Technical Summary
Documents containing legal or complex terminology, such as contracts, often confuse individuals unfamiliar with the terminology, leading to misinterpretation and potential liability, especially across language barriers.
A document analysis system utilizing artificial intelligence to identify document elements, generate textual representations, and create audio-visual summaries in selected languages, simplifying complex content into layman's terms.
Enhances understanding of complex documents by transforming legal or technical language into everyday language, reducing misinterpretation and facilitating informed decision-making across language barriers.
Smart Images

Figure US2025016191_04092025_PF_FP_ABST
Abstract
Description
Systems and Methods for Analyzing a Document Cross Reference to Related Applications
[0001] Not Applicable. Field of the Invention
[0002] The present disclosure relates generally to the field of digital document processing. More particularly, this disclosure relates to methods and systems for digitally analyzing documents to make the document contextually understandable to the reader. Background
[0003] Documents such as contracts and other materials containing legal or complex terminology can be confusing to people, especially for those not versed or experienced with the particular terminology or legal provisions. Misinterpretation of explicit terms can lead to risk and liability for the persons relying on the particular document to establish a legally binding agreement.
[0004] In many business transactions, contract documents cannot be avoided. For example, in the real estate industry, many of the documents used for transactions are standard form documents with expressly binding clauses and legal terminology. Consumer misunderstanding of such documents can also be compounded by language barriers when the documents are not rendered in the consumer’s spoken language. A need remains for improved tools to enable better understanding of complex documents and bring clarity to the parties relying on such documents to make informed decisions. Summary
[0005] According to an aspect of the invention, a document analysis system includes at least one processor in communication with non-transitory memory having instructions, which when executed cause the processor to perform functions including to: input data representing a document; identify different elements in the document using the input data; process the identified different elements using an artificial intelligence engine to generate a textual representation of at least one of the identified different elements; digitally generate a video including audio data based on the generated textual representation, wherein the video is generated with audio data in a selected language; digitally create a text file representing a summarization of the elements in the document; and digitally create a video file representing a summarization of the elements in the document.
[0006] According to another aspect of the invention, a computer-implemented method for analyzing a document includes receiving, at one or more processors, data representing a document; identifying, at the one or more processors, different elements in the document using the received data; applying an artificial intelligence process to the identified different elements to generate a textual representation of at least one of the identified different elements; generating, at the one or more processors, a video including audio data based on the generated textual representation, wherein the video is generated with audio data in a selected language; creating, at the one or more processors, a text file representing a summarization of the elements in the document; and creating, at the one or more processors, a video file representing a summarization of the elements in the document.
[0007] Aspects of the invention may also include a tangible, non-transitory computer-readable medium including instructions for analyzing a document that, when executed by one or more processors, cause the one or more processors to: receive data representing a document; identify different elements in the document using the received data; apply an artificial intelligence process to the identified different elements to generate a textual representation of at least one of the identified different elements; generate a video including audio data based on the generated textual representation, wherein the video is generated with audio data in a selected language; create a text file representing a summarization of the elements in the document; and create a video file representing a summarization of the elements in the document. Brief Description of the Drawings
[0008] The following figures form part of the present specification and are included to further demonstrate certain aspects of the present disclosure and should not be used to limit or define the claimed subject matter. The claimed subject matter may be better understood by reference to one or more of these drawings in combination with the description of embodiments disclosed herein. Consequently, a more complete understanding of the present embodiments and further features and advantages thereof may be acquired by referring to the following description taken in conjunction with the accompanying drawings, in which like reference numerals may identify like elements, wherein:
[0009] FIG. 1 shows a communication network configuration according to an example embodiment of this disclosure.
[0010] FIG.2 shows a system for analyzing documents according to an example embodiment of this disclosure.
[0011] FIG. 3 shows a screen shot according to an example embodiment of this disclosure.
[0012] FIG. 4 is a flow chart of a method for document analysis via a software as a service model according to an example embodiment of this disclosure. Detailed Description
[0013] The foregoing description of the figures is provided for the convenience of the reader. It should be understood, however, that the embodiments are not limited to the precise arrangements and configurations shown in the figures. Also, the figures are not necessarily drawn to scale, and certain features may be shown exaggerated in scale or in generalized or schematic form, in the interest of clarity and conciseness.
[0014] While various embodiments are described herein, in the interest of clarity all features of an actual implementation may not be described in this specification. In the development of any such actual embodiment, numerous implementation-specific decisions may need to be made to achieve the design-specific goals, which may vary from one implementation to another. It will be appreciated that such a development effort, while possibly complex and time-consuming, would nevertheless be a routine undertaking for persons of ordinary skill in the art having the benefit of this disclosure.
[0015] FIG. 1 shows a system 100 consistent with example embodiments of this disclosure. The system 100 includes a communication network 110 that provides communication links between one or more computing devices such as a mobile smart phone 10A, a tablet computer 10B, and a desktop or laptop computer 10C. The computing devices are conventional devices equipped with a visual display. The communication network 110 may be the Internet, an intranet, a wired or wireless network, a Wi-Fi network, a cellular network, or any combination thereof. The system 100 includes architecture 112 including an application programming interface 114, a server 116 configured with one or more processors 118 and a memory module 120 (transitory and non- transitory). Some embodiments may also include a database 122. The architecture 112 may be implemented as a unitary structure (e.g., central server at a main office) or as a cloud-based architecture. Use of the term "cloud" in this context refers generally to conventional cloudcomputing, which is a paradigm of computing in which dynamically scalable and often virtualized resources may be provided as a service over the network 110.
[0016] Some embodiments of this disclosure may also be implemented for operation without direct use of the communication network 110. In such embodiments, the application programming interface 114 is resident in the non-transitory memory of the individual computing devices 10A, 10B, 10C and executable via the internal processor(s) in the devices. Updates to the application programming interface 114 on the computing devices 10A, 10B, 10C could be obtained via the communication network 110 if desired. Embodiments may be implemented using conventional memory constructs (e.g., local memory, virtual memory, and / or cloud-based memory).
[0017] The software constructs enabling the embodiments of this disclosure reside in the application programming interface 114. Embodiments of the software code may be implemented using conventional programming languages as known in the art (e.g., JAVA™, PYTHON™, C, C++, etc.). It will be appreciated by those skilled in the art that the application programming interface 114 may be implemented with a single software program or a group of programs designed to perform the activities of the disclosed embodiments. The architecture 112 may be implemented with conventional computer hardware (e.g., server systems) situated in one location or via a distributed cloud-based network.
[0018] Those skilled in the art will appreciate that the embodiments of this disclosure may be implemented in network 110 computing environments with many types of computer system configurations, including, desktop computers, laptop computers, personal computers, message processors, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, and conventional cellphones. Embodiments may be implemented in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network 110, both perform tasks. In a distributed system environment, the application programming interface 114 may be located in both local and remote memory storage devices. Embodiments may also be implemented in software as a service (SAAS) models wherein the software is centrally hosted (e.g., on architecture 112) and remotely accessed by users via the communication network 110 as disclosed herein.
[0019] FIG. 2 shows an example system 200 embodiment of this disclosure. A document 30 is input into a computing device such as a computer 10C configured with an application programming interface 114 or linked through the communication network 110 to the interface 114 as disclosed herein. The document 30 may be digitally transmitted as data input to the computing device 10C via the communication network 110, uploaded locally by a user 32 (e.g., via flash drive, CD, scanner, etc.), or accessed from a repository in the database 122 locally or via the network 110. The document 30 may be in a format from a conventional processing application (e.g., Word™, PDF™, etc.) and may represent generic documentation such as contracts, forms, addenda, personal documentation, etc. In some embodiments, the document 30 may be standardized form or contract documentation used in a particular industry or endeavor. In some embodiments, an SAAS model may be implemented wherein a user 32 is allotted credits or points for use to perform the disclosed operations on uploaded or input documents 30.
[0020] Once the document 30 data is input, the interface 114 calls a program module 202 configured to parse the document 30 and identify the different elements in the document. The module 202 processes the document 30 data to parse and identify the different elements, such as the title 204, document identifier 206, paragraphs 208, headings 210, section delineators 212, clauses 214, subclauses 216, etc. In some embodiments, the program module 202 is configured to tag document 30 elements identified as non-textual (e.g., schedules, tables, figures, etc.).
[0021] After the document 30 elements are parsed and identified, the document 30 element data 220 is passed to an artificial intelligence engine (AIE) 300. The AIE 300 may be linked with the interface 114 or linked into the system 200 architecture 112 via the communication network 110. The AIE 300 is configured to perform directed tasks on the input element data 220 to generate tailored outputs in phases. In a first phase 302, the AIE 300 is configured with instructions to process the element data 220 to generate textual representations Tn of the specific elements. For example, an embodiment may be implemented where the AIE 300 processes the element data 220 and renders a textual explanation or description T1 of a clause (e.g. 214) in the document 30, a textual hypothetical example T2of the element described in T1, and one or more textual frequently asked questions (FAQS) T3 associated with the element described in T1. In some embodiments, the AIE 300 is configured to produce the textual representations Tn using layman’s terms or more common everyday language representative of the often complex terminology in the original document 30.
[0022] In a second phase 304, the AIE 300 is configured with instructions to process the outputs from the first phase 302 to generate a video 306 with audio based on the generated textual representations Tn. In some embodiments, the AIE 300 is configured to provide a selection of languages from which to select the language in which the video 306 audio file will be generated. The textual representations Tn and the video 306 are presented in a viewing page 308 (further described below). It will be understood by those skilled in the art that conventional commercial artificial intelligence software may be used to implement embodiments of this disclosure.
[0023] In some embodiments, the system 200 is also configured to generate a summary document 310, providing a written summary (in layman’s terms) of the contents of the original document 30. Embodiments may also be implemented wherein the system 200 generates a separate video 312, with audio in a selected language, summarizing the contents of the original document 30. In some embodiments, a user 32 can set up a document 30 as a template to share with others, providing shared access using a secure login portal via the network 110. System 200 embodiments may also be implemented to incorporate and / or link with conventional software and apps as known in the art (e.g., DocuSign™, ChatBot™, etc.). With such embodiments, a user 32 can open a chat to communicate with an administrator or others linked into the system 200 via the network 110 (e.g., to make specific inquiries regarding document 30 contents via a chat portal). Other embodiments may also be implemented with the AIE 300 configured to convert written text to speech, or with the interface 114 linked with conventional software to perform the conversion as known in the art.
[0024] FIG.3 shows a sample screen shot of a viewing page 308 according to an embodiment of this disclosure. The page 308 displays a visual copy 320 of the original document 30 as input to the system 200 for analysis. A user can navigate through or “flip” the pages in multi-page documents 30. The page 308 also displays an element tag 322 with a pull down menu option for the user to select a specific element in the displayed document 30 for analysis. Another tag 324 provides a pull down menu option for the user to select a specific sub-element in the displayed document 30 for analysis. For example, a user can select to analyze element subclause B1(e.g., see 216 FIG. 2) in the displayed document 30. Upon such election, the textual explanation or description T1 (see FIG. 2) pertaining to that element is displayed. The written example T2 of the elected element is also displayed. The example T2provides a hypothetical example of a potential fact pattern specifically relating to the document 30 element selected via the tags 322, 324. Alanguage selection tag 326 provides a pull down menu offering different audio language options for the video 306 (see FIG. 2), which is also displayed in the viewing page 308. The FAQS (see FIG. 2) are also accessible from a tag T3 in the viewing page 308. It will be understood by those skilled in the art that other embodiments of the viewing page 308 can be rendered with the displayed items in different arrangements.
[0025] FIG.4 shows a flowchart of a method 400 for analyzing a document in accordance with embodiments of this disclosure. The example embodiment of FIG.4 is described as a SAAS model 405. It will be appreciated, however, that embodiments of this disclosure are not limited to any particular type of method or process model. At step 410 a user establishes a profile or logs in to access the system 200 via a computing device over the network 110 (see FIG. 1). At this step, some embodiments may provide the user with options to select from among different document analysis preferences. At step 415, the user inputs the document 30 to be analyzed. As described herein, a user can select from among the documents 30 stored in a database 112 or input their own document 30 to be analyzed. In some embodiments, the user can select documents 30 pertaining to a specific business, enterprise, endeavor, and / or jurisdiction (e.g., real estate documentation authorized by a particular jurisdiction). In such embodiments, a user can select the document category from a library bank in the database 112.
[0026] At step 420, the input document 30 is parsed to identify the different elements as described herein. At step 425, the identified elements are processed by the AIE 300. At step 430, the videos 306 are generated using the AIE 300 outputs. At step 435, the document 30 summaries (310, 312 FIG. 2) are generated. Users can select to have the summaries of step 435 sent to them via email if desired. At step 440, the user can navigate through the document 30 and rendered analysis aspects as disclosed herein (see FIG. 3). All of the files and generated work product can be securely stored in a personal file in the database 112 for user access.
[0027] Advantages of the disclosed invention include an artificial intelligence enabled means to simplify the complexity of materials containing legal or complex terminology with easy to understand, multi-language guidance and contextual understanding. The disclosed systems and methods transform complex documents into everyday language or layman’s terms to empower clarity and understanding. System embodiments may be implemented to permit document sharing between selected users across the network. For example, embodiments of the disclosed systemsand methods can be implemented to arrange groups or teams of authorized users that can perform the documentation techniques disclosed herein on selected documents or groups of documents.
[0028] It will be understood by those skilled in the art that embodiments of this disclosure may be implemented using conventional software and computer systems programmed to perform the disclosed processes and operations. A usable computer-readable medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer- readable medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, and any suitable combination of the foregoing.
[0029] In light of the principles and example embodiments described and illustrated herein, it will be recognized that the example embodiments can be modified in arrangement and detail without departing from such principles. For example, alternative embodiments may include processes that use fewer than all of the disclosed operations, processes that use additional operations, and processes in which the individual operations disclosed herein are combined, subdivided, rearranged, or otherwise altered. In view of the wide variety of useful permutations that may be readily derived from the example embodiments described herein, this description is intended to be illustrative only, and should not be taken as limiting the scope of the invention. What is claimed as the invention, therefore, are all implementations that come within the scope of the following claims, and all equivalents to such implementations.
Claims
CLAIMS What is claimed is:
1. A document analysis system comprising: at least one processor in communication with non-transitory memory having instructions, which when executed cause the processor to perform functions including to: input data representing a document; identify different elements in the document using the input data; process the identified different elements using an artificial intelligence engine to generate a textual representation of at least one of the identified different elements; digitally generate a video including audio data based on the generated textual representation, wherein the video is generated with audio data in a selected language; digitally create a text file representing a summarization of the elements in the document; and digitally create a video file representing a summarization of the elements in the document.
2. The system of claim 1 wherein the generated textual representation comprises a description of the at least one identified element converted to layman’s terms.
3. The system of claim 2 wherein the generated textual representation comprises a hypothetical example relating to the at least one identified element.
4. The system of claim 3 wherein the generated textual representation comprises at least one frequently asked question relating to the at least one identified element.
5. The system of claim 4 wherein the functions performed by the at least one processor further include a function to display on a screen the generated video, the description, the hypothetical example, and the at least one frequently asked question.
6. The system of claim 2 wherein the functions performed by the at least one processor further include a function to display on a screen the created text file and video file.
7. The system of claim 2 wherein the functions performed by the at least one processor further include a function to display on a screen the generated video and / or electronically transmit the generated video.
8. The system of claim 2 wherein the identified different elements comprise a title, paragraph, clause, or subclause.
9. The system of claim 1 wherein the data representing a document is associated to a document input by a user.
10. The system of claim 1 wherein the data representing a document is input from a database in communication with the at least one processor.
11. A computer-implemented method for analyzing a document, comprising: receiving, at one or more processors, data representing a document; identifying, at the one or more processors, different elements in the document using the received data; applying an artificial intelligence process to the identified different elements to generate a textual representation of at least one of the identified different elements; generating, at the one or more processors, a video including audio data based on the generated textual representation, wherein the video is generated with audio data in a selected language; creating, at the one or more processors, a text file representing a summarization of the elements in the document; and creating, at the one or more processors, a video file representing a summarization of the elements in the document.
12. The computer-implemented method of claim 11 wherein the generated textual representation comprises a description of the at least one identified element converted tolayman’s terms.
13. The computer-implemented method of claim 12 wherein the generated textual representation comprises a hypothetical example relating to the at least one identified element.
14. The computer-implemented method of claim 12 further comprising displaying on a screen the created text file and video file.
15. The computer-implemented method of claim 11 wherein the data representing a document received at the one or more processors is associated to a document input by a user or a document input from a database in communication with the one or more processors.
Citation Information
Patent Citations
Computing device and corresponding method for generating data representing text
US20160004681A1
Techniques for generating media content for storyboards
US20210304468A1
Collaborative learning of question generation and question answering
US20220067486A1
Systems and methods for machine content generation
US20230252224A1