Controlled summarization and structuring of unstructured documents

A question-answer framework processes unstructured documents using a trained language model to generate structured outputs tailored to user interests, addressing inefficiencies in summarizing and structuring unstructured data by classifying and clustering content.

JP2025526284APending Publication Date: 2025-08-13PRYON INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025501302
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-11
Filing Date
2023-07-07
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

Users face significant challenges in efficiently summarizing and structuring vast amounts of unstructured data to meet their specific needs, particularly when the data lacks metadata or organizational information, leading to inefficiencies in compiling reports and missing relevant information.

Method used

A question-answer framework using a set of predetermined questions applied through a trained language model processes unstructured documents to generate structured output, such as summaries or reports, tailored to the user's interests by classifying and clustering document content, and generating relevant insights.

Benefits of technology

Automatically identifies and processes unstructured documents to generate customized outputs, reducing user intervention and ensuring that reports and insights are relevant and organized according to the user's needs, even when the documents lack metadata.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025526284000001_ABST
    Figure 2025526284000001_ABST
Patent Text Reader

Abstract

Methods, systems, devices, apparatus, media and other implementations are disclosed, including methods that include obtaining a query set (e.g., a universal set of questions); conducting a question-answer (QA) search on one or more documents using the query set to generate answer data responsive to one or more questions included in the query set, the answer data characterizing concepts related to the one or more documents; and deriving structured output information for the one or more documents based on the answer data generated in response to conducting the QA search.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is an international application claiming priority to U.S. Provisional Application No. 63 / 388,012, entitled "Supervised Summarization and Structuring of Unstructured Documents," filed July 11, 2022, the entire contents of which are incorporated herein by reference.

[0002] The present invention relates to summarizing and structuring unstructured documents. [Background technology]

[0003] Computer users often have access to vast amounts of data, whether accessible through public networks (such as the Internet) or private networks, that the user can search to find answers and information to specific or general queries about a topic or issue. It is often up to the user to search for the specific data they need and then compile the resulting output data into a meaningful output document or report. For example, a financial analyst may need to compile daily financial reports with a level of detail that may depend on the amount of information available to the user from the input data and the previously determined answers to some initial queries. If the user needs to periodically update the reports or generate new or follow-up reports of new events, the effort to search for and compile the information can be enormous. Making matters worse is the fact that most input sources of data may be unstructured (e.g., lacking metadata, summary data, or any type of organizational information), resulting in the user not even realizing that new data relevant to their tasks and responsibilities is available. Summary of the Invention [Means for solving the problem]

[0004] In a broad aspect, an approach to summarizing and structuring unstructured documents involves applying a question-answer process to the unstructured documents using a set of questions, a "query set," that is applied to the unstructured documents to generate answer data in response to the questions (preferably encompassing questions pertaining to multiple concepts or subject areas). This answer data characterizes concepts associated with the documents, and these concepts are used for further processing of the documents. For example, document classification, retrieval, and downstream processing can be based on these concepts.

[0005] The present disclosure is directed to guided intelligent document processing through automated question answering. Intelligent document processing involves the generation of structured representations of unstructured documents. This representation may be some type of report, such as a summary, table, alert, trend analysis, or other informative or actionable insight. The information to be included in a results report (e.g., in a financial report where a user may want to know the total value of their portfolio on a daily basis) is often known a priori. In other cases, the content of the report depends on the user's personal interests. For example, someone interested in sports may want to have a news summary including recent scores, while someone interested in movies may want to know if new movies have been released. The framework described and proposed herein guides the generation of reports (or the generation of other types of output by one or more downstream processes in communication with the question-answering system) for any source document containing content important to the user, where the reports are determined based on answer data generated through the use of target questions submitted to a question-answering system that processes any source document.

[0006] Under the proposed framework, a question-answering system trained based on one or more language models (e.g., a Bidirectional Encoder Representations from Transformers (BERT) language model, a GPT3 language model, or any other type of language model transformation) is used to process unstructured documents with unknown content (whether original source documents or already-transformed, searchable documents) by applying a set of predetermined questions (defining a universe of questions) to the documents. For example, the set of predetermined questions may include a list of questions related to many different concepts or subject areas. The question-answering system returns answer data to submitted questions that indicate (explicitly or inferentially) the relevance of the document's content to the question asked. For example, if the answers returned for a particular pre-defined question (from a library of pre-defined questions) are associated with a low consistency (relevance) score, this score indicates that the document to which the pre-defined questions were applied likely contains content unrelated to the particular question submitted. It can consequently be inferred that the document being processed has low relevance to concepts or subject matter related to the particular question. On the other hand, answers with a high consistency (relevance) score or a high level of detail may indicate that the content is relevant to the question asked, and consequently, the subject or concept of the document's content can be classified / determined. Classification of the document's content may trigger downstream processes (e.g., report generation processes, metadata generation processes) to generate resulting structured output (e.g., reports with output data organized according to some relevant format). For example, a particular document (e.g., an SEC report, a newspaper business article, etc.) discussing the financial performance of a particular company may be identified as a financial report document based on the answer data responses to finance-related questions posed through the question-answering system of the proposed framework, and may trigger a downstream financial report summarization process.The downstream financial report summarization process analyzes the particular document and (in response to initial questions and predetermined follow-up questions posed in response to the initial classification of the particular document as a financial document) generates a report that organizes the data in a specific format related to the document (e.g., the name of the company on line 1, the nature of the report on line 2, an arbitrary monetary value (profit, loss, etc.) on line 3), etc.

[0007] The framework described herein thus enables a user to guide the results of automated processing of a set of documents toward key insights of general and / or personal interest to the user. In the proposed solution, a set of a priori key questions is constructed (e.g., according to specific downstream processes that may be evoked based on an initial classification of the content of the documents being analyzed) so that relevant output data is targeted to the content for which it is generated. Questions within the a priori question set may be personalized toward the user's specific interests. These questions are sent to an automated question-answering system to generate appropriate output data (reports, classifications, alerts, etc.).

[0008] Advantageously, the proposed techniques and solutions described herein can automatically identify and perform applicable processing of any unstructured document with little or no guidance or intervention from the user, and generate customized / specialized output (which may take into account the special needs or desires of a particular user). Thus, upon receipt of any document, automatic classification can be performed through a question-and-answer process, and specialized report creation and output generation functions can automatically generate the required summary, report, or other type of output.

[0009] Accordingly, in some aspects, a method is provided that includes obtaining a query set; conducting a question-answer (QA) search on one or more documents using the query set to generate answer data responsive to one or more questions included in the query set, the answer data characterizing concepts related to the one or more documents; and deriving structured output information for the one or more documents based on the answer data generated in response to conducting the QA search.

[0010] Some embodiments of the method include one or more of the following features and may include at least some of the features described in this disclosure.

[0011] Deriving structured output information for one or more documents may include, for example, one or more of the following: determining classification information for one or more documents that represent at least one of the concepts; performing data clustering for one or more documents based on the answer data; applying a data discovery process to the answer data to determine one or more labels related to one or more concepts associated with the one or more documents; generating an output report based on the answer data; and / or deriving supplementary data related to at least some of the answer data.

[0012] Deriving supplemental data related to at least some of the answer data may include determining supplemental concepts related to at least some of the answer data, e.g., accessing at least one of one or more documents and / or other data sources, and determining supplemental information related to the supplemental concepts from the accessed at least one of the one or more documents or other data sources.

[0013] Determining supplemental concepts may include determining and applying supplemental questions to at least one of the one or more documents or another data source.

[0014] Generating the output report may include, for example, one or more of the following: generating a summary report to be provided to the user based on at least some of the response data arranged in one or more predetermined templates, generating an alert to be communicated to the user, and / or entering at least some of the response data into a database table.

[0015] Generating the output report may include determining a score for the answer data generated in response to performing the question-answer search using the query set, and including a predetermined number N1 of the highest-scoring answers determined from the answer data in the output report.

[0016] The method may further include identifying additional answers from the resultant answer data whose respective scores exceed a predetermined score threshold, and selecting up to N2-N1 selected answers (where N2>N1) from the additional answers whose respective scores exceed the predetermined score threshold for inclusion in an output report.

[0017] Generating the structured output information may include generating the structured output information based on the response data and further based on user information associated with the user.

[0018] The user information may include, for example, one or more of the user's personal preferences, network access controls associated with the user, and / or network groups with which the user is associated.

[0019] The method may further include determining an additional query based on at least some of the answer data, and conducting an additional question-answer search of one or more documents using the additional query.

[0020] Determining the follow-up queries may include using one or more ontologies that define relationships and associations between the concepts identified from at least some of the response data and other distinct concepts, and deriving follow-up questions for the follow-up queries based on the other distinct concepts determined using the one or more ontologies.

[0021] The method may further include determining a score for the answer data generated in response to conducting the question-answer search using the query set. Generating structured output information for the one or more documents may include generating structured output information for the one or more documents based on the determined score for the answer data.

[0022] The query set may include a universe of questions related to a plurality of different content subject areas. Generating the structured output information may include determining that one or more documents are unrelated to one or more of the plurality of different content subject areas based on determined scores for answer data generated from the plurality of questions related to one or more of the plurality of different content subject areas.

[0023] Determining the score of the answer data may include calculating a score for a particular answer in response to a particular question of one or more questions in the query set that represents, for example, one or more of the following: the similarity of the particular answer to the particular question; the similarity of the combination of the particular question and the particular answer to pre-defined question-answer pairs for one or more documents; the similarity of the particular answer to previously selected answers provided to a particular user; the relative location of the particular answer in one or more documents; and / or the level of detail contained in the particular answer.

[0024] Generating the structured output information may include applying one or more machine learning models to at least some of the response data.

[0025] Obtaining the set of queries may include tailoring the set of pre-defined questions based on user information associated with the user.

[0026] The user information may include, for example, one or more of the user's personal preferences, network access controls associated with the user, and / or network groups with which the user is associated.

[0027] The method may further include receiving one or more source documents and converting the one or more source documents into one or more documents on which the QA search is performed.

[0028] Transforming the one or more source documents may include applying one or more pre-segmentation processes to the one or more source documents to transform the one or more source documents; and applying one or more vector transforms to the one or more segmented documents to transform the one or more segmented documents into respective vector answers in one or more vector spaces.

[0029] Applying the one or more vector transforms may include transforming the one or more segmented documents according to, for example, one or more of the following: a Bidirectional Encoder Representations from Transformers (BERT) language model, a GPT3 language model, a T5 language model, a BART language model, a RAG language model, a UniLM language model, a Megatron language model, a RoBERTa language model, an ELECTRA language model, an XLNet language model, and / or an Albert language model.

[0030] Deriving the structured output information may be further based on interaction data provided by a user.

[0031] The interactive data may be provided in response to prompt data generated by the QA system and may include disambiguation data for selecting an answer from among multiple matches in the answer data related to one or more similar concepts.

[0032] In some aspects, a system is provided that includes one or more memory storage devices for storing executable computer instructions and data, and a processor-based controller electronically coupled to the one or more memory storage devices, where the controller is configured to: obtain a query set; conduct a question-answer (QA) search on one or more documents using the query set to generate answer data responsive to one or more questions included in the query set, the answer data characterizing concepts related to the one or more documents; and derive structured output information for the one or more documents based on the answer data generated in response to conducting the QA search.

[0033] In some aspects, a non-transitory computer-readable medium is provided that is programmed with instructions executable on one or more processors of a computing system for: obtaining a query set; conducting a question-answer (QA) search on one or more documents using the query set to generate answer data responsive to one or more questions included in the query set, the answer data characterizing concepts related to the one or more documents; and deriving structured output information for the one or more documents based on the answer data generated in response to conducting the QA search.

[0034] Embodiments of any of the above systems and / or computer-readable media may include at least some of the features described in this disclosure, including the above features of the present methods, and may be combined with any other embodiments or variations of the methods, systems, media, and other implementations described herein.

[0035] Other features and advantages of the invention will become apparent from the following detailed description and claims.

[0036] These and other aspects will now be described in detail with reference to the following drawings. [Brief explanation of the drawings]

[0037] [Figure 1] 1 is a diagram of an exemplary system for unsupervised question-answer document processing.

[0038] [Figure 2] FIG. 2 is a diagram of another exemplary system for unsupervised question-answer document processing.

[0039] [Figure 3] FIG. 1 is a schematic diagram of a document processing framework.

[0040] [Figure 4] 1 is an exemplary diagram of a document capture procedure.

[0041] [Figure 5] 1 is a flowchart of a procedure for guided intelligent document processing through automated question answering.

[0042] Like reference numbers in different drawings indicate like elements. DETAILED DESCRIPTION OF THE INVENTION

[0043] Disclosed is an implementation of a document processing system that can automatically process any unstructured document (e.g., process the document without necessarily having any a priori information about its content) to generate structured output (e.g., customized reports based on predetermined templates or scripts, metadata, alerts, etc.). The generation of structured output is achieved, at least in part, by applying a set of questions that suitably cover a range of topics, concepts, and subject areas (e.g., finance, business, sports, defense, and security, national and international news covering different news categories, etc.) to any unstructured document (which may initially be processed by a QA system that performs an ingestion operation to convert the document into a searchable document, as described in more detail below). The set (or library) of questions can be supplemented or customized based on a specific user's identifier, allowing for automated iterations of the initial QA on behalf of the specific user, thereby taking into account previously determined areas of interest or specific information needs associated with the user.

[0044] Submission of a set of questions covering a range of different topics, concepts, and subject areas results in answer data that can be processed by downstream processing to generate structured data. For example, the answer data can be used to perform one or more of the following: i) a classification process (determining the nature of the document and what content is contained within the document (e.g., is a particular document a financial statement? Is a particular document a medical record? Is a particular document a legal document such as an NDA or contract?), ii) a data clustering process, iii) a data discovery process to determine one or more labels associated with concepts related to the document, iv) generating an output report (which can be customized according to predetermined and optionally adjustable templates), and / or v) deriving supplemental data related to at least some of the answer data (e.g., performing multi-hop concept discovery in which additional data (not included in the content of any document) is accessed from other sources to provide a recipient user with information that they would not otherwise have obtained if the original document had been available to the user). Other types of downstream processing for generating other types of structured output can also be implemented.

[0045] Thus, in a broad example approach, a method for facilitating structuring of unstructured documents is provided, including obtaining a query set (e.g., a library of questions defining a universe of questions spanning a range of topics and concepts) and using the query set to conduct a question-answer (QA) search on one or more documents (which may be ingested to convert the one or more documents into one or more respective QA-searchable documents) to generate answer data responsive to one or more questions included in the query set. The answer data generated through performance of the QA search using the predetermined query set characterizes concepts associated with the one or more documents (e.g., indicates the concepts, topics, subject matter, general nature, and other characteristics of the one or more documents). The one or more documents upon which this global QA search is performed are generally unstructured documents, which may lack any a priori information regarding the nature of the one or more documents or their content. The method further includes deriving / generating structured output information for the one or more documents based on the answer data generated in response to conducting the question-answer search. In some embodiments, deriving structured output information (through application of one or more downstream processes) may include, for example, one or more of the following: determining classification information for one or more documents that represent at least one of the concepts; performing data clustering for one or more documents based on the answer data; applying a data discovery process to the answer data to determine one or more labels associated with concepts associated with one or more documents; generating an output report based on the answer data; and / or deriving supplementary data associated with at least some of the answer data.

[0046] Structured output information (which may be classification information, metadata, output reports, alerts, control messages for updated databases, etc.) may be determined according to the relevance / consistency scores calculated for the answers returned to the questions. Such relevance scores may be calculated based on, for example, a distance measure between the semantic content of the answers and the corresponding questions, the output of a trained machine learning engine for assessing relevance, etc. For example, one type of scoring process may be based on the transform-based distance (TBD) between the questions and answers, or the equivalent of the posterior probability of the TBD. A particular question and answer pair with a high relevance score may indicate that a particular document from which the question-answer pair was generated relates to a concept or topic related to the particular question. This may trigger a specific downstream process for generating structured output for the identified concept / topic. Therefore, in such an embodiment, the proposed framework may also be configured to determine a score for the answer data generated in response to performing a question-answer search by using a query set, and to generate structured output information for one or more documents based on the determined score of the answer data. As noted, the query set may include a universe of questions related to multiple different content subject areas, and generating the structured output information may include determining that one or more documents are unrelated to one or more of the multiple different content subject areas. Such a determination may be made based on determined scores of answer data generated in association with questions related to one or more of the multiple different content subject areas.

[0047] Further details regarding the proposed framework will now be provided with reference to FIG. 1. FIG. 1 shows a schematic diagram of an exemplary system 100 that processes unstructured data and generates structured output by using a question-answering system (schematically depicted in FIG. 1 as QA system 120). While some of the operations, functions, and / or processes discussed in this disclosure are described as applying to a particular document (e.g., a single document), it should be noted that the operations may be performed simultaneously or sequentially on a collection of documents. As shown, a document 102 is obtained. This document may be a news article, a personal report, or any type of document, and in the general case, may be received as an unstructured document with unknown content and no relevant information (e.g., metadata) regarding its nature or type. It should be noted that in some situations, at least some information may be known about the document (e.g., the network location from which the document was transmitted, authorship information, etc.). However, for purposes of illustrating the embodiments described herein, it is assumed that at least some characteristics and attributes of the document 102 or its content are not known a priori, and thus the document 102 is depicted as unstructured.

[0048] Preferably, before structured information is extracted from the raw content of unstructured source documents 102, the documents 102 are typically preprocessed by a preprocessing unit 110 that generates result documents 112 to which question-answering processes can be applied via a QA system 120. The preprocessing unit 110 may be part of a document processing platform that includes the QA system 120 and a communication interface (enabling interaction between various users and administrators and the document processing platform), as discussed in more detail below in connection with FIG. 3. The preprocessing unit 110 may be configured to perform ingest processing operations, such as data authentication and / or decryption, format conversion (if necessary), transformation operations, including document segmentation and vectorization (parameterization) operations, and storage operations that store the ingested documents in a repository of the system 100. For example, the segmentation operation may include dividing the document into portions (e.g., 200-word portions or any other word-based segments), where the segmentation is performed according to various rules that combine content from various portions of the document into distributed segments. An example of a preprocessing (i.e., pretransformation) rule is to construct a segment using a fixed-length or variable-length sliding window that combines one or more headings that precede the content to be captured, thereby generating a contextual association between the heading or headings and the captured content. Such a rule ensures that subsequent language model transformations performed on the segment combine important contextual information with content that is located far from the segment being processed (e.g., far away in the source document).

[0049] Another preprocessing step that may be optionally applied during segmentation of the source document 102 involves processing tabular information (e.g., when the original content is arranged in a table or grid). Such preprocessing is used to expand structured data arranged in a table (or other type of data structure) into a searchable format, such as a text equivalent. For example, if a portion of the source document 102 is identified as being a multi-cell table, a substitute portion is generated to replace the multi-cell table. Each of the multiple substitute portions includes content data as a subportion associated with the multi-cell table, and contextual information.

[0050] Following segmentation-related preprocessing operations, the resulting segmented preprocessed document can be submitted to one or more transformers. An initial language model transformation process is configured to reformat the document to include each sentence as a potential answer to the question. One example of a language model transformation that can be applied to the input content (which, as described above, may first be preprocessed for decoding and authentication and for segmenting the content into manageable chunks) is the Bidirectional Encoder Representations from Transformers (BERT) transformation. Briefly, under the BERT approach, the question and answer are concatenated (e.g., tokenized using WordPiece embedding, with suitable markers separating the question and the answer) and processed together in a self-attention-based network. The output of the network indicates a score for each possible starting position of the answer and a score for each possible ending position of the answer, and the overall score for the span of the answer is the sum of the corresponding starting and ending positions of the answer. That is, self-attention methods are used where the embedding vectors of a paragraph and the embedding vectors of a query are mixed through many layers prior to the decision-making and segmentation logic, providing an efficient way to determine whether a question can be answered by a paragraph, and if so, where exactly the answer span is within the paragraph.

[0051] In BERT-based methods, a network is first trained on a masked language model task in which words are omitted from the input, and predictions can be made by the network with an output layer that provides a probability distribution over the words in the vocabulary. Once the network has been trained on the masked language model task, the output layer is removed, and for question answering tasks, layers are added to produce start, end, and confidence outputs. The network is further trained (e.g., fine-tuned transfer learning) on supervised training data for the target domain (e.g., by using the Stanford Question Answering Dataset (SQuAD)). Once the network has been trained for question answering in the target domain, further training can be used to adapt the network to new domains. Another training strategy used for BERT is next-sentence prediction, in which a learning engine is trained to determine which of two input segments (e.g., such segments can be neighboring sentences in a text source) is the first of the two segments. When training the model, both the masked language training and next-sentence training procedures can be combined by using an optimization procedure that attempts to minimize a combined loss function. Alternatively or additionally, other training strategies (to achieve context awareness / understanding) may be used separately or in conjunction with one of the aforementioned training strategies for BERT.

[0052] An exemplary embodiment based on the BERT technique may use an implementation known as the Two-Leg BERT technique, in which much of the query processing is separated from the processing of portions of documents (e.g., paragraphs) where an answer to the query may be found. Generally, in the two-leg BERT technique, a neural network architecture has two "legs" (one leg for processing the query and one leg for processing the paragraph), and the outputs of the two legs are sequences of embeddings / encodings of the query words and the paragraph words. These sequences are passed to a question-answering network. A particular way this technique is used is to pre-compute BERT embedding sequences for paragraphs and complete the question-answering computation once the query is available. Advantageously, because much of the processing of the paragraphs is done before the query is received, the answer to the query can be computed with less delay compared to using a network in which the query and each paragraph are concatenated and processed together in sequence. Paragraphs are typically much longer than queries (e.g., 200-300 words vs. 6-10 words), so pre-processing is particularly effective. When sequential queries are applied to the same paragraph, the overall amount of computation can be reduced because the output of the paragraph's legs can be reused for each query. Low latency and reduced overall computation can also be advantageous in server-based solutions. BERT-based processing of a source document typically generates transformed content that is stored in a repository (such as DOM repository 340 in FIG. 3). The underlying documents from which the BERT-based transformed content is generated can likewise be retained and associated with the resulting transformed content (as well as with corresponding transformed content obtained via other transformations).

[0053] In some embodiments, a BERT-based transformer (e.g., used for fast / coarse transformation and / or for fine-detail transformation) may be implemented according to an encoder-based architecture. For example, the structure of a BERT-based transformer may include multiple stacked encoder cells, where an input encoder cell receives and processes an entire input sequence (e.g., a sentence). By processing the entire input sentence, a BERT-based implementation may process and learn contextual relationships between individual parts (e.g., words in the input sequence). The encoder layers may be implemented with one or more self-attention heads (e.g., configured to determine relationships between different parts of the input data, e.g., words in a sentence) preceding a feedforward network. The outputs of different layers in the encoder implementation may be directed to a normalization layer to appropriately configure the resulting output for further processing by subsequent layers.

[0054] In some embodiments, other language models may be used to transform the source document (in addition to or instead of the BERT-based transformation) as part of the preprocessing operations performed by preprocessor 110 of Figure 1. Examples of such additional language model transformations include: Autoregressive language model based on GPT3 language model-transformer implementation. T5 Language Model - A text-to-text converter-based framework that predicts output text from input text fed into the model. BART Language Model - a denoising autoencoder that uses a standard transformer architecture implemented with Gaussian Error Linear Units (GeLUs) as its activation function elements, is trained by corrupting text with an arbitrary noising function, and uses the learned model to reconstruct the original text. RAG Language Model - The Retrieval-Augmented Generation (RAG) language model combines pre-trained parametric and non-parametric memories for language generation. UniLM Language Model - A unified language model framework that is pre-trained using three types of language modeling tasks: one-way, two-way, and sequence-to-sequence prediction. Unified modeling is achieved by employing a shared transformer network and utilizing specific self-attention masks that control how the predictions are conditioned on the context. Megatron Language Model - a multi-parameter transducer framework. RoBERTa language model - a model similar to the BERT method, but with some of the BERT hyperparameters optimized. ELECTRA Language Model - This method uses a generator unit to replace (and therefore corrupt) tokens in the input data with plausible alternatives. The discriminator unit then attempts to predict which tokens in the modified (corrupted) input have been replaced by the generator unit. XLNet Language Model - An autoregressive pre-trained framework that uses context related to text to predict nearby text. The XLNet framework is constructed to maximize the expected log likelihood of a sequence over all possible permutations of the factorization order. Albert Language Model - This model is similar to the BERT model but has a smaller parameter size.

[0055] Other different language models implementing different prediction and training strategies may similarly be used in implementations of the proposed framework of FIG. 1.

[0056] The preprocessed source document (e.g., subjected to different secure communication processes (including authentication and decryption) prior to segmentation and language model transformation) results in the searchable document 112 depicted in FIG. 1 . Typically, the resultant document 112 is stored in a repository of preprocessed documents. For example, the preprocessed document generated by the preprocessing unit 110 is stored as a document object model (DOM) object in a repository configured to store and manage DOM records. The content of a DOM record typically depends on the transformations performed by the preprocessing unit 110. A DOM record may contain data items related to a particular source document or a particular portion of a source document. For example, a DOM record may be a collection of items including the original portion of the source document, metadata for the portion of the source document, contextual information related to the portion of the source document, and / or vectors (also called embeddings) resulting from one or more transformations applied to segments of the source content. A particular segment may be transformed by several different transformers (corresponding to different language models and / or different levels of resolution or coarseness of the segment being transformed) and therefore may be associated with several different vectors. Metadata associated with the transformed content may include contextual information related to the original source content and location information indicating the location or position of a portion within a larger source document. Such location information may be provided in the form of pointer information that points to a memory location (or memory offset location) where the source document is stored (e.g., so that when the pointer information is returned to a requesting user, it can be used to indicate a memory location where relevant content that constitutes an answer to the user's query can be found).

[0057] Continuing with reference to FIG. 1 , the system 100 includes a database of questions (also referred to as a query set) 104, the application of which to the documents 112 via the QA system 120 generates answer data. This answer data can then be utilized by the post-processing unit 130 to perform one or more downstream post-processing tasks (such as generating specific reports, generating alerts, performing data discovery tasks based on the answer data generated through the application of the query set to the preprocessed documents 112, etc.). As noted, an a priori set of questions (defining a library or universe of questions) is constructed to identify target content (identifying structured information about unstructured documents that would otherwise have unknown or poorly defined content) upon which post-processing is performed depending on the identified type and nature of the documents. In other words, the QA system is utilized to investigate / probe the nature of the documents 112. The query set 104 typically includes a set of questions (that can be applied to the preprocessed documents 112) that potentially cover a large range of topics, concepts, and subject areas. For example, a query set may include questions directed at legal concepts, business concepts, concepts trending on social media (sports, politics, and other current concepts), specialized subject areas related to science, engineering, and many other areas of knowledge.

[0058] The query set 104 is adjusted intermittently (periodically or irregularly) to update the query set according to changing characteristics of various popular concepts (as may be determined according to social media trends). In some embodiments, the question set may be personalized or customized either before or after the initial application of the query set 104 according to an identifier of the user overseeing the processing and / or other contextual information associated with the source document 102 (including where the source document 102 was originally stored, where the query request arrives from (e.g., geographic or network address), an identifier of the entity for which the structured discovery processing is performed, network access control information (e.g., network access permissions) associated with a particular user or a larger group of users, etc.). For example, if a source document or a request to perform a processing operation as described herein arrives from a legal services entity (e.g., a law firm), the query set may be dynamically adjusted to include additional legal questions (e.g., questions regarding non-disclosure agreements, leases, asset transfer agreements, etc.).

[0059] As described in more detail below, in some embodiments, the query set may be iteratively adjusted according to the resulting answer data generated from the application of an initial (or prior) set of questions within the query set. For example, in response to the application of the query set, answer data is generated that may indicate the level of responsiveness / relevance of the contents of the source documents 102 or the preprocessed (ingested) resulting documents 112 to different questions within the initial query set (e.g., through relevance scores associated with text-based representations or parameterized representations of answers to the applied initial set of questions). Thus, the answer data may be used to identify relevant concepts or subject areas to which the documents are more likely to be related, and to eliminate from further consideration concepts, topics, and subject areas that yield answers with relatively weak relevance / consistency scores. Based on the answers deemed more relevant to the questions within the query set, supplemental queries / questions may be determined that can be used to conduct subsequent question-answer searches of the documents 102 or 112.

[0060] For example, the framework described herein may generate follow-up questions as part of a question expansion procedure. The QA system 120 of FIG. 1 may access ontology and synonym datasets to construct new questions that are semantically similar to questions deemed appropriate for the concepts / topics of the document, or that are determined to be appropriate follow-up questions given the structured information discovered about the document 102 or 112 and / or contextual information related to the user and entity to which the initial research question set was submitted. The determination of follow-up questions may be performed, for example, via a rule-based process or a machine learning process (e.g., generating labels representing follow-up or follow-up questions for previously asked questions). The new questions are processed by the QA system 120 to generate additional answer data, and this process may be repeated again (e.g., new question generation may continue for a fixed number of iterations or until some condition is met, such as reaching a level of answer responsiveness associated with some specified confidence level).

[0061] Thus, the proposed framework is configured to use an initial query set to query documents such as document 112 (e.g., by QA system 120) and generate answer data in response to one or more questions included in the query set. The initial query set may be organized as a database of questions, such as customized query set 104, based on contextual data associated with the user submitting the query set. As noted, in some embodiments, the proposed framework may be further configured to determine a follow-up query (in the form of a follow-up question) based on at least some of the answer data and to perform additional question-answer searches on document 102 or 112 (and / or the additional document) by using the follow-up query (e.g., by submitting the follow-up query to the QA system). To determine the follow-up query, the proposed framework may be configured to determine the follow-up query by using one or more ontologies that define relationships and associations between concepts identified from at least some of the answer data and other distinct concepts, and to derive follow-up questions for the follow-up query based on the determined other distinct concepts by using the one or more ontologies.

[0062] As discussed herein, the determined query set (and any subsequent supplemental query sets) submitted to the QA system are processed and applied to the preprocessed documents 112, and the determined query set is used to derive structured output information based on answer data resulting from performing QA searches on the documents 112. More specifically, as further discussed in connection with Figures 2 and 3 below, the QA system 120 is configured to transform the query data of the query set (i.e., multiple questions covering a range of subject areas, concepts, and topics) into transformed query data that is compatible with the content of the transformed source (e.g., compatible with one or more of the transformed content records in the DOM repository). For example, if the preprocessed documents 112 include a data representation compatible with BERT-based transformed data, the query set 104 is similarly processed to transform it into a BERT-based representation (e.g., to generate a parameterized / vector representation of the question that includes the query set). The QA system 120 launches a search process, in which the QA system 120 identifies one or more candidate portions within the transformed content of the document 112 that match the transformed query data according to one or more matching criteria. For example, the matching operation may be based on some closeness or similarity criterion corresponding to some calculated distance between, for example, a calculated vector derived from the transformation applied to the query data and the vector record comprising the document 112. In some embodiments, a matching / relevance score is derived to represent the relevance of each question in the query set to the content of the document 112 (and thus to the content of the original source document 102).

[0063] The matching / searching process applied to the query set and documents 112 generates answer data (which may optionally include a relevance score) that indicates the relevance of each question. In the exemplary embodiment of FIG. 1 , the answer data and correspondence scores are represented as data 122. The answer data 122 (whether provided in a user-readable semantic format or as a vector / parameter representation) represents structured information related to documents 112 or 102, as the matching process identifies predefined questions from the query set 104 related to a given concept or subject area for which an answer with some relevance is found in documents 112. For those predefined questions in the query set that did not produce a possible matching answer or that produced an answer with a low matching / relevance score, it is assumed that the question is not relevant to documents 112, and therefore documents 112 (or 102) are not considered to be related to the concept or subject area associated with those non-matching or low-relevance questions.

[0064] As additionally shown in FIG. 1 , upon generating answer data 122 (with or without respective scores), structured output information is derived for document 112 (by extension source document 102) by applying one or more post-processing (i.e., post-QA operations) processes (executed on post-processing unit 130) to answer data 122. In some embodiments, specific downstream post-QA processes may be automatically invoked or launched in response to answer data indicating that the content of document 112 (or 102) is of a type (e.g., related to a particular subject) for which a specific structured output generating process is to be performed. For example, consider a situation in which a QA process applied to document 112 determines that the content of document 112 (and thus source document 102) contains financial performance data for a particular company. Suppose that, upon recognizing that document 112 contains financial data, post-processing unit 130 performs a data extraction process to generate a financial report summary to locate the portion of answer data (generated by QA system 120 in response to overall query set 104) that pertains to the financial data. For example, the query set 104 may include questions such as: · What is the name of the company? · What industry does the company operate in? · What was the daily change in the company's stock price? · What caused this change? · What is the name of the company's CEO? · How many employees does the company have? · Who are the company's largest shareholders? · and so on.

[0065] If posing those questions yields meaningful / relevant answers, post-processing unit 130 may determine that the analyzed document is a financial news item or report and consequently launch a process to prepare a financial report (e.g., following some predetermined template) to be provided to a user (such as a stockbroker or financial advisor) interested in the financial data contained in document 102 or 112. In some embodiments, the report / summary (depicted in FIG. 1 as summary document 132) may include fields requesting information that may not have been generated through application of QA system 120 using the initial query set (e.g., because the query set did not include a specific question corresponding to the missing report field). In this case, post-processing unit 130 (which implements the specific financial reporting process) may interact with the QA system to pose additional questions (searched for either within document 112 or within other documents (maintained locally or remotely in one or more data repositories)) corresponding to the unknown information. Upon receiving such a follow-up question, the QA system may convert the question into a representation compatible with the data representation (e.g., a vector representation and / or a semantic representation) of document 112 and perform a search on document 112 and / or other documents that may be accessed (locally or remotely) via QA system 120 to determine the missing information. In the financial reporting example, consider a situation in which a report generated by post-processing unit 130 includes a field for CEO age. In this example, a question regarding the CEO's age is not included in the initial query set 104, and therefore the financial reporting process may be configured to send a query request to QA system 120 (or to some other search platform with which post-processing unit 130 may communicate) to determine the age of the CEO identified for the company discussed in document 112 (or document 102).

[0066] In addition to organizing reports or summaries according to concept and subject area identifiers through QA processing of one or more unstructured documents using the entire list of questions, many other types of downstream processing can be performed following the QA processing (e.g., on post-processing unit 130 or at some other local or remote computing node). Some illustrative, non-exhaustive examples of such downstream post-QA processing include document classification (e.g., used as a trigger for the report organizing process), data clustering processing, data discovery processing, generating alerts, generating database access requests to add data (determined from the resulting answer data) to a database, generating specific documents based on the answer data using information determined from the answer data if the QA processing indicates that such documents are needed (e.g., preparing various legal documents such as non-disclosure agreements, contract clauses, deeds, etc.). Various downstream processing can be performed using machine learning models (e.g., for classification and data clustering processes), rule-based processes, or any type of algorithmic process for processing or further analyzing the answer data generated from the QA processing performed on unstructured documents 112 (or documents 102).

[0067] More specifically, referring to FIG. 2 , a schematic diagram of another exemplary system 200 for processing unstructured documents using a QA system is shown. System 200 is similar in its general configuration and operation to system 100 of FIG. 1 . However, FIG. 2 provides more detail regarding downstream post-processing implementations and depicts several different processes that may be performed in response to answer data generated by the QA system. Accordingly, system 200 includes a repository 202 of documents, which may include raw (unprocessed) documents and at least some preprocessed documents (e.g., source documents that have undergone secure communication processing, initial formatting processing, and / or conversion to a language model transformation representation, such as BERT). The preprocessed documents may have been processed using a preprocessing module, such as preprocessing module 110 of FIG. 1 . One or more of the documents in repository 202 are then provided to a question-answering (QA) system 220, similar to question-answering system 120 of FIG. 1 . The QA system may be part of a document processing framework, such as the framework described in connection with FIG. 3 . As described in more detail below, such a framework may be implemented privately on a local network operated by an individual entity (e.g., a company) and protected by a firewall from unauthorized remote access attempts, or alternatively, may be implemented as a cloud server serving multiple entities that perform at least some of their document processing operations on the cloud server.

[0068] Document D, whether operated locally or on a cloud server x A query set containing multiple questions covering multiple concepts, topics and subject areas is documented as D x For illustrative purposes, a single document D x Refer to D xIt is understood that the query set maintained and / or retrieved from library 204 may optionally be customized according to the preferences of a particular user (e.g., user ID, network location, and document D). x may be determined based on the user's preferences (as may be determined based on the user's preferences, or any other contextual information indicative of the user or group of users being analyzed), and thus the set of initial questions may be expanded explicitly or inferentially through the user's preferences (by using optional question expansion process 206). Question expansion may be achieved through algorithmic and / or machine learning processes to identify potential additional questions related to the set of initial questions as informed by the user's preferences.

[0069] Therefore, Document D x The set of questions that apply to is questions Q1-Q N Answer A x 1-A x N. Generally, document D x Without any a priori knowledge of the content of document D, and therefore by using a broad set of questions that span a potentially large range of concepts, topics and subject areas, x It is therefore possible to explore and determine the nature of the content of document D x For example, it is possible to derive the structured information of document D. x Some of the questions applied to the document D x Some questions may result in no answer (e.g., returning a blank answer), indicating that the document is unlikely to have any relevance to the question. Other questions may result in answers that may have varying levels of completeness (e.g., as may be indicated by the calculation of a consistency / relevance score). Generally, the more complete the answer or the higher the relevance score of the answer, the more likely the document is to have subject matter that overlaps with the question that generated the answer / score.

[0070] As further shown in FIG. 2, once the set of answers (which may not include any answers for some of the questions) and / or relevance scores are determined, the answer, document D x , and the questions contained within the query set are typically forwarded for downstream post-processing by performing one or more of the downstream processes 230-236 depicted in Figure 2. While only four processes are shown in Figure 2, any number of downstream processes may be available for execution (such processes may be performed by a post-processing unit similar to post-processing unit 130 of Figure 1, implemented on one or more local or distributed computing devices / nodes). The invocation of a particular downstream process may be in direct response to some subject area identifier as a result of applying the query set 204. For example, x As a result of asking a question intended to ascertain "whether the transaction involves a real estate or business transaction requiring the production of a corresponding document or contract," the QA system 220 x generated response data indicating that document D truly pertains to such transactions (as determined by the answers and / or scores associated with questions designed to identify these types of transactions). x A special contract generation application / process can be implemented to generate the necessary documents based on the content of the contract.

[0071] As shown in FIG. 2 and also discussed in connection with FIG. 1, some non-exhaustive examples of post-QA processes that may be launched in response to generated answer data in response to questions asked include the following examples: Report generation process 230 (also identified as process 1) is configured to generate a report or summary (which may be similar to report 132 of FIG. 1) according to an optionally customized template or format. In some embodiments, the query set is converted into documents D xThe answer data (in semantic, vector, or some other representation) generated from applying the process to the answer data is arranged in a summary report according to scores associated with the answers that contain the answer data (identified as processes) and according to various rules for selecting and presenting the answer data. For example, after asking each of the predefined questions (and / or any extended or custom questions) that resulted in answer data (such as actual semantic content, vector representations, location information identifying paragraphs of DOM objects or corresponding paragraphs in remotely stored source documents, etc.) and / or obtaining their respective scores, the best-scoring answers to the predefined questions asked are output to the report. More specifically, if the summary to be generated needs to include N1 to N2 sentences, then the top N1 scoring answers are selected as the basis for the summary; if the scores of the (N2 - N1) additional sentences are greater than a certain threshold, then the (N2 - N1) additional sentences may optionally be included in the summary. Thus, in such an embodiment, generating the output report (e.g., by the report generation process 230) may include determining a score for the answer data generated in response to performing the question-answer search by using the query set, and including in the output report a predetermined number N1 of the answers determined from the answer data having the highest scores. The report generation process may further include identifying additional responses from the answer data results whose respective scores exceed a predetermined score threshold, and selecting up to N2-N1 (where N2>N1) selected answers for inclusion in the output report from the additional answers whose respective scores exceed the predetermined score threshold. In this example, document D xThe calculation of the score of an answer determined from may be based on a distance measure, such as the answer's distance to the beginning of the document (the answer's similarity to previously selected answers, etc. In one example, the score used to select an answer for inclusion in a report or summary may be calculated based on a Transform-Based-Distance (TBD) scoring process (or the TBD's posterior probability equivalent) between the question and the answer, and the TBD (or posterior probability equivalent score) is combined with various other heuristics (distance from the beginning of the document, similarity to other answers, etc.). For example, in one embodiment of a summary generation process, individual sentences are clustered within a document into several centroids. For each answer to each question in the question set, a TBD score (or some other metric representing the quality of the match) for all centroids is derived, and the minimum TBD is included in the heuristic combination of "things" (various scores) used to determine the final score. For the heuristic combination process, additional score components may be introduced in some specified order, and the decision / choice regarding how to combine such additional score components may be made at different steps.

[0072] In another exemplary embodiment, the generated report provides a user-friendly and informative format for the document being analyzed (e.g., document D). x ) may be formatted / structured according to a predetermined template that presents the content available in the current events news item. For example, upon determining that the document being analyzed is a current events news item, a summary is generated that reproduces at least some of the questions that yielded answer scores indicating high relevance of the document to the current events item, along with corresponding answers to those questions (presented in a user-readable semantic format). For example, consider the following set of questions that would generate relevant answers if the document were related to a current events news item: · What happened? · What is the most likely cause? · Who is doing this? · What are the consequences of this? · and so on.

[0073] The above example may provide a general analysis of any news item (covering any subject area, such as politics, world affairs, sports, finance, etc.). For entities or users with special interests in more specific subject areas, the general news item question set may be customized to make it more specific to that particular subject area. For example, for a user with an interest in finance, the above question set may be personalized to include: · What are the financial implications of this? · How much does this cost? · Who is paying for this? · and so on.

[0074] Alternatively, as mentioned above, the query set may include more specific questions related to a particular subject (in addition to or instead of general news item discovery questions). Again, not all answers (and / or corresponding questions) necessarily need to be generated in the report, but instead of only the top N1 scoring answers being selected for reporting, additional optional answers (having lower scores) may be included in the report if they meet a minimum relevance score determination criterion.

[0075] In some embodiments, the report generation process 230 may be configured to generate legal or administrative documents based on information extracted through the QA process (e.g., entity names, relationships, agreement terms, etc.). Examples of legal or administrative documents may include real estate documents, business documents (contracts), non-disclosure agreements, etc. Alternatively, the QA system may be configured to recognize certain legal or administrative content and generate a report that summarizes key information that may be gleaned from the analyzed documents (e.g., providing a summary of the transaction including the names of the transaction parties, the nature of the transaction, the key terms of the transaction, etc.).

[0076] As further shown in FIG. 2, any document (document D xAnother example of a downstream process that may be applied to the answer data generated by the QA system 220 through the application of a universe of questions to the QA system 220 (e.g., a dataset of questions), is a data mining process 232 (also identified as process 2) configured to perform data clustering and data store analysis based on the answer data. The data clustering process may be implemented according to one or more clustering techniques that use information extracted by the QA system 220 to perform document-level clustering (e.g., document D x and / or data store level clustering to determine equivalent entity names contained within the document D analyzed by the QA system 220. x (or portions thereof) to identify relationships and / or associations (e.g., according to one or more similarity determination criteria) between the analyzed data / content and other documents stored in the local repository or a remote data repository. Clustering techniques implemented may include algorithmic clustering (e.g., calculating similarities or distances between the analyzed data / content), machine learning clustering (based on machine learning models trained to identify grouping relationships), rule-based clustering techniques, etc.

[0077] The query set 204 is used to organize the QA process into documents D x Another exemplary downstream process that may be invoked in response to answer data resulting from applying the QA process is a database management process 234 configured to update a database (or data repository) with information extracted through the QA process. The high-scoring answers to the various questions contained in the query set 204 may be used (perhaps after performing an early report generation process or data mining / clustering operations) to identify databases / tables that need to be maintained (or created) in response to information contained in newly received unstructured documents.

[0078] Consider an example where the report generation process is used to present questions and answers in the form of a table. This can be done, for example, by adding an attached "header" field to each question. To generate a leadership table for a corporate document, the following questions can be included in the query set: Who is the CEO?:CEONAME What is the CEO's name?: CEONAME · How much does the CEO get paid?: CEOCOMP Who is the CFO?: CFONAME · How much does the CFO get paid?: CFOCOMP

[0079] Multiple questions with the same tag (CEONAME in the example above) are filtered out so that only the top scoring answer is retained for that tag. After all of the answers have been retrieved, a table is created. The table may be formatted as a CSV or TSV file for import into a spreadsheet or into a relational database. In this example, the headers may be "Name, Role, Compensation" and the table in the relational database may be constructed by using the following example instructions: · INSERT INTO leadership(id,Name,Role,Compensation)VALUES(1,CEONAME,'CEO',CEOCOMP); · INSERT INTO leadership(id,Name,Role,Compensation)VALUES(2,CFONAME,'CFO',CFOCOMP);

[0080] Yet another exemplary downstream process that may be implemented is the secondary source derivation process 236 (also identified as Process K) depicted in Figure 2. In some exemplary embodiments, the QA process is x) are used to determine the general concepts and subject areas to which the questions in the query set relate based on the level of completeness of the answers and / or the relevance / consistency scores calculated for the answers. However, it may occur that for some of the questions in the query set (including questions related to other questions that have generated highly relevant answers), no answers are generated, or the generated answers have low relevance scores. If information corresponding to the unanswered questions is needed (to generate a report or summary), system 200 (more specifically, the post-processing section of system 200) is configured to launch a search for information on locally or remotely available secondary sources. This search, in some examples, results in documents D submitted through QA system 220, applied to local documents (e.g., maintained in a repository of source documents and / or DOM documents), and processed / analyzed by QA system 220. x Determine what information is found to be unavailable from

[0081] The exemplary systems 100 and 200 for processing unstructured data using a question-answering system may, in some embodiments, be implemented on a general data processing system adapted to identify or perform broad searches of documents. Accordingly, with reference to FIG. 3, a diagram of an exemplary system 300 for document processing and response generation (which may also be adapted to derive structured output information) is provided. Further details regarding document processing implementations are provided in International Application No. PCT / US2021 / 039145, entitled "DOCUMENT PROCESSING AND RESPONSE GENERATION SYSTEM," the entire contents of which are incorporated herein by reference.

[0082] System 300 is configured to ingest source documents (e.g., a large library of customer documents or other repository of data, such as email data, collaboration platform data, etc.) or newly generated input documents (e.g., news items) and convert them into document objects (called Document Object Model, or DOM, documents) that represent a mapping from the source documents to searchable result objects (transformed documents). These document objects may be stored in a DOM repository (also called a knowledge distillation, or KD, repository). Users associated with the customer that provided the document library (e.g., employees of the customer) subsequently submit queries (e.g., natural language queries) that are processed by system 300. In situations where a quick answer is not available from a cache for commonly asked questions, the queries are processed and converted into a format compatible with the format of the ingested documents. The system then identifies portions within one or more of the ingested documents that contain answers to the user's queries. Output data is then returned to the user, including, for example, pointers to locations within one or more of the source documents that correspond to the identified one or more ingested documents. The user may then directly access it to retrieve the answer to their query. In some embodiments, the output may alternatively or additionally include the answer to the user's query and / or a portion of the document (e.g., a paragraph) containing the answer. Advantageously, the output returned to the user need not include the specific information sought by the user (although in some instances it may, if desired), but rather merely includes a pointer to a portion of the source document stored within a secure site that cannot be accessed by parties not authorized to access the source document. This answer determination technique therefore provides enhanced security for transmitting sensitive information (e.g., confidential or private information).As discussed herein, the system 300 may also be configured to automatically pose a set of predetermined questions (spanning multiple subject areas and concepts), for example via the query processing module 336, and generate answers with varying degrees of relevance and completeness that indicate the subject matter to which the content of the particular document being searched pertains.

[0083] In some embodiments, searching a document object repository to find answers to a query typically involves two operations: (1) first, a process called a fast lookup or fast match (FM) process is performed; and (2) next, a process called a detailed lookup or detailed match (DM) process (also referred to herein as a “dense detailed” lookup) follows the fast match process. Both the FM and DM processes may be based on the Bidirectional Encoder Representations from Transformers (BERT) model or any of the other models described herein (e.g., UniLM, GPT3, RoBERTa, etc.). For FM, the model (in some implementations) reduces to, for example, one vector for the query and one vector for a paragraph (e.g., a 200-word window that may also include contextual data). For DM, there are typically multiple vectors per query or paragraph, for example, proportional to the number of words or subwords in the query or paragraph. Alternatively, the data processing platform represented by system 300 may implement only a single language model transformation of fixed variable semantic resolution.

[0084] In some embodiments, transformation of the query and / or source document may occur on the customer's network, and the transformed query and / or transformed content is then transmitted to a central server. Such embodiments may improve privacy and security for transmitting sensitive data across a network, since the resulting vector (derived via transformation of the content or query data) is generated within the customer's (client's) secure space, and thus only the resulting transformed vector (rather than the actual content or query data) is available or present on the centralized cloud server. Transformation of the content or query data on the client's device may serve as a type of encryption applied to the transformed data, thus resulting in secure processing that protects the data from attacks on the server cloud. In some embodiments, data transformed on the client's network may additionally be encrypted to provide further enhanced secure communication of the client's data (whether source data or query data).

[0085] As depicted in FIG. 3 , system 300 typically includes document processing agent 310 (which may be an AI-based agent) in communication with customer network 350 a (one of n customer networks / systems accessing document processing agent 310 in exemplary system 300). Document processing agent 310 may be implemented as an independent remote server serving multiple customers, such as customer systems 350 a–350 n, and may communicate with such customers via network communications (either a private network or a public network such as the Internet). Communication with customer units is achieved via a communications unit including one or more communications interfaces (such as server interface 320, management interface 325, interactive user interface 330, and / or expert interface 332, all of which are represented schematically in FIG. 3 ), which typically include communications modules (e.g., transceivers for wired and / or wireless network communications configured according to various suitable types of communications protocols). Alternatively, the document processing agent 310 need not be located at a remote location, but may be a dedicated node within the customer network (e.g., the document processing agent 110 may be implemented as a process running on one of the customer's one or more processor-based devices, or may be a logically remote node implemented on the same computing device as a logically local node; i.e., it should be noted that the term "remote device" may refer to the customer station, while "local device" may refer to the document processing agent 310, or vice versa). An arrangement in which the agent 310 runs from the customer's network (such as any of the customer networks 350a-n) may improve data security, but may be more expensive to run personally.

[0086] Furthermore, in other alternative embodiments, some parts of the system (e.g., an ingestion unit configured to perform preprocessing and vectorization (parameterization) operations on source documents and / or user-submitted queries) may be located inside the customer's network firewall, while the storage of the captured documents (and optionally, a search engine for searching the captured content) may be located outside the customer's network firewall (e.g., on a centralized cloud server). In such alternative embodiments, data sent to the cloud server (e.g., for performing searches at a centralized location) may already be processed (e.g., through vector processing performed via coarse transformations (e.g., applied to fixed-size input segments) and / or fine-grained numerical transformations applied to smaller portions than those processed by the coarse transformations) into coded (captured) content that is incomprehensible to third parties not authorized to use the data, thus adding another measure of privacy and security protection to data processed using system 300. In these alternative embodiments, an initial portion of the processing of input queries may also be processed inside the customer's network firewall. In addition to performing transformations (of source content and / or queries) within the client's firewall, such transformed data may be further encrypted (using symmetric or asymmetric encryption keys) before being sent to the document processing agent 310, thus increasing the level of security / privacy achieved for communications between the customer's network and a centralized document processing agent (serving multiple customers).

[0087] An exemplary customer network 350a may be a distributed set of stations, potentially with dedicated secure gateways (protected by firewalls and / or other security measures) that can be controlled (from station 352) by an administrator. In one example, a customer has typically accumulated a large amount of electronic documents (e.g., technical documents appropriate to the customer's operations, administrative documents such as human resources documents, and all other types of written documents in electronic form). These documents are located in document library 360 (which may be a computing portion of customer network 350a) and are accessible by various authorized users at user stations 354a-c within network 350a and administrators (via administrator station 354). Any number of stations may be deployed within any particular customer network / system. Administrator station 352 may control access to documents in library 360 by controlling privileges or by managing the documents (e.g., access to specific documents in library 360, content management to obscure portions that do not comply with privacy requirements, etc.).

[0088] In addition to library 360 (which contains documents related to the operations of entities operating on the network), other sources of data or information may be available from various applications (e.g., email applications, chat applications such as Slack, customer relationship applications such as Salesforce, etc.) employed by customers for processing through the document processing implementations described herein. In yet additional embodiments, documents may be submitted to document processing agent 310 from third-party data providers (e.g., financial services providers, news agencies, etc.), and the content (possibly unstructured data) stored within the provider's document repository may be processed to generate meaningful structured output information that can be provided to one or more of the customers associated with networks 350a-n.

[0089] The administrator station 352 is configured to communicate with the document processing agent 310, for example, via the administration interface 325. Among other functions, the administrator may provide the document processing agent 310 with information identifying the location of source documents within a repository (library) 360 that maintains multiple source documents, the location of third-party document repositories that the customer wants to monitor and process, control the configuration and operation of the document processing agent 310's functionality related to the customer network 350a, scrutinize data generated by the agent 310 (e.g., ignore some responses), provide training data for the document processing agent 310, etc. Communications between the station 352 and the administration interface 325 may be established based on any communications technology or protocol. For enhanced security, communications between the document processing agent 310 and the administrator station 352 may include authentication and / or encryption data (e.g., by using symmetric or asymmetric encryption keys provided to the document processing agent 310 and the administrator station 352). Using the communications link established between the administrator station 352 and interfaces 320 and 325, the administrator provides the information necessary for the document processing agent 310 to access the document library. For example, the administrator station may send a message to the document processing agent 310 providing the network address of the document library 360 (and / or an identifier for a document within that library that agent 310 will access and process), or a message providing location information (e.g., a network address) of an individual document (or a third-party repository containing multiple documents) whose contents will be processed. The administrator station, in turn, may receive an encryption key (e.g., a private symmetric key or a public key corresponding to a private asymmetric key used by agent 310) used to encrypt the contents of the document to be transferred to agent 310.Communications between administrator station 352 and administrative interface 325 (or any of the other interfaces with which the administrator may communicate, such as interfaces 320 and 330) may also be used to establish other configuration settings that control the exchange of data and information between customer network 350a and document processing agent 310.

[0090] Once the document processing agent has been provided with the location (e.g., expressed as a network address) of the document library 360 or other source document, the agent 310 may begin receiving data transmissions of the document to be processed. The administrator station 352 controls the transmitted content and may perform any pre-transmission processing on the documents sent to the document processing agent 310, including removing sensitive content (e.g., private details), encrypting the content (e.g., by using a public key corresponding to a private key in the document processing agent 310), authenticating the data being transmitted, etc. The document processing agent 310 receives transmitted data from the customer network 350a via the server interface 320 and performs data pre-processing on the received data, including authenticating and / or decrypting the data, format conversion (if necessary), etc. The server interface 320 then sends data corresponding to the submitted document (subject to any preprocessing performed by the interface 320, which may include at least some of the preprocessing performed by the preprocessor 110 of FIG. 1 ) to the document ingestion engine 326, which processes the received document to convert it into a representation that enables the determination and generation of answers to queries provided by users of the network 350 a or queries of other sources (e.g., unstructured documents analyzed using a query set with a universe of questions). Typically, prior to applying the transformation, the source document undergoes preprocessing, such as segmentation into several parts (e.g., 200-word parts, or any other word-based segments). Segmentation is performed according to various rules that adjoin content from various parts of the document into discrete segments. As noted, one example of a preprocessing (i.e., pre-transformation) rule is to build segments using a fixed-length or variable-length sliding window, where one or more preceding headings are combined with the content captured by the fixed-length or variable-length sliding window. This creates a contextual association between one or more headlines and the content captured by the window.Such rules ensure that important contextual information and content located far from the segment being processed (e.g., far away in the source document) is combined by the transformations performed on the segment.

[0091] After segmenting the source documents and / or performing other types of preprocessing (such as that described above in connection with the preprocessing processor 110), the document ingestion engine 326 is configured to apply one or more types of transformations to the document segments to convert them into searchable segments (e.g., segments capable of question / answer search). As noted, one type of transformation that may be applied to the segments may be based on converting fixed-size (or approximately fixed-size) segments, typically containing multiple words / tokens, into numeric vectors to implement a fast search process. Such searches are typically coarse because they typically return a relatively large number of results (hits) (in response to queries submitted by users) because they are based on matching vectors generated from input data containing a relatively large number of words (tokens or features). As a result, the resolution achievable from such transformations is lower than that achievable from converting smaller segments. Therefore, results based on coarse vector transformations may not provide as accurate a representation of the textual meaning of the converted content as other transformations applied to smaller segments. A fast search can be performed relatively quickly and can therefore be used to differentiate possible candidates for possible answers (to a posed query) into a size or number that can then be more carefully searched (perhaps through a search based on another type of transform). Another transform that can be applied by the ingestion engine is to generate a dense-detail vector transform used to more narrowly pinpoint the location of an answer with a sequence of answer words specific to a segment (e.g., a paragraph) of text. Typically, the document segment to which the dense-detail transform is applied can be of finer granularity (resolution) (fast search segments are typically of a fixed size, e.g., 200 words, and therefore cannot typically pinpoint the exact location of an answer within the segment, if any). Any of the above types of transforms (to perform a fine or coarse search) can be implemented using one or more types of language model transforms, including BERT, GPT3, UniLM, and others.

[0092] With respect to the dense and detailed transformation performed by the document ingestion engine 326, source data (e.g., text-based portions segmented from a source document according to one or more rules or decision criteria; the segmented portions are typically smaller in size than the source segments used for fast search transformation) is typically transformed into multiple vectorized (numerical / parameterized) transformation contents. The dense and detailed transformation may also be performed according to any of the language model transformations described herein (including BERT). Processing by the document ingestion engine 326 may include natural language preprocessing to determine at least some language-based information (such as detecting and recording the location of named entities (e.g., people's names and company names) within the document), expanding structured data such as tables into searchable text equivalents, transforming information into knowledge representations (such as predetermined frameworks), extracting semantic meaning, and so on. In some embodiments, the resulting dense and detailed transformed data may be combined with the original content being transformed along with derived or provided metadata (such metadata is not critical but may facilitate intelligent search and question-answering performance of the document). In some examples, the combination of the transformed content with the source segment is further extended with automated questions that are germane to the source segment, and these generated questions may be paired with specific segments (or at specific locations within the entire document that includes the entire source content and the corresponding transformed content) or specific information fields. When processing a user's query, the similarity between the user's question and such automatically generated questions may be used to answer the user's question by returning output information (e.g., a pointer or actual user-understandable content).

[0093] 3, the content generated and captured by the document capture engine 326 is stored in a document object model (DOM) repository 340. The repository 340 is typically implemented on one or more data storage devices (distributed or available in a single local location) that are accessible from multiple access / interface points between the repository 340 and other modules / units of the document processing agent 310. 3, the repository 340 is depicted as having two access points: one that is a unidirectional link between the ingestion engine 326 and the repository 340 (i.e., a link that allows content to be written from the engine 326 to the DOM repository 340), and a bidirectional access point connected to the query processing module 336 that provides query data to the DOM repository 340 (to search DOM records stored in the repository) and receives search results (optionally forwarded after further processing, such as placing the searched content into a predetermined report content) that are forwarded to the user who submitted the query. In some embodiments, the access point to the repository may be implemented as a single point connected to a module configured to perform query processing and document ingestion operations.

[0094] The DOM repository 340 (which may be implemented similarly to repository 202 of FIG. 2 ) is configured to store, manage, and search DOM records 342 a-n (in conjunction with the document ingestion engine 326 and / or the query processing module 336). The content of a DOM record typically depends on the transformations performed by the document ingestion engine 326. A DOM record may contain data items related to a particular source document or portion of a source document. For example, a DOM record may be a collection of items including the original portion of the source document, metadata for that portion of the source document, contextual information related to that portion of the source document, corresponding coarse vectors resulting from transformations applied to one or more fixed-size (or nearly fixed-size) segments of the original portion of the source document (facilitating a fast search process), corresponding fine-grained transformation results resulting from fine-grained transformations (facilitating more accurate and refined text searches), and so on. Thus, if the conversion results in a vector of values representing the textual content of the segment, that vector is stored in the repository (appended or embedded in the vector), possibly in association with metadata and / or in association with the original content (in situations where the actual original textual content is preserved). In some embodiments, for security or privacy reasons, the source content may be ignored during ingestion or may be available only at the customer site. The metadata associated with the converted content may include contextual information related to the original source content and document location information that indicates the location or position of the source content that gave rise to the converted content within a larger source document. Such document location information may be provided in the form of pointer information that points to a memory location (or memory offset location) of the source document stored within the customer network; that is, when the pointer information is returned to a requesting user, it may be used to find a memory location where relevant content can be found that constitutes an answer to the user's query.It should be noted that some documents stored in DOM repository 340 (or in repository 202 of FIG. 2) may be substantially unstructured and may lack metadata or other information that indicates the nature of the content of the source that produced the DOM document. As noted, in such situations, the framework implemented by system 300 can be used to determine the structured information of documents that lack such information by applying a query set having predetermined questions covering a wide range of subject areas and concepts to the documents, and using the resulting answer data (level of completeness and relevance score) to determine the document's structured information.

[0095] The transformed content (which may include several transformed content items resulting from various transformations applied to the segmented content), metadata, various representations of the document's structural information, and / or source content stored in repository 340 may together define an integrated record structure, in which each of the transformed content, metadata, and / or original source content is a field or segment of the integrated record structure. Individual records may be related to each other (e.g., by arranging them consecutively or through logical or actual links / pointers) to define larger document portions (e.g., chapters of a particular document) when they correspond to discrete document segments of a larger source document, or to define the entire original document that was captured as a segment.

[0096] 3 , the document processing agent 310 further includes a query unit (also referred to as a query stack) configured to receive input (data representing queries from one or more users authorized to submit queries related to at least some of the ingested documents disposed in the DOM repository 340) and, in turn, provide output data returned to the initiating user. The query stack includes an interactive user query interface 330 in communication with a query processing module (also referred to as a query engine) 336. (The interactive user query interface may be similar to the server interface 320 and may be implemented using the same hardware and software as the server interface 320.) The query processing module may include a transformation engine that applies similar transformations to queries submitted by users and generates transformed query data that is compatible with the transformed content in the DOM records 342a-n maintained in the DOM repository 340. The transformed query may include coarse numeric-to-vector transformed data that can be used to search for numeric-to-vector transformed content in the repository 340, dense-detail transformed queries (that can be used to search for similarly transformed dense-detail transformed content in the repository 340), or any other transformed format that can be used to ingest the source document (using one or more source language model transformations).

[0097] Interactive interface 330 may be configured not only to receive process query data from a user and provide query output back to the user, but also to determine (by itself or in combination with modules other than agent 310) disambiguation information. Such disambiguation information may include disambiguation information originally provided (with the query) to aid in initial search / match operations (e.g., pre-filtering operations) performed on searchable content (either in DOM repository 340 or cache 335) managed by agent 310. Disambiguation information may also include post-filtering dynamically generated disambiguation information presented to the user to solicit the user to provide disambiguation information to resolve ambiguities present in two or more of the query results. For example, if two answers relate to the same or similar concept / category of information (such as an entity name, relevant contextual information, or some abstract concept derived using natural language processing or a learning machine implementation) but have different concept / category values, an intermediate output may be provided to the user (e.g., as a visual disambiguation prompt or an audio disambiguation prompt) requesting the user to provide disambiguation information defining which of the identified concepts is most relevant to the user's query. The disambiguation information returned by the user is then used to select one or more of the initial matches (and eliminate some other matches) and / or to rank the initial or remaining matches (based on relevance determined and calculated using input returned by the user). Further details regarding disambiguation processing are provided in International Application No. PCT / US2022 / 053437, entitled "Contextual Clarification and Disambiguation for Question Answering Processes," the entire contents of which are incorporated herein by reference.

[0098] In some embodiments, the interactive interface 330 is also configured to allow the user to provide personalization information that can be used to customize / personalize the query set applied to a particular (typically unstructured) document, revising or supplementing the library of questions to include more specific questions or to include questions covering additional concepts that may not have been covered (or not sufficiently covered) within the initial predefined set of questions intended to determine the concepts and subject areas to which the particular document pertains. The user's interactive input can also be obtained to control or personalize one or more of the post-search processes performed on the answer data determined from the application of the entire question set to a particular document. For example, the answer data resulting from the application of the entire question set may indicate that the document pertains to a company's financial data, thus triggering a process that generates multiple types of available reports. The user may then be prompted via the interactive interface 330 to select from available processes and / or to select options to customize the report format.

[0099] 3 , in embodiments in which repository 340 contains multiple types of transformed source content, a search of repository 340 may be implemented as a multi-pronged search. For example, because a coarse-valued numeric vector representation is generally more compact and easier to search (but may not be as accurate as a dense-detailed transformed representation, whether implemented by a BERT-based transform or some other transform), a first prong of a search to determine an answer to a submitted query may be to transform the query data into a coarse-valued vector representation and use that first transformed query expression to search for records in repository 340 that match the coarse-valued numeric base transform of the query data (e.g., according to some closeness criterion that may represent the distance or difference between the transformed vector query data and the transformed vector captured content data). This type of initial search may be referred to as a fast search. The results of the search may result in the identification of one or more answer candidates (e.g., identifying 1000 or any other number of possible segments that may contain the answer word sequence in response to the query submitted by the user). The identified first batch of possible results may then be used to conduct a second stage of search by converting the query into a dense-detailed transformed query and searching the dense-detailed transformed content associated with the search results identified in the first stage of the search process. This search stage may be referred to as a detailed (or fine-grained) search. It should be noted that in some embodiments, a fast search may be used to identify original portions of the source content associated with the identified candidates, which may then be converted into dense-detailed transformed content. In such embodiments, the repository 340 need not maintain the dense-detailed transformed content, but rather, the source content is converted based on which portions the fast search identifies as containing the answer to the query.In an alternative example, searching for answers to a query may be performed directly on the entire dense-detailed transformed content record without first identifying possible candidate portions of the source content through a fast search of the transformed content record with a fast search. In some embodiments, the fast search and the dense-detailed search may be performed on different transformed representations of the content of the document being searched. For example, the fast search may be performed on content transformed using one type of language model transformation (e.g., BERT), while the dense-detailed search may be performed on a different content transformation (e.g., according to a GPT3 language model).

[0100] Thus, in some embodiments, the query stack (e.g., the query processing module 336) is configured to transform the query data into transformed query data that is compatible with the transformed source content (e.g., compatible with one or more of the transformed content records in the DOM repository 340). For example, a fast search-compatible transformation can be a coarse BERT-based transformation (e.g., using a learning engine implementing the same or similar trained learning model used to generate the searchable transformed content from the source data) that is applied to the entire query data (e.g., natural language question) and generates a single-vector resultant. The query processing module can, for example, launch a fast search process (each numerical vector resulting from the coarse transformation) that identifies one or more candidate portions within the transformed source content that match the transformed query data according to a first determination criterion. For example, the matching operation can be based on a closeness or similarity determination criterion corresponding to a calculated distance metric between the calculated vector-transformed query data and various vector-transformed content records in the repository 340. As described herein, in some embodiments, the transformed content may include vectors corresponding to possible questions the user may ask to which source content provides possible answers. Thus, in some embodiments, fast search may compare the transformed query results (typically, the resulting vector records) with searchable vector records that represent possible questions that may be asked in connection with the source content from which those searchable vectors were generated.

[0101] The query processing module 336 may further be configured to determine, from the one or more dense-detail transformed content records corresponding to the one or more candidate portions identified based on those coarse-transformed vectors, at least one dense-detail transformed content record that matches the dense-detail transformed data of the query data according to a second decision criterion (e.g., some other closeness or similarity metric, or the same decision criterion applied with respect to the coarse-transformed data). Alternatively, in embodiments where fast searching is not performed, the query processing module 336 may be configured to identify one or more candidate portions of the transformed source content having respective dense-detail transformed content records that match the transformed query data according to the second decision criterion.

[0102] In some embodiments, interface 330 and / or the query processing module may be coupled to a query cache 335 and a question generation unit (which may be part of cache 335 or query processing module 336, or may be a separate unit). Query cache 335 stores answers / content corresponding to frequently asked questions, among other things. Such answers / content may include content previously retrieved from DOM documents (and / or from their corresponding raw source content) in response to previously submitted queries. Counters associated with such cached answers may track how frequently particular questions and answers are submitted and / or retrieved. Cache 335 may also be configured to ignore stale cache content that has not been accessed within a threshold time interval. Content in the answer cache may also be populated by an administrator (e.g., operating from a station such as station 352 via administration interface 325). The administrator may anticipate several plausible questions that users of customer system (network) 350a are expected to submit or ignore (e.g., determined to be inaccurate or non-compliant with the submitted query based on subsequent user feedback) content retrieved from DOM repository 340. Accordingly, in some embodiments, the query stack is configured to determine whether the received query data matches one of predefined questions (which may be stored in an answer cache), and, in response to determining that the received query data matches one of the predefined questions, generate output data based on one or more answer data records (possibly stored in the answer cache). In some embodiments, matching of the query data against previous questions and related answers stored in the cache is performed by calculating a score based on the combination of the question and its answer and ranking the calculated scores to identify one or more plausible matching candidates.

[0103] The query processing module may also include a question generation engine that may determine follow-up or related questions to one or more questions submitted through the query data (e.g., based on the trained learning engine and / or by using a repository of question data). Follow-up questions may be generated by paragraphing the submitted query (e.g., by submitting and / or standardizing the query to modify the submitted question, e.g., using the trained learning engine). In some embodiments, the answer data determined for the submitted query (e.g., based on content retrieved from the DOM repository 340 via the query processing module 336) may be processed (by a separate module) to formulate further questions from the answers. Such derived questions may then be resubmitted to the query processing module to retrieve follow-up answers. This process may be repeated iteratively up to a predetermined number of times. In some situations, the content stored in the DOM repository 340 may associate multiple questions (represented in any transformation format applied during the document ingestion phase) with processed segments of the source document. As noted, generating transformed content may include, for each processed segment, data representing a question associated with the processed segment, metadata, content that may be provided in a transformed format, and / or content from the original source. Thus, when a query is submitted (typically in a transformed format computed according to one or more language model transformations and / or different levels of content processing granularity), at least one DOM record / element is identified. The search results may potentially be associated with multiple questions, including questions that may have resulted in a match between the identified results and the submitted query. One or more of the additional questions (i.e., questions other than those that were matched with the query) may be used as separate queries to resubmit for searching to identify additional content that may be germane to the original query submitted by the user.

[0104] As noted, the generation of supplemental questions (also called query expansion) may be performed with respect to the overall question set that is applied to the document to determine structured information. For example, upon retrieval of the query set (including the overall question set) from cache 335 (or from some other storage device), at least some of the questions may be processed by the query processing module to generate follow-up questions or documents to be processed (e.g., document D in FIG. 2). x), or formulate synonymous questions (that may better align with the specific content of the document analyzed by the global question set) that can be submitted by the query processing module to apply supplemental questions to other documents in repository 340 (or in remote repositories accessible by document processing agent 310). Additionally, in some embodiments, the query processing module can be configured to generate supplemental questions based on answer data generated in response to applying the initial global question set to the particular document being processed. For example, based on identifying relevant answers to one or more of the questions (indicating the relevance of the document to related concepts or subject areas), the answer data of the identified related questions or related concepts or subject areas can be used to generate follow-up questions that may be related to the determined answer data, concepts, and subject areas, or to generate questions to determine missing data not included in the answer data to the identified relevant answers (e.g., using a machine learning model to generate labels corresponding to the identified answers, concepts, and / or subject areas). Additionally, the query processing module (or other unit of the document processing agent 310) may be configured to determine questions needed to complete information missing from a report or other structured output generated in response to answer data determined from applying the universe of questions (and / or additional iterations of question / answer processing following the initial application of the universe of questions). For example, in a situation involving the production of a financial report requesting, among other things, the salary of the CEO of a company referenced in an original source document being processed, the CEO's name and salary may not have been included in the source document. In this situation, the report generation process may cause the query processing module of the system 310 to generate questions (as may be directed by entity metadata and other contextual information associated with the document being sought) that are applied to the same and / or other documents in the repository 340 (or to remote repositories) related to the company in question to determine the CEO's name and CEO salary.

[0105] As further shown in FIG. 3 , determination of an answer to a query can be initiated by a user submitting a query 372 via a link 370 established between station 354a and interface 330. (As noted with respect to links established to transfer source documents for ingestion, the link can be based on any type of communication technology or protocol, including wireless and wireless communication protocols.) The query 372 can be the actual raw question submitted by the user, or it can be partially or completely transformed (e.g., for privacy and security reasons). For example, station 354a can apply a transformation commensurate with that applied by ingestion engine 326 (in which case, performing a similar transformation in the query stack may be unnecessary). Alternatively or additionally, authentication and encryption processing can be performed on the query 372. The query (question data) 372 is sent to document processing agent 310 and received at user query interface 330. Upon receiving the query, a determination can be made as to whether an appropriate answer is available within cache 335 of predetermined answers. If there are predefined questions / answers (e.g., if the query data matches one or more predefined questions), one or more of the predefined answers are used to generate output data (shown as output data 374) that is returned to the user via link 370 (or via some other link). As noted, in some embodiments, the submitted query may be a set of predefined questions configured to help determine structured information for an otherwise unstructured (or poorly structured) document. In such embodiments, the query set may be stored locally in the document processing agent (e.g., in a cache) and automatically applied to certain documents, for example, newly arriving documents that satisfy some conditions (e.g., documents arriving from some pre-specified third-party data provider or vendor).

[0106] Generally, the query data is transformed by the query stack (if not already transformed at station 354a) into transformed query data. The transformed data may provide the query in one or more transformed formats that are compatible with the format of the transformed source content stored in the DOM repository 340. In some embodiments, the query data may also be used to generate one or more additional questions (e.g., follow-up questions or questions related to the original query submitted by the user). In situations where an answer to a query is available from the answer cache, the answer itself may be used as a basis for generating one or more further questions related to the cached answer. The query or transformed query is used to search the DOM repository 340 via the query processing module 336. As noted, the search may be performed as a multi-way process that conforms to the multiple transformation formats used to store the data in the DOM repository 340.

[0107] The output generated in response to a submitted query may include a pointer to source content available on customer network 350a. In such an embodiment, because the data stored in repository 340 is populated based on source documents maintained in a document library available on the customer network accessible to the user submitting the query, the source documents may not have been stored in their original format in document processing agent 310 (e.g., for security reasons to protect sensitive data from being compromised), and the output returned to the user does not require the actual answer data to be sent back to the user. Instead, the pointer returned as query output may identify the address or location of the answer within an appropriate document available to the user on user network 350. For example, in the example shown in FIG. 3, output data 374 is shown as a pointer to the specific location of the answer within document 362a (stored in library 360 along with documents 362b-d). Such a pointer may therefore include data representing the document 362a, e.g., data representing the network address or memory location where the beginning of the document is found, and the specific location (e.g., a relative offset from the beginning of the beginning location of the document 362a or from the actual address or memory location where the beginning of the identified portion is found) of the portion of the document that represents the answer to the question asked by the user at station 354a. The pointer data provided in the output data may be included in a metadata field of a DOM record that includes data of the transformed content that is determined (e.g., by the query processing module 336) to match (according to one or more applied matching criteria) to the query submitted by the user. In some embodiments, the output data may include, in addition to or instead of the pointer data, at least a portion of the source content that corresponds to at least a portion of the transformed content and / or a summary of the source content that corresponds to at least a portion of the transformed content.

[0108] As discussed in connection with Figures 1 and 2, the output generated in response to answer data determined through application of the overall question set to one or more unstructured documents includes different types of structured output data generated by corresponding post-QA processing. Such post-QA processing may occur in the document processing agent 310 (e.g., by the query processing module 336 and / or the interactive query interface 330) or may occur on another computing device located remotely from the document processing agent 310. Such post-QA processing (i.e., downstream processing) may include report / summary generation processes, document generation processes (e.g., for automatically generating various required legal or administrative documents), database management processes (for automatically populating a database or repository with structured data records), data mining processes, and other types of processes based on information derived from the analyzed documents, as discussed in connection with systems 100 and 200 of Figures 1 and 2, respectively.

[0109] As discussed in connection with FIG. 1 , source documents typically undergo preprocessing, including secure communication processing (decryption and / or authentication), document formatting, initial context discovery, content transformation (e.g., converting text to a vector representation), and other preliminary processes to convert the source document into a QA-searchable document. One example of a document formatting preprocessing procedure is segmenting the source content of a source document into multiple document segments. Such segmentation may be performed according to hierarchical rules that semantically relate one portion of the source document to one or more other portions of the source content. For example, a sliding window of fixed or variable size (e.g., 200 words) may be applied to the source content to generate manageable-sized segments onto which content transformations are applied. However, when segmented into small chunks, the content segments may lose important contextual information that would otherwise be available for larger-sized segments. For example, a passage in the middle of a section of a document may not, by itself, contain important contextual information, such as the section heading, the location of the passage relative to an earlier passage within the section (e.g., if the current passage is a footnote), or font size relative to other passages not captured by a particular segment. Thus, in some embodiments, contextual information (e.g., section heading, chapter heading, document title, location, font type and size, etc.) may be combined with one or more of the document segments. This preprocessing procedure is shown in FIG. 4 , which provides a diagram of an exemplary document preprocessing (also called ingestion) procedure 400. In FIG. 4 , source content 410 (which may be a portion of the source document) has been segmented into segments 420 a-n. Each segment has its own individual segmented content (resulting from applying a segmentation window to the source content) that may be combined with contextual information (which may be textual information, numerical information, or both) associated with each segment. As will be appreciated, at least some of the contextual information, for example, the document identifier ("Doc a"), the chapter information (Chapter S), and the heading information (Section x), is common to the segments shown in FIG.This allows a transformation subsequently applied to the segment to preserve at least some of the contextual information and thus preserve some of the relevance of the segment being transformed to the subject.

[0110] In some examples, to simplify the segmentation process (to facilitate more efficient search and retrieval), a source document may be segmented to generate overlap between consecutive document segments (without adding contextual information to each segment separately). Thus, for example, in situations where segments are generated by windows of some particular (fixed or variable) size, the windows may be shifted from one position to the next by a predetermined fraction of the window size, e.g., ¾ (150 words for a 200-word window). As a result of the fractional shift, a transform (e.g., a language model transform) applied to the overlapped segments may generate correlations between the segments, which may preserve the associations between consecutive segments for subsequent QA searches. In some embodiments, headline information (and other contextual information) may be added directly to the segmented segments. Alternatively, the headline and contextual information may be converted into vectors that are added to vectors resulting from transform operations applied to the content extracted by the sliding window, or may be combined with the content extracted by the window before the transform is applied to the resulting combined data. Correlating neighboring segments with each other (e.g., via slight shifting of a window over a document to form a segment) improves the identification of relevant paragraphs (in response to a submitted query) for retrieval and presentation of lead paragraphs and related answer snippets.

[0111] Another preprocessing step that can be applied during source document segmentation involves handling tabular information (i.e., when the original content is arranged in a table or grid). This preprocessing step is used to expand structured data arranged in a table (or other type of data structure) into a searchable format, such as a text equivalent. For example, if a portion of the source document is identified as a multi-cell table, multiple substitute portions are generated to replace the multi-cell table, each of the multiple substitute portions including respective subcontent data and contextual information related to the multi-cell table. Additional example preprocessing steps include steps for associating contextual information with one or more portions of the source document, for example, based on a) information provided by a user in response to one or more questions related to the source document presented to the user, and / or b) one or more ground truth examples of question / answer pairs.

[0112] In some instances, contextual information may not be explicitly included with the segment, but instead may need to be discovered and included with the document segment as extended information (in this case, extended contextual information). For example, entity discovery (determining identifiers of relevant entities referenced within a document) can be used to help speed up searches and improve search accuracy.

[0113] Consider the following example implementation: Each search unit (e.g., 200-word window, paragraph, document, etc.) is analyzed for unique entities associated with the search unit and for metadata associated with task-specific entities (e.g., HR, author, organization, etc.). Each search unit is tagged with appropriate intrinsic and metadata entities. During the search, different heuristics may be used that may remove many of these search units by briefly identifying them as irrelevant to the query. For example, the user's question may be determined with high confidence to relate to some specific subject matter (because, for example, of a user-explicit identification of the subject matter (e.g., a question stating "I have a financial question"), or the subject matter may be inferred via rules or a classification engine to relate to a specific subject matter.) In one use case, all documents / document objects of other subjects (HR, security, etc.) may be removed from further consideration; these documents do not need to be searched in response to the submitted query. A by-product of such filtering is to speed up, for example, FM and DM searches. In addition, potential answer units from irrelevant categories will not generate false recognition errors, and consequently, this helps to improve the accuracy of the search.

[0114] Information about the particular entity (or entities) relevant to the user's search can also be used to generate more precise follow-up questions (e.g., to determine different ways to paragraph the input query so that additional potential question / answer pairs can be generated), and to provide additional context that can be used to search the repository of data (whether DOM objects in a transformed format or in a user-readable data format).

[0115] In some embodiments, document preprocessing can be performed as two separate tasks. In one processing task, the source document is appropriately segmented and organized into small chunks (e.g., paragraphs) with additional extensions (e.g., a vector sequence representing a section heading can be added to the vector of every paragraph in that section). These extensions are used to improve search accuracy. In a parallel task, the document is segmented in a manner that is most appropriate for presentation purposes. The results of the two different segmentation outputs need to be correlated so that leading paragraphs and relevant answer fragments are identified during the search process, but what is presented to the user is presentation content related to the identified answer fragments (rather than the identified answer fragments). In other words, the system can capture specific passages to facilitate search operations and separately capture specific passages to facilitate presentation operations. In this example, once a passage is identified as a result of matching a query with searchable content, presentation content related to the identified passage is output.

[0116] Once the source document is segmented into multiple segments, each segment may be provided to one or more content transformers (or transformers) 430 a-m, which transform the segment (content and optional contextual information; however, in some embodiments, the contextual information may be preserved without transformation) into resulting transformed content associated with questions and answers related to the original content of the respective segment. The example of FIG. 4 shows m transforms, each applied to any one of the segments (e.g., segment 420 j). While the same segment (e.g., 420 j) is shown to be provided to each of the transforms, in some embodiments, different segmentation procedures may be applied to obtain segments of different sizes and configurations as required by the individual transforms (e.g., the coarse-fast-search BERT-based transform 430 a may be configured to apply to segments of a first segment size, while the fine-fine BERT-based transform 430 b may be configured to apply to segments of a second, different size (e.g., strings of several words)).

[0117] The transformation module may be implemented via a neural net pre-trained to generate transformed content associated with question-answer pairs. Other transformation implementations may be achieved through the use of filters and algorithmic transformations. Training of the neural net implementation may be achieved with a large training sample of question-answer ground truth, which may be publicly available, or may be developed internally / privately by a customer that uses a document processing system (such as system 300 of FIG. 3) to manage its document library.

[0118] 5, a flowchart of an exemplary procedure 500 for guided intelligent document processing via automated question answering is shown, which may be implemented for exemplary systems such as systems 100, 200, and 300 of Figures 1-3, respectively. Procedure 500 includes obtaining a query set 510, conducting a question-answer (QA) search on one or more documents using the query set 520 to characterize concepts associated with the one or more documents, where the answer data is responsive to one or more questions contained in the query set, and deriving structured output information for the one or more documents based on the answer data generated in response to conducting the question-answer search 530.

[0119] As noted, the query set may be a global collection of questions related to multiple different content subject areas. That is, the query set may include a wide range of questions covering multiple topics, concepts, and subject areas (including financial issues, legal matters, sports, national and international affairs, etc.). In some examples, obtaining the query set may include tailoring the predetermined question set based on user information associated with the user. Such user information may include, for example, one or more of the user's personal preferences, network access controls associated with the user, and / or network groups to which the user is associated. In some embodiments, procedure 500 may further include determining follow-up queries based on at least some of the answer data and conducting follow-up question-answer searches of one or more documents using the follow-up queries. In such embodiments, determining the follow-up queries may include using one or more ontologies that define relationships and associations between concepts identified from at least some of the answer data and other distinct concepts, and deriving follow-up questions for the follow-up queries based on the other distinct concepts determined using the one or more ontologies.

[0120] As discussed herein, structured output data is generated by using one or more downstream post-QA processes (e.g., to generate a report or summary, determine additional information from secondary sources, perform data mining and clustering, etc.). Thus, deriving structured output information for one or more documents may include, for example, one or more of the following: i) determining classification information for one or more documents that represent at least one of the concepts; ii) performing data clustering of the one or more documents based on the answer data; iii) applying a data discovery process to the answer data to determine one or more labels associated with concepts associated with the one or more documents; iv) generating an output report based on the answer data; and / or v) deriving supplemental data associated with at least some of the answer data. For example, deriving supplemental data associated with at least some of the answer data may include determining supplemental concepts associated with at least some of the answer data, accessing at least one of one or more documents or another data source, and determining supplemental information related to the supplemental concepts from the accessed at least one of the one or more documents or other data sources. In such embodiments, determining supplemental concepts may include determining supplemental questions to apply to at least one of the one or more documents or other data sources.

[0121] In some examples, generating the output report may include, for example, one or more of the following: i) generating a summary report to be provided to the user based on at least some of the answer data arranged in one or more predetermined templates, ii) generating an alert to be communicated to the user, and / or iii) entering at least some of the answer data into a database table. In some examples, generating the output report may include: determining a score of the answer data generated in response to conducting a question-answer search using the query set, and including in the output report a predetermined number N1 of the highest-scoring answers determined from the answer data. In such examples, procedure 500 may further include identifying additional responses from the answer data results, whose respective scores exceed a predetermined score threshold, and selecting up to N2-N1 (where N2>N1) selected answers for inclusion in the output report from the additional answers, whose respective scores exceed the predetermined score threshold.

[0122] Generating the structured output information may include generating the structured output information based on the response data and further based on user information associated with the user, which may include, for example, one or more of the user's personal preferences, network access controls associated with the user, and / or network groups with which the user is associated.

[0123] In some embodiments, deriving the structured output information may include applying one or more machine learning models to at least some of the response data.

[0124] Procedure 500 may further include determining a score for the answer data generated in response to performing a question-answer search using the query set. In such an example, generating structured output information for one or more documents may include generating the structured output information for the documents based on the determined scores for the answer data. Generating the structured output information may include determining that one or more documents are unrelated to one or more of the different content subject areas (related to the questions constituting the entire set of the query set) based on the determined scores for the answer data generated from the plurality of questions related to one or more of the different content subject areas. Determining the score for the answer data may include calculating a score for a particular answer in response to a particular question from one or more questions in the query set, representing, for example, one or more of the following: similarity of the particular answer to the particular question; similarity of the combination of the particular question and the particular answer to a predetermined question / answer pair for one or more documents; similarity of the particular answer to a previously selected answer provided to a particular user; the relative location of the particular answer in one or more documents; and / or the level of detail included in the particular answer.

[0125] As noted, the framework described herein may perform preprocessing (also referred to as ingestion) on received source documents. In such an embodiment, procedure 500 may further include receiving one or more source documents and transforming the one or more source documents into one or more documents on which QA searches are performed. Transforming the one or more source documents may include applying one or more segmentation preprocessing processes to the one or more source documents to generate one or more segmented documents and applying one or more vector transforms to the one or more segmented documents to transform the one or more segmented documents into respective vector answers in one or more vector spaces. Applying the one or more vector transforms may include transforming the segmented one or more documents according to, for example, one or more of the Bidirectional Encoder Representations from Transformers (BERT) language model, GPT3 language model, T5 language model, BART language model, RAG language model, UniLM language model, Megatron language model, RoBERTa language model, ELECTRA language model, XLNet language model, and / or Albert language model.

[0126] In some embodiments, deriving the structured output information may be further based on interaction data provided by a user, for example, the interaction data may be provided in response to prompt data generated by the QA system and may include disambiguation data for selecting an answer from among multiple matches in the answer data related to one or more similar concepts.

[0127] In the learning machine-based implementations described herein, different types of learning architectures, configurations, and / or implementation techniques may be used. Examples of learning machines include neural networks, including convolutional neural networks (CNNs), feedforward neural networks, recurrent neural networks (RNNs), etc. A feedforward network includes one or more layers of nodes ("neurons" or "learning elements") connected to one or more portions of input data. In a feedforward network, the connectivity between the input and the layers of nodes is such that the input data and intermediate data propagate in a forward direction toward the output of the network. There are typically no feedback loops or cycles within the configuration / structure of a feedforward network. Convolutional layers enable the network to efficiently learn features by applying the same learned transformation to subsections of data. Other examples of learning engine techniques / architectures that can be used include generating an autoencoder; using dense layers of a network to correlate with the probability of future events via an auxiliary vector machine; building a regression or classification neural network model that dictates a particular output from the data (based on training that reflects the correlation between similar records and the output to be identified); and so on.

[0128] Neural networks (and other network configurations and implementations for realizing the different procedures and operations described herein) may be implemented on any computing platform, including computing platforms that include one or more microprocessors, microcontrollers, and / or digital signal processors that provide processing functionality as well as other computational and control functionality. The computing platform may include one or more CPUs, one or more graphics processing units (e.g., GPUs (such as NVIDIA GPUs) that can be programmed according to the CUDA C platform), and may also include special purpose logic circuitry (e.g., FPGAs (field programmable gate arrays), ASICs (application-specific integrated circuits), DSP processors, accelerated processing units (APUs), application processors, customized special purpose circuitry, etc.) for at least partially implementing the processes and functionality of the neural networks, processes, and methods described herein. Computing platforms used to implement neural networks also typically include memory for storing data and software instructions for executing the programmed functionality within the device. Generally speaking, a computer-accessible storage medium may include any non-transitory storage medium that can be accessed by a computer during use to provide instructions and / or data to the computer. For example, a computer-accessible storage medium may include a magnetic or optical disk, a semiconductor (solid-state) memory, DRAM, SRAM, etc.

[0129] The different learning processes described herein implemented through the use of neural networks may be configured and programmed using TensorFlow, an open source software library used for machine learning applications such as neural networks. Other programming platforms that may be employed include keras (an open source neural network library) building blocks, NumPy (an open source programming library useful for implementing modules to process arrays) building blocks, etc.

[0130] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly or customarily understood. As used herein, the articles "a" and "an" refer to one or more than one (i.e., at least one) of the grammatical object of the article. By way of example, "an" means one element or more than one element. "About" and / or "approximately," as used herein when referring to a measurable value such as a quantity, a time duration, or the like, encompasses a deviation of ±20%, ±10%, ±5%, or ±0.1% from the stated value. Such deviations are appropriate in the context of the systems, devices, circuits, methods, and other implementations described herein. "Substantially," as used herein when referring to a measurable value such as a quantity, a time duration, a physical attribute such as frequency, or the like, also encompasses a deviation of ±20%, ±10%, ±5%, or ±0.1% from the stated value. Such deviations are appropriate in the context of the systems, devices, circuits, methods, and other implementations described herein.

[0131] As used herein (including in the claims), "or" used in a list of items preceded by "at least one of" or "one or more of" indicates a disjunctive list, such as the list "at least one of A, B, or C" meaning A or B or C or AB or AC or BC or ABC (i.e., A and B and C), or a combination of two or more features (e.g., AA, AAB, ABBC, etc.). Also, as used herein, unless otherwise specified, a statement that a function or operation is "based on" an item or state means "the function or operation is based on the listed item or state, and may be based on one or more items and / or states in addition to the listed item or state."

[0132] While specific embodiments have been disclosed in detail herein, this has been done by way of example for purposes of illustration only and is not intended to limit the scope of the present invention as defined by the appended claims. Any of the features of the disclosed embodiments described herein can be combined with one another, rearranged, etc., within the scope of the present invention to create more embodiments. Certain other aspects, advantages, and modifications are believed to be within the scope of the claims provided below. The presented claims represent at least some of the embodiments and features disclosed herein. Other unclaimed embodiments and features are also contemplated.

Claims

1. Obtaining a query set; conducting a question-answer (Q-A) search of one or more documents using the query set to generate answer data responsive to one or more questions contained in the query set, the answer data characterizing concepts related to the one or more documents; deriving structured output information for the one or more documents based on the answer data generated in response to performing the Q-A search; and A method comprising:

2. Deriving the structured output information for the one or more documents comprises: determining classification information for the one or more documents that represent at least one of the concepts; performing data clustering for the one or more documents based on the response data; applying a data discovery process to the response data to determine one or more labels associated with the one or more concepts associated with the one or more documents; generating an output report based on the response data; and deriving supplemental data associated with at least some of the response data; The method of claim 1 , comprising one or more of:

3. Deriving the supplemental data associated with at least some of the response data includes: determining complementary concepts related to at least some of the response data; accessing at least one of the one or more documents or another data source; determining supplemental information related to the supplemental concept from the accessed at least one of the one or more documents or the other data source; The method of claim 2 , comprising:

4. The method of claim 3 , wherein determining the supplemental concepts comprises determining and applying supplemental questions to the at least one of the one or more documents or the other data source.

5. generating the output report comprises: generating a summary report to be provided to a user based on at least some of the response data arranged within one or more predetermined templates; generating an alert that is communicated to the user; or entering at least some of said response data into a database table; The method of claim 2 , comprising one or more of:

6. generating the output report comprises: determining a score for the answer data generated in response to performing the question-answer search using the query set; A predetermined number N having the highest score determined from the response data 1 including the answer in said output report; The method of claim 2 , comprising:

7. identifying additional answers from the response data results whose respective scores exceed a predetermined score threshold; From the additional answers whose respective scores exceed the predetermined score threshold, select a maximum of N answers for inclusion in the output report. 2 -N 1 (where N 2 >N 1 ) selected answer and The method of claim 6 further comprising:

8. 8. The method of claim 1, wherein generating the structured output information comprises generating the structured output information based on the response data and further based on user information associated with the user.

9. The user information is The user's personal preferences, network access controls associated with the user, or network groups associated with the user The method of claim 8 , comprising one or more of:

10. determining additional queries based on at least some of the response data; conducting an additional question-answer search for the one or more documents using the additional query; The method of any one of claims 1 to 9, further comprising:

11. Determining the additional query comprises: determining, using one or more ontologies defining relationships and associations between the concepts identified from at least some of the response data and other distinct concepts; deriving additional questions for the additional query based on the other different concepts determined using the one or more ontologies; and The method of claim 10, comprising:

12. determining a score for the answer data generated in response to performing the question-answer search using the query set; further comprising 12. The method of claim 1, wherein generating the structured output information for the one or more documents comprises generating the structured output information for the one or more documents based on the determined score for the answer data.

13. the query set includes a universe of questions related to a plurality of different content subject areas; 13. The method of claim 12, wherein generating the structured output information includes determining that the one or more documents are unrelated to one or more of the plurality of different content subject areas based on the determined scores for the answer data generated from the plurality of questions related to the one or more of the plurality of different content subject areas.

14. Determining the score for the response data includes: For a particular answer in response to a particular question of the one or more questions in the set of inquiries, the similarity of the particular answer to the particular question, the similarity of the combination of the particular question and the particular answer to predefined question-answer pairs for the one or more documents, the similarity of the particular answer to previously selected answers provided to a particular user, the relative location of the particular answer in the one or more documents, and the level of detail included in the particular answer. The method of claim 12 , comprising calculating a score representing one or more of:

15. generating the structured output information includes: applying one or more machine learning models to at least some of the response data; The method according to any one of claims 1 to 14, comprising:

16. Obtaining the query set includes: Tailoring the set of predefined questions based on user information associated with the user The method according to any one of claims 1 to 15, comprising:

17. 17. The method of claim 16, wherein the user information includes one or more of the user's personal preferences, network access controls associated with the user, or network groups with which the user is associated.

18. receiving one or more source documents; Transforming the one or more source documents into the one or more documents on which the QA search is performed; The method of any one of claims 1 to 17, further comprising:

19. Transforming the one or more source documents includes: applying one or more pre-segmentation processes to the one or more source documents to generate one or more segmented documents; applying one or more vector transformations to the one or more segmented documents to transform the one or more segmented documents into respective vector answers in one or more vector spaces; 20. The method of claim 18, comprising:

20. applying the one or more vector transformations Transforming the one or more segmented documents according to one or more of a Bidirectional Encoder Representations from Transformers (BERT) language model, a GPT3 language model, a T5 language model, a BART language model, a RAG language model, a UniLM language model, a Megatron language model, a RoBERTa language model, an ELECTRA language model, an XLNet language model, or an Albert language model.

20. The method of claim 19, comprising:

21. The method of any one of claims 1 to 20, wherein deriving the structured output information is further based on interaction data provided by a user.

22. 22. The method of claim 21, wherein the interactive data is provided in response to prompt data generated by a Q-A system and includes disambiguation data for selecting an answer from among a plurality of matches in the answer data related to one or more similar concepts.

23. one or more memory storage devices for storing executable computer instructions and data; a processor-based controller electrically coupled to the one or more memory storage devices; the controller comprising: Obtaining a query set; conducting a question-answer (Q-A) search on one or more documents using the query set to generate answer data responsive to one or more questions contained in the query set, the answer data characterizing concepts related to the one or more documents; deriving structured output information for the one or more documents based on the answer data generated in response to performing the Q-A search; and A system that is configured to:

24. The controller configured to derive the structured output information for the one or more documents, determining classification information for the one or more documents that represent at least one of the concepts; performing data clustering for the one or more documents based on the response data; applying a data discovery process to the answer data to determine one or more labels associated with the concepts associated with the one or more documents; generating an output report based on the response data; and deriving supplemental data associated with at least some of the response data; 24. The system of claim 23, configured to do one or more of the following:

25. The controller further comprises: determining a follow-up inquiry based on at least some of the response data; The system of any one of claims 23 to 24, configured to perform an additional question-and-answer search for the one or more documents by using the additional query.

26. The controller configured to determine the further query comprises: using one or more ontologies that define relationships and associations between the concepts identified from at least some of the response data and other distinct concepts; 26. The system of claim 25, configured to derive additional questions for the additional inquiry based on the other different concepts determined using the one or more ontologies.

27. The controller further comprises: configured to determine a score for the answer data generated in response to performing the Q-A search using the query set; the controller configured to generate the structured output information for the one or more documents is configured to generate the structured output information for the one or more documents based on the determined scores of the answer data. A system according to any one of claims 23 to 26.

28. the query set includes a universe of questions related to a plurality of different content subject areas; the controller configured to generate the structured output information is configured to determine that the one or more documents are unrelated to one or more of the plurality of different content subject areas based on the determined scores for the answer data generated from the plurality of questions related to the one or more of the plurality of different content subject areas.

28. The system of claim 27.

29. The controller configured to obtain the query set comprises: configured to tailor the set of predetermined questions based on user information associated with the user, the user information including one or more of personal preferences of the user, network access controls associated with the user, or network groups associated with the user; A system according to any one of claims 23 to 28.

30. Obtaining a query set; conducting a question-answer (Q-A) search on one or more documents using the query set to generate answer data responsive to one or more questions contained in the query set, the answer data characterizing concepts related to the one or more documents; deriving structured output information for the one or more documents based on the answer data generated in response to performing the Q-A search; and A non-transitory computer-readable medium programmed with instructions executable on one or more processors of a computing system to perform the steps of: