A system and method for training machine learning models to recognize concepts in multimedia documents through natural language interaction and mixed-instructive learning.
The system addresses inefficiencies in legal pipelines by training machine learning models through natural language interactions, enabling dynamic knowledge integration and automation, thus improving communication and reducing errors across legal processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- RARE AI INC
- Filing Date
- 2024-05-17
- Publication Date
- 2026-06-03
Smart Images

Figure 2026518079000001 
Figure 2026518079000002 
Figure 2026518079000003
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 503,346, filed on May 19, 2023, the content of which is incorporated herein by reference.
[0002] This disclosure generally relates to training machine learning models, and more specifically, to training machine learning models through natural language interactions.
Background Art
[0003] The current legal pipelines surrounding corporate, investigative, and litigation environments have numerous challenges and inefficiencies. These environments often classify documents according to special concepts specific to a company or task, using either rules (e.g., keyword-based rules) or datasets for training machine learning models. However, modern tools require user proficiency and typically adopt a single method of training the system, which requires the user to provide labeled examples that match and do not match concepts from the documents, or requires the user to provide decision rules such as boolean queries.
[0004] Examples of such problems are found in litigation, governance, and compliance. In an exemplary governance / compliance scenario, a company hires a team of compliance experts who use a dedicated tool that combines keyword searches and machine learning from examples. These experts devise search terms to identify content that violates corporate governance rules or regulatory requirements in order to avoid litigation. In some cases, the experts may also tag exemplary cases of violations to train a machine learning model for automatic identification.
[0005] Ultimately, violations can escalate to litigation, triggering the export of corporate data to a law firm. This firm employs document reviewers to find documents that corroborate litigation against the company, a process often similar to that performed by corporate compliance specialists. Throughout the litigation, other team members, such as associates and partners, scrutinize the data found by the document reviewers for use in later stages, including depositions and court preparation.
[0006] However, the knowledge gathered during litigation is not used to improve machine learning models or to inform corporate compliance / governance experts. This valuable information, which could improve corporate data governance and document review, is effectively wasted due to poor communication between different groups of experts using different tools. Two possible reasons for this situation are (1) that companies and law firms use different tools, and (2) that companies and law firms employ different experts for tasks ranging from governance to document review and depositions.
[0007] As a result, information flows linearly through the process, and most of it remains inaccessible until it has been processed in a preceding stage. This leads to three notable consequences: (1) time delays, (2) communication barriers due to different tools, and (3) the propagation of errors.
[0008] Regarding time delays, human intervention at each stage slows down information dissemination, and it takes time for data changes to reach later stages. For example, changes in a document submission request take time to be reflected in the document review protocol, and reviewers need time to find the corresponding new documents. As a result, an associate preparing an impression may not have access to the most relevant documents at a critical time.
[0009] Regarding communication barriers, the use of different tools and experts at different stages of legal work creates communication barriers, leading to fragmentation and expertise bubbles. For example, when associates acquire new information during depositions, they typically lack the training or resources to use document review tools to add information by tagging additional documents. Similarly, corporate compliance and governance experts employ different tools, making it difficult to share or adapt predictive models built during litigation in a timely manner. This hinders the effective transfer of knowledge about legal issues accumulated during litigation in order to minimize future risks.
[0010] Regarding the propagation of errors, current tools for litigation and compliance / governance primarily consist of keyword searches and machine learning-based predictive tools that learn from examples of tagged documents. These tools introduce the possibility of either Type 1 (false positive) or Type 2 (false negative) errors. Because current workflows proceed linearly from compliance / governance to early assessment of litigation cases, document review, and depositions, errors accumulate at each stage.
[0011] Legal work ranging from compliance / governance to early assessment of litigation cases, document review, and depositions shares at least a few attributes: namely (1) custom specifications, (2) natural language, and (3) dynamic specifications / queries.
[0012] Regarding custom specifications, every company needs its own specifications that define the information that needs to be identified as part of standard governance and compliance due diligence. Each company has its own guidelines, rules for governance, and evolving administrative regulations.
[0013] Regarding natural language, custom specifications are now often defined in natural language for the professionals who implement them today, and communication involves various types of documents and people.
[0014] Regarding dynamic specifications / queries, the process of defining the specification is often dynamic, highly interactive, and linked to queries for data that will be obtained by applying that specification. Specifications and queries evolve across the organization over time. Initial specifications yield initial results, but other team members use them to redefine or expand upon the specification. This takes the form of continuous two-way communication among team members at different stages of the pipeline.
[0015] Some of the challenges and barriers created by some of the existing solutions used in these legal processes include (1) pipeline workflows that restrict and delay information sharing, and (2) inflexible, narrow-field tools that hinder the specification and querying of knowledge.
[0016] Regarding the pipeline workflow, information propagates linearly through the process, with each stage employing specialized experts using narrowly focused tools. New information obtained at later stages does not easily or quickly propagate to earlier stages for specification updates. Experts must communicate ad-hoc, leading to delays in information transfer and limited specification changes based on new information. Associates cannot directly input new information into the system; therefore, they must communicate with other members of the pipeline.
[0017] Regarding inflexible, narrow-field tools, some existing solutions target narrow expertise and are not universally usable by everyone involved in the legal process. Various experts must be contacted to request specific changes to the specification or to formulate information query requests. However, current tools support only limited forms of interaction, which is also problematic. To train new information, existing tools can only learn from many examples of the desired new information. To define new information as an updated specification, reviewers may need to tag the new examples and update the machine learning model, which can be laborious and time-consuming. For querying, most queries are formulated using Boolean searches, which are highly narrow and accurate, but often miss many relevant documents.
[0018] Figure 4 is an illustrative Figure 400 showing a legal workflow with several existing solutions. Early assessment of litigation cases is performed by creating keyword search filters used by attorneys and litigation analysts to identify relevant documents. Document review is typically performed by another group of document reviewers who manually review many documents to find potentially relevant ones. Documents found during document review are used during litigation proceedings, such as depositions. Compliance reviewers continuously manually review internal documents to identify and flag risks.
[0019] Therefore, it would be advantageous to offer solutions that will overcome the aforementioned challenges. [Overview of the project]
[0020] A summary of some exemplary embodiments of this disclosure is provided below. This summary is provided for the convenience of the reader to provide a basic understanding of such embodiments and does not fully define the scope of this disclosure. This summary is not intended to be a comprehensive overview of all conceivable embodiments, nor to identify any key or definitive elements of all embodiments, nor to definitively describe the scope of any or all embodiments. Its sole purpose is to present some concepts of one or more embodiments in a simplified form as a prelude to the more detailed descriptions that will be presented later. For convenience, “some embodiments” or “specific embodiments” may be used herein to refer to a single or more embodiments of this disclosure.
[0021] Certain embodiments disclosed herein include a method for training a machine learning model using natural language interactions. The method includes applying a language model to the text of a set of natural language interactions to output a set of domain-specific language (DSL) data, the set of natural language interactions being between a user and at least one other entity, the set of natural language interactions representing at least one user-defined concept, querying a knowledge base based on the set of DSL data to obtain at least one DSL query result, integrating at least one DSL query result with a structured representation of the natural language interactions to create at least one contextualized DSL query result, and training a language model using at least one contextualized DSL query result.
[0022] Certain embodiments disclosed herein include a stored non-temporary computer-readable medium that causes a processing circuit unit to execute a process, the process being to apply a language model to the text of a set of natural language interactions in order to output a set of domain-specific language (DSL) data, the set of natural language interactions being between a user and at least one other entity, the set of natural language interactions representing at least one user-defined concept, querying a knowledge base based on the set of DSL data in order to obtain at least one DSL query result, integrating at least one DSL query result with a structured representation of the natural language interactions in order to create at least one contextualized DSL query result, and training a language model using at least one contextualized DSL query result.
[0023] Certain embodiments disclosed herein also include a system for training a machine learning model using natural language interactions. The system includes a processing circuit and a memory which, when executed by the processing circuit, applies a language model to the text of a set of natural language interactions in order to output a set of domain-specific language (DSL) data, the set of natural language interactions being between a user and at least one other entity, the set of natural language interactions representing at least one user-defined concept, query a knowledge base based on the set of DSL data to obtain at least one DSL query result, integrate at least one DSL query result with a structured representation of the natural language interaction in order to create at least one contextualized DSL query result, and train a language model using at least one contextualized DSL query result.
[0024] Certain embodiments disclosed herein include, or are configured to perform, one or more of the following steps, the method, non - transient computer - readable medium or system as described above. The steps are to create a structured representation of a set of natural - language interactions, the structured representation including a set of fields and corresponding values, and the values of the structured representation including values expressed in natural - language interactions.
[0025] Certain embodiments disclosed herein include, or are configured to perform, one or more of the following steps, the method, non - transient computer - readable medium or system as described above. The steps are to create a knowledge base based on a dataset associated with a set of files, and the created knowledge base represents a plurality of entities and a plurality of relationships among the entities of the plurality of entities shown in the set of files.
[0026] Certain embodiments disclosed herein include the method, non - transient computer - readable medium, or system as described above, where at least one contextualized DSL query result indicates at least one of a plurality of entities.
[0027] Certain embodiments disclosed herein include, or are configured to perform, one or more of the following steps, the method, non - transient computer - readable medium or system as described above. The steps are to update a knowledge base according to at least a portion of a natural - language interaction, and the step of updating the knowledge base includes adding a representation of a new entity shown in that portion of the natural - language interaction.
[0028] Certain embodiments disclosed herein include the method, non - transient computer - readable medium, or system as described above, where the new entity added to the knowledge base is represented in a domain - specific language of the entity.
[0029] Certain embodiments disclosed herein include the methods, non-transitory computer-readable media or systems described above, where natural language interaction exhibits at least one specification for at least one user-defined concept.
[0030] Certain embodiments disclosed herein include the methods, non-transitory computer-readable media or systems described above, where natural language interaction includes at least one natural language query.
[0031] Certain embodiments disclosed herein further include or are configured to perform one or more of the following steps: applying a trained language model to at least one electronic document to identify at least one of at least one user-defined concept within the at least one electronic document.
[0032] Certain embodiments disclosed herein further include the system described above, which includes a knowledge base (KB) that stores entities, relationships, and concepts, a natural language processing (NLP) component configured to translate at least one natural language query into at least one DSL query, where the natural language processing component includes a natural language processing model, a structured domain-specific language (DSL) component configured to query the knowledge base using at least one DSL query, and a composition component configured to integrate the query results of at least one knowledge base into a structured representation of natural language interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The subject matter disclosed herein is specifically identified and explicitly claimed in the last claim herein. The aforementioned and other purposes, features and advantages of the disclosed embodiments will become apparent from reading the following description in conjunction with the accompanying drawings. [Figure 1] This is a network diagram used to illustrate the various embodiments disclosed. [Figure 2] This is a flowchart illustrating a method for training a machine learning model using natural language interaction, according to one embodiment. [Figure 3] This is a schematic diagram of a trainer according to one embodiment. [Figure 4] This is an illustrative diagram of a legal workflow. [Figure 5] This is an illustrative diagram of a workflow that may be realized by various embodiments disclosed. [Figure 6] This is a diagram illustrating a system flow. [Figure 7] This table is used to describe certain use cases related to mixed concept or strategic education that may be realized by the various embodiments disclosed. [Figure 8] This is an illustrative diagram illustrating the data flow for the process of training a machine learning model using natural language interaction. [Figure 9] This is an illustrative diagram illustrating the creation of a structured representation. [Modes for carrying out the invention]
[0034] Given the challenges described above, solutions that enable improvements to existing workflows are desirable. Automated solutions for improving such workflows have been identified as particularly desirable. However, existing manual solutions cannot be automated simply by having a computer perform the same process. Therefore, the disclosed embodiments provide solutions that enable the effective automation of these workflows by leveraging machine learning models in processes that overcome various technical challenges toward achieving this automation. More specifically, the various disclosed embodiments enable the training of machine learning models to recognize concepts within data, such as multimedia documents, using natural language interactions. These natural language interactions will be able to bridge the gap between human education and machine learning without requiring process management by data scientists.
[0035] The various embodiments disclosed include methods and systems for training machine learning models using natural language interaction, as well as techniques that also utilize machine learning trained using natural language interaction. The disclosed embodiments may be used to train machine learning models to identify concepts taught by a user via natural language interaction with respect to electronic documents, such as (but not limited to) documents containing multimedia content. Such documents may be based on or not based on digitally created documents (e.g., electronic documents created with a word processor or image processing program) based on user input, or they may represent real documents, for example, by including scanned images of real documents in which text may be identified using an optical character recognition (OCR) system.
[0036] One of the benefits of at least some embodiments disclosed is to enable untrained users to interact seamlessly with the integrated system. In one embodiment, the system allows users to (1) educate the system on a wide range of custom concepts applicable to different use cases (e.g., compliance, governance, litigation, investigation, and general corporate data management use cases), (2) query the knowledge, including concepts previously educated by other users of the system, and (3) interact in natural language, enabling users to educate and query knowledge through mixed-intuitive natural language dialogue, as well as communicate tasks relating to the system itself, such as metaknowledge and its performance. In some embodiments, queries may include compound queries such as searching for documents tagged with a certain concept, or documents tagged with a certain concept, constrained by other criteria such as their content.
[0037] In one embodiment, the system is designed to communicate naturally and directly with different users throughout all stages of the process. This supports mixed-input interaction, allowing users to query the system in natural language. The system can also initiate interaction by asking questions or presenting new information. The types of natural language interactions the system may support include, but are not limited to, (1) teaching new knowledge, (2) querying knowledge, and (3) meta-level interaction.
[0038] Teaching new knowledge may include, but is not limited to, receiving user input (such as input specifications) in the form of natural language descriptions or instructions, which can be used to understand specific issues or specifications. The system may be configured to interact with the user, clarifying information throughout the natural language dialogue, for example, by asking questions and presenting the information it finds. The concepts being taught may include, but are not limited to, document-level, intra-document, and inter-document concepts. An unrestricted embodiment of the system flow is further described below with reference to Figure 6.
[0039] The concept of a document level may include, but is not limited to, classification categories that apply to the entire document. In a non-restrictive legal embodiment, such classification categories may include categories indicating that a given document is privileged or responsive to a particular issue within a specification document.
[0040] Document concepts may include, but are not limited to, concepts at the subdocument level, such as specific facts or information explicitly stated within a document. In a non-restrictive legal embodiment, such document concepts may include certain types of personal information (PH) contained within a document, or certain types of information, such as parameters of contracts, invoices, and other types of documents.
[0041] Interdocument concepts may include, but are not limited to, broad concepts that include elements that are not strongly linked to a document but are substantiated by or within a document. As a non-restrictive example of a legal embodiment, a list of lawyers or persons involved in a particular transaction may be used as an interdocument concept, even though not all of them may have been known prior to the case, but their names, titles, and roles are mentioned in numerous documents that serve as evidence. A system can learn to extract such facts. Such interdocument concepts may be expressed between pairs of documents or within a given document.
[0042] A broad concept may be expressed across many documents (e.g., three or more), or multiple types of documents (e.g., one or more Type 1 documents and one or more Type 2 documents). In an unrestrictive embodiment, a broad concept may be a fundraising event supported by various documents such as term sheets and closing sets.
[0043] With regard to querying knowledge, in one embodiment, a user can query both newly learned and existing knowledge within data (e.g., case data) by natural language input in the form of a question. With regard to meta-level interactions, the system may be configured to handle interactions related to issues surrounding the process itself, such as querying the system's predicted performance, querying the status of data within a workflow pipeline (e.g., a litigation pipeline), and defining annotation instructions.
[0044] In some embodiments, the system may include, but is not limited to, the following components: (1) a knowledge base (KB); (2) a structured domain-specific language (DSL) component configured to query the knowledge in the knowledge base using domain-specific language queries and update the knowledge base using data in the domain-specific language; (3) a natural language processor or other natural language processing (NLP) component configured to translate natural language queries and conversational contexts into the domain-specific language by applying a trained natural language processing model to the text of a set of natural language interactions; and (4) a synthesis component configured to integrate the results of performing domain-specific language operations on the knowledge base into contextual data in the form of structured representations of natural language interactions used to train a natural language processing model for generating natural language responses to user natural language queries.
[0045] In some embodiments, the system operates sequentially in a series of stages. In another embodiment, in the first stage, initial data is imported into the system. The initial data may include, but is not limited to, text data, image data, video data, audio data, metadata, or a combination thereof, associated with files. In the second stage, the system transfers data to build an internal knowledge base (KB) of base entities, including people, places, organizations, and relationships such as employment, titles, and positions. However, the system is not limited to extracting only this set of entity and relationship types. In later stages, new entities and relationships may be taught by the user. In the third stage, after processing the initial data, the system becomes active and supports asynchronous natural language interaction from the user.
[0046] Natural language interactions may include, but are not limited to, (1) specifications of concepts to be taught, (2) natural language queries, both, and similar elements. Specification natural language interactions include interactions in which users can submit specifications through a conversational chat interface or by submitting documents containing specifications, such as document review protocols or internal rulebooks. During natural language query interactions, users can submit natural language queries in the form of questions. These questions may include, or relate to, data, document queries, conceptual queries, and meta queries.
[0047] Examples of different types of concepts that users may be taught by the various embodiments disclosed include (1) document-level concepts, (2) inter-document-level concepts, and (3) intra-document-level concepts. Document-level concepts teach that a certain document is an example of a given concept, and these concepts are based on the content of a single document. Inter-document-level (scope) concepts teach that a certain scope or reference in a certain text is an example of a given concept (e.g., PH), and these concepts are based on the content of a given scope / reference. Intra-document-level concepts teach concepts that do not directly rely on a single scope or reference but may rely on multiple documents or texts.
[0048] The disclosed embodiments may be used to train machine learning models for embodiments such as (but not limited to) legal processes. As an unrestricted embodiment, the disclosed embodiments may be used to enable users to seamlessly interact with an integrated system at various stages of legal processes such as (but not limited to) compliance, governance, litigation, investigation, and general corporate data management.
[0049] In an unrestrictive, exemplary embodiment, the system acts as a hub throughout all stages of the legal process, from governance / compliance to litigation. In one embodiment, each stage of the process queries knowledge from a centralized, integrated system that provides and expands upon knowledge, making that knowledge accessible to the entire group of experts throughout the litigation process.
[0050] It should be noted that the disclosed embodiments are not necessarily limited to legal embodiments, and that the disclosed embodiments may be similarly used in other embodiments without departing from the scope of the disclosure. The various embodiments disclosed are described for illustrative purposes in terms of legal process embodiments, but are not limited to at least some of the disclosed embodiments.
[0051] Figure 1 shows an exemplary network diagram 100 used to illustrate various embodiments disclosed. In this embodiment, the network diagram 100, a user device 120, a trainer 130, and several databases 140-1 to 140-N (hereinafter, for the sake of brevity, individually referred to as database 140 and collectively as the database group 140) communicate via network 110. Network 110 may be, but is not limited to, a wireless network, a cellular or wired network, a local area network (LAN), a wide area network (WAN), a metro area network (MAN), the Internet, the World Wide Web (WWW), similar networks, and any combination thereof.
[0052] The user device (UD) 120 may be a personal computer, laptop, tablet computer, smartphone, wearable computing device, or any other device capable of receiving natural language input, such as but not limited to text, voice, and the like. For this purpose, the user device 120 may include, or communicate with, one or more input devices, such as but not limited to, a keypad, mouse, touchscreen, microphone, combination thereof, and the like, configured to capture natural language input from the user.
[0053] The trainer 130 is configured to train machine learning models according to various embodiments disclosed, and in some embodiments may be further configured to utilize the machine learning models.
[0054] The database group 140 stores data used to train machine learning models, including, but not limited to, exemplary documents, user-instructed concepts, combinations thereof, and similar items. In some embodiments, database 140 may further include data analyzed by trainer 130 using the machine learning models trained as described herein.
[0055] While Figure 1 illustrates an exemplary network diagram 100, it should be noted that the disclosed embodiments are not limited to the network environment 100 illustrated in Figure 1.
[0056] Here, we describe an exemplary process that may be performed with the setup illustrated in network diagram 100. The system (e.g., trainer 130 in Figure 1) is educated on a wide range of custom concepts applicable to different use cases such as compliance, governance, litigation, investigation, and general corporate data management (but not limited to these). Queries are made to the knowledge, including concepts previously educated by other users of the system. Natural language interactions are received, enabling users to educate and query knowledge through mixed-initiated natural language dialogue, as well as communicate tasks related to meta-knowledge and the system itself (such as its performance). A trained machine learning model is applied as part of the education.
[0057] Figure 2 is an exemplary flowchart illustrating a method for training a machine learning model using natural language interaction according to one embodiment. In one embodiment, the method is performed by the trainer 130 in Figure 1.
[0058] In S210, a knowledge base is created based on the dataset associated with the set of files. In one embodiment, the knowledge base represents multiple entities and multiple relationships between entities among the multiple entities shown across the set of files.
[0059] In S220, a language model is applied to a set of natural language interactions to output a set of domain-specific language (DSL) data. In one embodiment, the set of natural language interactions is between a user and at least one other entity (e.g., an artificial intelligence assistant). The set of natural language interactions represents at least one user-defined concept, i.e., the user communicates a concept that they want to teach the system during the natural language interaction.
[0060] In one embodiment, the natural language interaction represents at least one specification for at least one user-defined concept. The specification's natural language interaction includes an interaction that allows the user to submit the specification either through a conversational chat interface or by submitting a document containing the specification, such as a document review protocol or an internal rulebook.
[0061] In another embodiment, a natural language interaction includes at least one natural language query. During a natural language query interaction, the user can submit natural language queries in the form of questions. These questions may include, or relate to, facts within data, document queries, conceptual queries, and meta queries.
[0062] In S230, a structured representation of the set of natural language interactions is created. In one embodiment, the structured representation includes a set of fields and corresponding values, the values of the structured representation include values represented in the natural language interactions.
[0063] In S240, a knowledge base is queried based on a set of DSL data to obtain at least one DSL query result. In one embodiment, querying the knowledge base involves translating at least one natural language query into at least one domain-specific language (DSL) query formatted according to the domain-specific language of the knowledge base. In one embodiment, the at least one query result represents at least one of a plurality of entities.
[0064] In S250, at least one DSL query result is integrated with a structured representation of a natural language interaction in order to create at least one contextualized DSL query result.
[0065] In S260, the language model is trained using at least one contextualized DSL query result. In embodiments where the natural language interaction includes one or more references to electronic documents, the language model is trained using the contextualized DSL query result and the electronic document. In this regard, the language model may be improved to recognize user-defined concepts expressed in the natural language interaction within the electronic document.
[0066] In S270, a trained language model is applied to at least one electronic document to identify at least one user-defined concept within at least one electronic document.
[0067] In S280, the knowledge base is updated according to at least a portion of the natural language interaction. In one embodiment, updating the knowledge base includes adding representations of new entities shown in that portion of the natural language interaction. In another embodiment, the representations of new entities added to the knowledge base are expressed in a domain-specific language.
[0068] Figure 3 is an exemplary schematic diagram of a trainer 130 according to one embodiment. The trainer 130 includes a memory 320, a storage device 330, and a processing circuit unit 310 connected to a network interface 340. In one embodiment, the components of the trainer 130 may be connected communicably via a bus 350.
[0069] The processing circuit section 310 may be implemented as one or more hardware logic components and circuits. Examples of hardware logic components that may be used, as an example and not an limitation, include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), graphics processing units (GPUs), tensor processing units (TPUs), general-purpose microprocessors, microcontrollers, digital signal processors (DSPs), and similar devices, or any other hardware logic components capable of performing calculations or other operations on information.
[0070] Memory 320 may be volatile (e.g., random access memory), non-volatile (e.g., read-only memory, flash memory), or a combination thereof.
[0071] In one configuration, software for carrying out one or more embodiments disclosed herein may be stored in the storage device 330. In another configuration, memory 320 may be configured to store such software. Software should be broadly interpreted to mean any type of instruction, whether called software, firmware, middleware, microcode, hardware description language, or otherwise. Instructions may include code (e.g., in source code format, binary code format, executable code format, or any other suitable code format). When executed by the processing circuit unit 310, the instructions cause the processing circuit unit 310 to perform the various processes described herein.
[0072] The storage device 330 may be a magnetic storage device, an optical storage device, or similar, and may be implemented as, for example, flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital general-purpose disc (DVD), or any other medium that can be used to store desired information.
[0073] The network interface 340 allows the trainer 130 to communicate with, for example, user devices 120, database groups 140, and similar devices.
[0074] It should be understood that the embodiments described herein are not limited to the specific architecture illustrated in Figure 3, and other architectures may also be used without departing from the scope of the disclosed embodiments.
[0075] The disclosed embodiments may be used to implement new workflows, such as the workflow illustrated in Illustration 500 of Figure 5. In contrast to a linear workflow, the workflow illustrated in Figure 5 demonstrates how the system can integrate user input across various parts of a non-restrictive illustrative use case, such as the legal process represented in Illustration 500.
[0076] The various embodiments disclosed enable a natural blend of multiple concepts and educational strategies, allowing users to effectively address a wide range of practical use cases. For example, users may combine previously taught concepts and educational strategies, such as by providing both explanations and examples.
[0077] Figure 6 illustrates Figure 600, an unrestricted example of a system flow illustrating a cycle of updating specifications and queries in which new specifications update system knowledge, enabling updates to specifications when new queries are received from users. As illustrated in Figure 6, new specifications in the set of natural language (NL) specifications 610 may be used to update system knowledge, which may further lead to new natural language queries 620 from users, which may further lead to specification updates.
[0078] Figure 7 is Table 700 of unrestricted examples illustrating several exemplary user-led and system-led teaching methods, including (but not limited to) instructions, examples, and explanations.
[0079] The disclosed embodiments may be used to facilitate the effective training of machine learning models through natural language interaction and to enable users to effectively train them. More specifically, user input in the form of natural language samples, such as text, speech, and similar, is used as training data to train a machine learning model. Natural language interaction may be aided by one or more Large Language Models (LLMs) to allow the user to interact with a trainer system that trains the model, enabling the collection of training data in an intuitive process without the need for explicit programming, and may achieve similar performance with a smaller amount of training data compared to at least some existing solutions.
[0080] Figure 8 is an exemplary Figure 800 illustrating the data flow for the process of training a machine learning model using natural language interaction, illustrating an exemplary natural language interaction between a user and an artificial intelligence system.
[0081] As illustrated in Figure 8, the natural language interaction session 810 is input as a set of text representing the natural language interaction between a natural language processing component and a domain-specific language (DSL) component, in order to translate text 810 from natural language into domain-specific language data (820) and represent at least a portion of the natural language interaction in domain-specific language format. In the exemplary embodiment illustrated in Figure 8, the natural language processing component is a large-scale language model (LLM). Using the domain-specific language data, a query is generated to query a knowledge base using the DSL interaction data to produce a DSL query result (830). The DSL query result is inserted into the data structure of the natural language interaction, thereby creating a structured representation of at least a portion of the natural language interaction.
[0082] The knowledge base query results 840 are synthesized in the context synthesis process 850. The context synthesis process 850 integrates the results 840, obtained by querying the knowledge base using domain-specific language, with a structured representation of at least some of the natural language interaction. In the exemplary embodiment illustrated in Figure 8, the context synthesis process 850 is an LLM context synthesis process.
[0083] The results of context synthesis 850 may be used to train a language model (860) such as an LLM for natural language processing components (e.g., a generative pre-trained transformer, or a GPT model). The trained language model may be queried, the query results of the language model may be used to update a knowledge base, provide a response 870 to a user query, or both.
[0084] Figure 9 is an exemplary figure illustrating the creation of structured representations through such natural language interaction.
[0085] It is important to note that the embodiments disclosed herein are merely examples of many advantageous uses of the innovative teachings herein. In general, the descriptions set forth in this specification are not necessarily limited to any of the various embodiments claimed. Furthermore, some descriptions may apply to some inventive features but not to others. In general, unless otherwise indicated, singular elements may be plural and vice versa without loss of generality. In the drawings, similar numbers refer to the same part through multiple drawings.
[0086] The various embodiments disclosed herein can be implemented as hardware, firmware, software, or any combination thereof. Software may also be implemented as an application program tangibly embodied in a program storage device or computer-readable medium consisting of components, or certain devices and / or combinations of devices. The application program may be uploaded to and executed by a machine having any suitable architecture. Preferably, the machine is implemented on a computer platform having hardware (such as one or more central processing units ("CPU"), memory, and input / output interfaces). The computer platform may also include an operating system and microinstruction code. The various processes and functions described herein may be any part of microinstruction code, or part of an application program, or any combination thereof, executed by a CPU, whether such a computer or processor is explicitly indicated. In addition, various other peripheral devices, such as additional data storage devices and printing devices, may be connected to the computer platform. Furthermore, non-transient computer-readable medium is any computer-readable medium except for transient propagating signals.
[0087] All embodiments and conditional language described herein are intended for educational purposes to help the reader understand the principles of the disclosed embodiments and the inventors' concepts for the advancement of the art, and should be construed as not being limited to the embodiments and conditions described herein. Furthermore, all descriptions herein describing the principles, aspects and embodiments of the disclosed embodiments, as well as specific examples thereof, are intended to encompass both their structural and functional equivalents. In addition, such equivalents include both currently known equivalents and those to be developed in the future, i.e., any elements developed to perform the same function, regardless of their structure.
[0088] It should be understood that references to elements using markings such as “First,” “Second,” etc., in this specification do not generally limit the number or order of those elements. Rather, these markings are commonly used in this specification as a convenient way to distinguish between two or more elements or between two or more instances of an element. Therefore, references to the First and Second elements do not mean that only two elements can be adopted, or that the First element must somehow precede the Second element. Also, unless otherwise stated, a set of elements includes one or more elements.
[0089] As used herein, the phrase “at least one of” following an enumeration of items means that any of the enumerated items may be used individually, or any combination of two or more of the enumerated items may be used. For example, if it is stated that a system includes “at least one of A, B, and C,” the system may include A alone, B alone, C alone, 2A, 2B, 2C, 3A, a combination of A and B, a combination of B and C, a combination of A and C, a combination of A, B, and C, a combination of 2A and C, a combination of A, 3B, and 2C, and so on.
Claims
1. A method for training machine learning models using natural language interaction, Applying a language model to the text of a set of natural language interactions in order to output a set of domain-specific language (DSL) data, wherein the set of natural language interactions is between a user and at least one other entity, and the set of natural language interactions represents at least one user-defined concept. To obtain at least one DSL query result, the knowledge base is queried based on the set of DSL data, To create at least one contextualized DSL query result, the at least one DSL query result is integrated with the structured representation of the natural language interaction, A method comprising training the language model using the at least one contextualized DSL query result.
2. The method according to claim 1, further comprising creating a structured representation of the set of natural language interactions, wherein the structured representation includes a set of fields and corresponding values, the values of the structured representation include values represented in the natural language interactions.
3. The method according to claim 1, further comprising creating the knowledge base based on a dataset associated with a set of files, wherein the created knowledge base shows a plurality of entities and a plurality of relationships between the entities among the plurality of entities shown between the set of files.
4. The method according to claim 3, wherein the at least one contextualized DSL query result indicates at least one of a plurality of entities.
5. The method according to claim 1, further comprising updating the knowledge base based on at least a portion of the natural language interaction, wherein updating the knowledge base includes adding new representations of the entities shown in the portion of the natural language interaction.
6. The method according to claim 5, wherein the representation of the new entity added to the knowledge base is expressed in a domain-specific language.
7. The method according to claim 1, wherein the natural language interaction provides at least one specification for the at least one user-defined concept.
8. The method according to claim 1, wherein the natural language interaction includes at least one natural language query.
9. The method according to claim 1, further comprising applying the trained language model to at least one electronic document in order to identify at least one of the at least one user-defined concept within at least one electronic document.
10. A non-temporary computer-readable medium storing instructions for executing a process in a processing circuit unit, wherein the process is: Applying a language model to the text of a set of natural language interactions in order to output a set of domain-specific language (DSL) data, wherein the set of natural language interactions is between a user and at least one other entity, and the set of natural language interactions represents at least one user-defined concept. To obtain at least one DSL query result, the knowledge base is queried based on the set of DSL data, To create at least one contextualized DSL query result, the at least one DSL query result is integrated with the structured representation of the natural language interaction, A non-temporal computer-readable medium comprising training the language model using the at least one contextualized DSL query result.
11. A system for training machine learning models using natural language interaction, Processing circuit section, It includes memory, and the memory is used when executed by the processing circuit unit. Applying a language model to the text of a set of natural language interactions in order to output a set of domain-specific language (DSL) data, wherein the set of natural language interactions is between a user and at least one other entity, and the set of natural language interactions represents at least one user-defined concept. To obtain at least one DSL query result, the knowledge base is queried based on the set of DSL data, To create at least one contextualized DSL query result, the at least one DSL query result is integrated with the structured representation of the natural language interaction, A system comprising instructions that configure the system to train the language model using the at least one contextualized DSL query result.
12. The aforementioned system, The system according to claim 11, further configured to create a structured representation of the set of natural language interactions, wherein the structured representation includes a set of fields and corresponding values, the values of the structured representation include values expressed in the natural language interactions.
13. The aforementioned system, The system according to claim 11, further configured to create the knowledge base based on a dataset associated with a set of files, wherein the created knowledge base represents a plurality of entities and a plurality of relationships between the entities among the plurality of entities shown between the set of files.
14. The system according to claim 13, wherein the at least one contextualized DSL query result indicates at least one of a plurality of entities.
15. The aforementioned system, The system according to claim 11, further configured to update the knowledge base based on at least a portion of the natural language interaction, wherein updating the knowledge base includes adding new representations of the entities shown in the portion of the natural language interaction.
16. The system according to claim 15, wherein the representation of the new entity added to the knowledge base is expressed in a domain-specific language.
17. The system according to claim 11, wherein the natural language interaction indicates at least one specification for at least one user-defined concept.
18. The system according to claim 11, wherein the natural language interaction includes at least one natural language query.
19. The aforementioned system, The system according to claim 11, further configured to apply the trained language model to at least one electronic document in order to identify at least one of the at least one user-defined concept within at least one electronic document.
20. The aforementioned system, The aforementioned entities, relationships, and concepts are stored in the aforementioned knowledge base (KB), A natural language processing (NLP) component configured to translate at least one natural language query into at least one DSL query, wherein the natural language processing component includes the natural language processing model, A structured domain-specific language (DSL) component configured to query the knowledge base using at least one of the DSL queries, The system according to claim 11, further comprising a synthetic component configured to integrate the queries of at least one knowledge base with the structured representation of the natural language interaction.