Conversation search with efficacy engine
The chatbot system with a two-pass categorization process and prompt optimization addresses the inefficiencies of current data search methods, achieving faster, more affordable, and transparent data retrieval in large email archives.
Patent Information
- Application Number
- PCT/US2025/032278
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-27
- Filing Date
- 2025-06-04
- Publication Date
- 2025-12-11
AI Technical Summary
Current methods for finding relevant data in large email archives are time-consuming, costly, and produce large amounts of irrelevant data, while predictive coding models are capital-intensive and require extensive training before providing effective results, lacking traceability.
A chatbot system utilizing a natural language interface and a two-pass categorization process, including a first pass for filtering irrelevant data and a second pass for predicting relevance, with optional prompt optimization by an efficacy engine, to streamline data search and provide rationale for categorization.
The system significantly reduces the time and cost of data retrieval, enhances traceability, and provides faster, more accurate results with reduced reliance on expert intervention, enabling quicker case resolution and potential cost savings.
Smart Images

Figure US2025032278_11122025_PF_FP_ABST
Abstract
Description
CONVERSATION SEARCH WITH EFFICACY ENGINECROSS-REFERENCE TO RELATED APPLICATION(S)This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 700,121 entitled CONVERSATION SEARCH WITH EFFICACY ENGINE filed September 27, 2024 and also claims the benefit of U.S. Provisional Patent Application No. 63 / 656,338 entitled CONVERSATION SEARCH filed June 5, 2024, each of which is hereby incorporated herein by reference in its entirety.FIELD OF THE INVENTIONThe invention generally relates to an integrated chatbot for performing email and other conversation search functions using natural language processing.BACKGROUND OF THE INVENTIONThere are many applications that require the ability to find relevant data within a large email archive in a quick, affordable, and efficient manner. For example, one might need to find emails or other messages or data relating to a specific topic such as in response to an e-discovery request, subpoena, internal investigation, etc. Without limitation, some examples include an HR officer searching for evidence of abusive behavior, an investigator searching for evidence of fraud, a compliance office searching for evidence of insider trading, etc. Most times, these searches are time sensitive, where getting an answer quickly is critical.Currently, finding relevant data within a large email or other archive is generally done through searches (e.g., mainly keyword / boolean searches) along with manual review, which can be time and cost intensive. Typically, with these types of searches, the search results contain large amounts of irrelevant data (‘noise’). This is a problem, as it slows down the process and increases costs (especially if using third party / counsel and tools).There are also predictive coding models, although such models are capital intensive and take time to fine-tune before they provide good results (e.g., requiring supervised / reinforcement learning to train these models). Additionally, they do not provide rationale for their output (e.g., “traceability” information as to how and why particular data was included in or excluded from search results) other than perhaps providing some type of confidence score (e.g., data X is included with Y% confidence).The following are some potentially relevant references: India Patent Application No. 1492 / CHENP / 2009 (METHOD, APPARATUS AND SYSTEM FOR SEARCHING EMAILS);European Patent Application No. 08005505 (Apparatus, method and computer program product for processing email, and apparatus for searching email); and United States Patent Application No. 12051444 (Apparatus, method and computer program product for processing email, and apparatus for searching email).SUMMARY OF VARIOUS EMBODIMENTSIn accordance with one embodiment of the invention, a conversation search system and method utilize a chatbot comprising a natural language user interface through which a user can enter natural language prompts for a conversation search of a reviewset containing data items for review, a plurality of scripts including at least a first pass script configured to characterize and label the data items of the reviewset to filter irrelevant data items out of the reviewset and a second pass script configured to predict relevance of each data item of the reviewset remaining after the filtering of the first pass script, and an orchestrator that selects and triggers execution of the scripts based on the natural language prompts.In various alternative embodiments, the data items may include at least one of messages or documents. The natural language user interface may use a large language model (LLM) to process user prompts. The natural language user interface may allow natural language prompts to be entered textually and / or verbally. The first pass script may use natural language processing (NLP) to characterize and label the data items. The second pass script may natural language processing to predict relevance of each data item, e.g., uses a Bert model such as a fraud detection model, a harassment model, or a compliance model.Additionally or alternatively, the chatbot may include an efficacy engine that automatically generates one or more additional prompts for a given user prompt. The number of additional prompts may fixed or user configurable. At least one prompts may be a variant of the given user prompt. The chatbot may be configured to use prompt engineering techniques to automatically generate the one or more additional prompts. The chatbot may execute the one or more additional prompts in addition to or in lieu of the given user prompt. The chatbot may execute a plurality of prompts from among the given user prompt and the one or more additional prompts and identifies a best prompt from among the plurality of executed prompts. The chatbot may identify the best prompt by comparing an output of each executed prompt to a baseline or by comparing outputs of the plurality of executed prompts.In yet further embodiments, the plurality of scripts may include an interrogation script to interrogate a specified data item, a report generation script, and / or a visualization script.Additional embodiments may be disclosed and claimed.BRIEF DESCRIPTION OF THE DRAWINGSThose skilled in the art should more fully appreciate advantages of various embodiments of the invention from the following “Description of Illustrative Embodiments,” discussed with reference to the drawings summarized immediately below.FIG. 1 is a schematic diagram of a conversation search system, in accordance with certain embodiments.FIG. 2 is a schematic diagram of a chatbot user interface, in accordance with certain embodiments.It should be noted that the foregoing figures and the elements depicted therein are not necessarily drawn to consistent scale or to any scale. Unless the context otherwise suggests, like elements are indicated by like numerals. The drawings are primarily for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein.DESCRIPTION OF ILLUSTRATIVE EMBODIMENTSEmbodiments implement a so-called “conversation search” that attempts to speed up search processes by performing a two-pass categorization and interrogation of data such as emails or other messages or data in one integrated chatbot with report and visualization capabilities. In certain embodiments, the first pass characterizes and labels the type of data, e.g., so that irrelevant data can be filtered out (e.g., newsletters, automated emails, etc.). Then, the second pass predicts if particular data is relevant to the search criteria. Finally, the chatbot provides interrogation capabilities to address specific questions around particular data such as an email and / or attachment. Among other things, embodiments are expected to reduce the time needed to find relevant data (which often can be characterized as a ‘needle in a haystack’) as well as being faster and more affordable than having users trawl through the data themselves. Some exemplary embodiments are described herein with references to categorization and interrogation of email messages and related data, although it should be noted that embodiments can utilizes other types of data (e.g., other types of messages) in addition to, or in lieu of, email messages.FIG. 1 is a schematic diagram of a conversation search system, in accordance with certain embodiments. Among other things, the conversation search system includes a chatbot, a large language model (FEM) agent (referred to herein as the “orchestrator”), a number of backend scripts, and a database system (DB - not explicitly shown in FIG. 1), e.g., a Dynamo DBsupported by a Kubemetes task. Each script is configured to perform a specific type of function relating to conversation search. Some example scripts are described below.In practice, a user generally would perform a preliminary search to produce a set of emails for review (referred to herein as a “reviewset”). For example, the user may search for emails sent and / or received by relevant users (e.g., users involved in a particular case or inquiry), which can produce search results containing an enormous volume of emails (e.g., 100,000+ emails) that now need to be culled down to find the relevant data. The reviewset can be stored or loaded in the DB for use by the conversation search scripts. This might involve, for example, obtaining and saving a list of message IDs for the emails in the reviewset and obtaining and saving message bodies for the emails in the reviewset.In order to perform conversation search operations on the reviewset, the user enters a prompt into the chatbot, e.g., through a user interface as depicted schematically in FIG. 2. Certain embodiments include a natural language (NF) interface through which the user can enter prompts (e.g., typed or spoken) that generally are converted to text or other format usable by the system. The user prompt is saved to the DB and is then provided to the orchestrator, which determines which type of script to run, saves the determination to the DB (e.g., for tracking purposes), and triggers execution of the determined script, e.g., by passing the user prompt to the selected script (which, in some embodiments, may be conveyed in JSON format). As depicted schematically in FIG. 1, in certain embodiments, the conversation search system includes the following types of scripts:Email Categorization - uses natural language processing (NEP) to process all of the emails in the reviewset to categorize and label each email including a rationale for the categorization / label (i.e., indicating why the script categorized and labeled each email);Auto Review- user initiated prompt run against all data and returned as labels / comments with rationale;Model categorization - select a ready-to-use ME Bert model (e.g., a fraud detection model, a harassment model, a compliance model, etc.) to make a determination against all data as labels / comments with score;Email Interrogation - user initiated prompt run against messagelD specified and returned in chatbot UI;Report Generation - Retrieval- Augmented Generation (RAG) gen Al generated report based on user prompt; andVisualization - Retrieval-Augmented Generation (RAG) gen Al generated visualization based on user prompt - in certain embodiments, this is returned as python script can be rendered in chatbot UI.Example 1: Email Categorization. The email categorization script might be triggered if the user enters a prompt such as “Please review all emails in the reviewset and label if the email is a newsletter, automated email, automated notification, company message, or external sales pitch, and please also provide your reasoning for each categorization label." As discussed above, the prompt is saved to the DB and then passed to the orchestrator. The orchestrator determines the relevant script for the prompt (in this case, the email categorization script), saves the determination to the DB, and triggers execution of the script. In certain embodiments, the email categorization script pulls every email body from the DB, processes each email body to categorize the email, labels the email, and saves the categorization and the rationale for the categorization to the DB. The chatbot is notified when the script execution is completed.Example 2: Model Categorization. The model categorization script might be triggered if the user enters a prompt such as “Please review all the emails and apply the XXXXXX model," where “XXXXXX” would indicate a specific model (e.g., fraud detection model, harassment model, compliance model, etc.). As discussed above, the prompt is saved to the DB and then passed to the orchestrator. The orchestrator determines the relevant script for the prompt (in this case, the model categorization script), saves the determination to the DB, and triggers execution of the script. In certain embodiments, the model categorization script pulls every email body from the DB, processes each email body to categorize the email, ingests the data into the relevant NLP BERT model, and saves the outcomes to the DB (e.g., scores between 0-1). All emails with scores above a predetermined threshold (e.g., 0.85) may be flagged (e.g., labeled) as relevant and the score may be posted as a comment in the reviewset. The chatbot is notified when the script execution is completed.Example 3: Auto Review. The auto review script might be triggered if the user enters a prompt such as “I’m looking for evidence of fraud, apply label if relevant and give reasoning as a comment." As discussed above, the prompt is saved to the DB and then passed to the orchestrator. The orchestrator determines the relevant script for the prompt (in this case, the auto review script), saves the determination to the DB, and triggers execution of the script. In certain embodiments, the auto review script pulls every email body with the user provided prompt from the DB to obtain the categorizations and rationales, processes each email body to categorize theemails, saves the results to the DB, adds labels and comments to the reviewset, and notifies the chatbot when completed.Example 4: Email Interrogation. The email interrogation script might be triggered if the user enters a prompt such as “Summarize the email with Message ID: 101." As discussed above, the prompt is saved to the DB and then passed to the orchestrator. The orchestrator determines the relevant script for the prompt (in this case, the email interrogation script), saves the determination to the DB, and triggers execution of the script. In certain embodiments, the email interrogation script queries the DB for the specific email ID, passes the message body with the user prompt to the LLM for interrogation, saves the response to the DB, and sends the response as a reply to chatbot.Example 5: Visualization. The visualization script might be triggered if the user enters a prompt such as “Create a conversation cluster visual for all users in the review set where each circle should represent a user, and the more commination to a user, the larger their ‘circle’ should be." As discussed above, the prompt is saved to the DB and then passed to the orchestrator. The orchestrator determines the relevant script for the prompt (in this case, the visualization script), saves the determination to the DB, and triggers execution of the script. In certain embodiments, the visualization script provides the LLM with access to the DB for Retrieval- Augmented Generation (RAG), passes the user prompt to the LLM, saves the response from the LLM to the DB (e.g., in Python code), and sends the response to the chatbot as a reply (e.g., in Python code) for rendering by the chatbot for the user to see the visualization.Example 6: Report Generation. The report generation script might be triggered if the user enters a prompt such as “Please create a report summarizing the key emails to the fraud case and give reasoning on why each email is relevant with evidence." As discussed above, the prompt is saved to the DB and then passed to the orchestrator. The orchestrator determines the relevant script for the prompt (in this case, the report generation script), saves the determination to the DB, and triggers execution of the script. In certain embodiments, the report generation script provides the LLM with access to the DB for Retrieval- Augmented Generation (RAG), passes the user prompt to the LLM, saves the response from the LLM to the DB, and sends the response to the chatbot as a reply.Some possible uses of the described embodiments include data visualization report generation, security anomaly categorization, etc.Some possible advantages of the described embodiments include that the search is faster than manual reviewing of data, searches are more affordable than manual reviewing of data, there is no upramp time training ML models, the system can provide rationale for itscategorizations in comparison to a ML model, provides more flexibility / agility versus use of a specific ML model, provides an integrated user experience with one to many backend services chat experience for the user, meta / mixture of LLM experts architecture, simplification of interaction with LLM, etc. The increased speed and simplicity can help resolve cases faster, which can minimize damage (e.g., emotional, financial, reputational, brand / goodwill, customer loss, data loss, compliance / penalties, etc.) for those involved. The LLM can be pretrained (e.g., for categorizing emails and for selecting scripts based on prompts) for increased efficacy of LLM and reduced hallucinations.In some embodiments, the chatbot includes an efficacy engine that can generate alternate and / or optimized prompts for a given user prompt (e.g., through “prompt engineering” techniques such as specifying assumptions, rewards, or constraints in the prompt), either automatically or when prompted by the user. For example, imagine emails are presented to the LLM with a prompt to categorize if the emails are relevant or not relevant for a given use case (for example, look at this email and indicate if it has something to do with fraud and reply either relevant or not relevant with a confidence score). The efficacy engine can evaluate the user prompt and generate one or more variants (where the number of variants may be fixed or configurable) to optimize the efficiency of the prompt to find relevant data.The variant prompt(s) can be executed in addition to, or in lieu of, the user prompt. The outputs of multiple prompts can be compared (e.g., against the other outputs and / or against a baseline) to identify a “best” prompt, which, for example, can then be either shown to the user to run on data or automatically executed on data, e.g., based on configuration. The user may be permitted to configure the number of variants it wants generated. The engine also can be configured to run each prompt multiple times (which may be fixed or configurable) and take the average to improve reliability. In essence, the efficacy engine is an automated “helper” that can speed up finding the ‘needle in the haystack’ without the user needing to be an expert in LLM / can trust the LLM more.The following scenario demonstrates some of the functionality and benefits of the described embodiments. Say that an investigation is started to find some insider fraud alleged between a few employees of an organization over the last three years. This is very much a classic ‘needle in a haystack’ E-discovery investigation. An investigation of this type generally would start with a search being run against all the alleged employees’ mailboxes for the last three years (say, for example, that the search results in 100,000 emails) then adding those emails to a reviewset for the legal department to process. The process below describes the traditionalapprove compared to the new approach. The following is a summary of this traditional approach to E-discovery:This traditional approach is time and cost intensive. Using the above example, assuming it takes 1 minute to review each email, this would result in -1600 hours of work to process the entire reviewset (or, put another way, this would take a team of 10 a month to get through the entire set). Assuming a per hour cost of the lawyers to be $50 (which is very conservative), the cost would be -$80,000. The following is an example of using conversation search tools to perform the same analysis:This is fast and affordable. Assuming it takes 1 second to review an email for categorization (which is a conservative estimate), this would result in -27 hours to categorize the 100,000 emails of the initial reviewset (not factoring in that the system can have parallel reviews with the LLM happening at once to make it much faster). Assuming 1 second to review each of the 20,000 filtered emails for relevance, this would result in ~5 hours (which, again, could be much faster with parallel jobs) if a subject matter expert reviewed just 500 emails in 100 data sets at a time for 5 prompt optimizing loops. Assuming LLM review is $0.1 per email (which is estimated to be at the high end), the total cost of this investigation would be -$13,600, which isaround 16% of the cost of the traditional way and would be done in around 3% of the time (just a matter of days).Thus, some benefits of the conversation search with efficacy engine include:1. Significant time savings through automation of initial review and prompt optimization.2. Reduced reliance on expert prompt engineers or extensive legal team for effective searches while allowing optimization of prompts.3. Improved consistency and reliability in document classification.4. Greater transparency and auditability of the eDiscovery process.5. Continuous improvement of search efficacy across multiple cases and over time.6. Potential for cost savings due to reduced manual labor and faster processing times.7. Ability to test different LLM models for accuracy and efficacy.These improvements could lead to faster resolution of cases, more thorough eDiscovery processes, and potentially better legal outcomes due to more comprehensive and accurate document review.It should be noted that, while various embodiments are described herein with reference to a reviewset including email messages and documents, alternative embodiments could apply to other types of data in addition to or in lieu of email messages and / or documents, such as, for example, text messages, social media messages, direct messages, in-app messages, etc. Thus, for example, embodiments can act on various types or combinations of types of data.Various embodiments of the invention may be implemented at least in part in any conventional computer programming language. For example, some embodiments may be implemented in a procedural programming language (e.g., “C”), or in an object-oriented programming language (e.g., “C++”). Other embodiments of the invention may be implemented as a pre-configured, stand-alone hardware element and / or as preprogrammed hardware elements (e.g., application specific integrated circuits, FPGAs, and digital signal processors), or other related components.In alternative embodiments, the disclosed apparatus and methods (e.g., as in any flow charts or logic flows described above) may be implemented as a computer program product for use with a computer system. Such implementation may include a series of computer instructions fixed on a tangible, non-transitory medium, such as a computer readable medium (e.g., a diskette, CD-ROM, ROM, or fixed disk). The series of computer instructions can embody all or part of the functionality previously described herein with respect to the system.Those skilled in the art should appreciate that such computer instructions can be written in a number of programming languages for use with many computer architectures or operating systems. Furthermore, such instructions may be stored in any memory device, such as a tangible, non-transitory semiconductor, magnetic, optical or other memory device, and may be transmitted using any communications technology, such as optical, infrared, RF / microwave, or other transmission technologies over any appropriate medium, e.g., wired (e.g., wire, coaxial cable, fiber optic cable, etc.) or wireless (e.g., through air or space).Among other ways, such a computer program product may be distributed as a removable medium with accompanying printed or electronic documentation (e.g., shrink wrapped software), preloaded with a computer system (e.g., on system ROM or fixed disk), or distributed from a server or electronic bulletin board over the network (e.g., the Internet or World Wide Web). In fact, some embodiments may be implemented in a software-as-a-service model (“SAAS”) or cloud computing model. Of course, some embodiments of the invention may be implemented as a combination of both software (e.g., a computer program product) and hardware. Still other embodiments of the invention are implemented as entirely hardware, or entirely software.Computer program logic implementing all or part of the functionality previously described herein may be executed at different times on a single processor (e.g., concurrently) or may be executed at the same or different times on multiple processors and may run under a single operating system process / thread or under different operating system processes / threads. Thus, the term “computer process” refers generally to the execution of a set of computer program instructions regardless of whether different computer processes are executed on the same or different processors and regardless of whether different computer processes run under the same operating system process / thread or different operating system processes / threads. Software systems may be implemented using various architectures such as a monolithic architecture or a microservices architecture.It should be noted that terms such as “computer,” “client,” “server,” may be used herein to describe devices or systems that may be used in certain embodiments of the present invention and should not be construed to limit the present invention to any particular device or system type unless the context otherwise requires. Thus, a device or system may include, without limitation, a node, server, computer (including desktop, laptop, tablet, portable, wearable, etc.), cloud computing platform, appliance, or other type of device or system. Such devices or systems typically include one or more network interfaces for communicating over a communication network and at least one processor (e.g., a microprocessor with memory and other peripherals and / or application-specific hardware) configured accordingly to perform device or systemfunctions. Communication networks generally may include public and / or private networks; may include local-area, wide-area, metropolitan-area, storage, and / or other types of networks; and may employ communication technologies including, but in no way limited to, analog technologies, digital technologies, optical technologies, wireless technologies (e.g., Bluetooth, WiFi, cellular, etc.), networking technologies, and internetworking technologies.It should also be noted that devices and systems may use communication protocols and messages (e.g., messages created, transmitted, received, stored, and / or processed by the device or system), and such messages may be conveyed by a communication network or medium. Unless the context otherwise requires, the present invention should not be construed as being limited to any particular communication message type, communication message format, or communication protocol. Thus, a communication message generally may include, without limitation, a frame, packet, datagram, user datagram, cell, or other type of communication message. Unless the context requires otherwise, references to specific communication protocols are exemplary, and it should be understood that alternative embodiments may, as appropriate, employ variations of such communication protocols (e.g., modifications or extensions of the protocol that may be made from time-to-time) or other protocols either known or developed in the future.It should also be noted that logic flows may be described herein to demonstrate various aspects of the invention, and should not be construed to limit the present invention to any particular logic flow or logic implementation. The described logic may be partitioned into different logic blocks (e.g., programs, modules, functions, or subroutines) without changing the overall results or otherwise departing from the true scope of the invention. Often times, logic elements may be added, modified, omitted, performed in a different order, or implemented using different logic constructs (e.g., logic gates, looping primitives, conditional logic, and other logic constructs) without changing the overall results or otherwise departing from the true scope of the invention.The present invention may be embodied in many different forms, including, but in no way limited to, computer program logic for use with a processor (e.g., a microprocessor, microcontroller, digital signal processor, or general purpose computer), programmable logic for use with a programmable logic device (e.g., a Field Programmable Gate Array (FPGA) or other PLD), discrete components, integrated circuitry (e.g., an Application Specific Integrated Circuit (ASIC)), or any other means including any combination thereof. Computer program logic implementing some or all of the described functionality is typically implemented as a set of computer program instructions that is converted into a computer executable form, stored as such in a computer readable medium, and executed by one or more processors optionally under thecontrol of an operating system. Hardware-based logic implementing some or all of the described functionality may be implemented using one or more appropriately configured FPGAs or other programmable logic devices.Computer program logic implementing all or part of the functionality previously described herein may be embodied in various forms, including, but in no way limited to, a source code form, a computer executable form, and various intermediate forms (e.g., forms generated by an assembler, compiler, linker, or locator). Source code may include a series of computer program instructions implemented in any of various programming languages (e.g., an object code, an assembly language, or a high-level language such as Fortran, C, C++, JAVA, or HTML) for use with various operating systems or operating environments. The source code may define and use various data structures and communication messages. The source code may be in a computer executable form (e.g., via an interpreter), or the source code may be converted (e.g., via a translator, assembler, or compiler) into a computer executable form.Computer program logic implementing all or part of the functionality previously described herein may be executed at different times on a single processor (e.g., concurrently) or may be executed at the same or different times on multiple processors and may run under a single operating system process / thread or under different operating system processes / threads. Thus, the term “computer process” refers generally to the execution of a set of computer program instructions regardless of whether different computer processes are executed on the same or different processors and regardless of whether different computer processes run under the same operating system process / thread or different operating system processes / threads.The computer program may be fixed in any form (e.g., source code form, computer executable form, or an intermediate form) either permanently or transitorily in a tangible storage medium, such as a semiconductor memory device (e.g., a RAM, ROM, PROM, EEPROM, or Flash-Programmable RAM), a magnetic memory device (e.g., a diskette or fixed disk), an optical memory device (e.g., a CD-ROM), a PC card (e.g., PCMCIA card), or other memory device. The computer program may be fixed in any form in a signal that is transmittable to a computer using any of various communication technologies, including, but in no way limited to, analog technologies, digital technologies, optical technologies, wireless technologies (e.g., Bluetooth), networking technologies, and internetworking technologies. The computer program may be distributed in any form as a removable storage medium with accompanying printed or electronic documentation (e.g., shrink wrapped software), preloaded with a computer system (e.g., on system ROM or fixed disk), or distributed from a server or electronic bulletin board over the communication system (e.g., the Internet or World Wide Web).
[0063] Hardware logic (including programmable logic for use with a programmable logic device) implementing all or part of the functionality previously described herein may be designed using traditional manual methods, or may be designed, captured, simulated, or documented electronically using various tools, such as Computer Aided Design (CAD), a hardware description language (e.g., VHDL or AHDL), or a PLD programming language (e.g., PALASM, ABEL, or CUPL).Programmable logic may be fixed either permanently or transitorily in a tangible storage medium, such as a semiconductor memory device (e.g., a RAM, ROM, PROM, EEPROM, or Flash-Programmable RAM), a magnetic memory device (e.g., a diskette or fixed disk), an optical memory device (e.g., a CD-ROM), or other memory device. The programmable logic may be fixed in a signal that is transmittable to a computer using any of various communication technologies, including, but in no way limited to, analog technologies, digital technologies, optical technologies, wireless technologies (e.g., Bluetooth), networking technologies, and internetworking technologies. The programmable logic may be distributed as a removable storage medium with accompanying printed or electronic documentation (e.g., shrink wrapped software), preloaded with a computer system (e.g., on system ROM or fixed disk), or distributed from a server or electronic bulletin board over the communication system (e.g., the Internet or World Wide Web). Of course, some embodiments of the invention may be implemented as a combination of both software (e.g., a computer program product) and hardware. Still other embodiments of the invention are implemented as entirely hardware, or entirely software.While the invention has been particularly shown and described with reference to specific embodiments, it will be understood by persons of ordinary skill in the art that various changes in form and detail may be made without departing from the spirit and scope of the invention as defined by the appended clauses. While some of these embodiments have been described in the claims by process steps, an apparatus comprising a computer capable of executing the process steps is also included in the present invention. Likewise, a computer program product comprising a tangible, non-transitory computer readable medium having embodied therein computer executable instructions for executing the process steps is included in the present invention. Data signals embodying computer program instructions and / or messages received or transmitted over a communication system are also included in the present invention. Unless the context requires otherwise, the various functions and features described herein can be used in combination even if disclosed or claimed individually. Thus, for example, it is contemplated that dependent claims included below could be rewritten into multiple dependent form to depend from the base claim and an intervening claim(s).Importantly, it should be noted that embodiments of the present invention may employ conventional components such as conventional computers (e.g., off-the-shelf PCs, mainframes, microprocessors), conventional programmable logic devices (e.g., off-the shelf FPGAs or PLDs), or conventional hardware components (e.g., off-the-shelf ASICs or discrete hardware components) which, when programmed or configured to perform the non-conventional methods described herein, produce non-conventional devices or systems. Thus, there is nothing conventional about the inventions described herein because even when embodiments are implemented using conventional components, the resulting devices and systems are necessarily non-conventional because, absent special programming or configuration, the conventional components do not inherently perform the described non-conventional functions.The activities described and claimed herein provide technological solutions to problems that arise squarely in the realm of technology. These solutions as a whole are not well- understood, routine, or conventional and in any case provide practical applications that transform and improve computers and computer routing systems.While various inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.Various inventive concepts may be embodied as one or more methods, of which examples have been provided. The acts performed as part of the method may be ordered in anysuitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements andnot excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.As used herein in the specification and in the claims, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.Although the above discussion discloses various exemplary embodiments of the invention, it should be apparent that those skilled in the art can make various modifications that will achieve some of the advantages of the invention without departing from the true scope of the invention. Any references to the “invention” are intended to refer to exemplary embodiments of the invention and should not be construed to refer to all embodiments of the invention unless the context otherwise requires. The described embodiments are to be considered in all respects only as illustrative and not restrictive.
Claims
1. What is claimed is:
1. A conversation search system comprising. a chatbot comprising a natural language user interface through which a user can enter natural language prompts for a conversation search of a reviewset containing data items for review; a plurality of scripts including at least: a first pass script configured to characterize and label the data items of the reviewset to filter irrelevant data items out of the reviewset; and a second pass script configured to predict relevance of each data item of the reviewset remaining after the filtering of the first pass script; and an orchestrator that selects and triggers execution of the scripts based on the natural language prompts.
2. The system of claim 1, wherein the data items include at least one of messages or documents.
3. The system of claim 1, wherein the natural language user interface uses a large language model (LLM) to process user prompts.
4. The system of claim 1, wherein the natural language user interface allows natural language prompts to be entered textually.
5. The system of claim 1, wherein the natural language user interface allows natural language prompts to be entered verbally.
6. The system of claim 1, wherein the first pass script uses natural language processing (NLP) to characterize and label the data items.
7. The system of claim 1, wherein the second pass script uses natural language processing to predict relevance of each data item.
8. The system of claim 7, wherein the natural language processing uses a Bert model.
9. The system of claim 8, wherein the Bert model includes one of a fraud detection model, a harassment model, or a compliance model.
10. The system of claim 1, wherein the chatbot includes an efficacy engine that automatically generates one or more additional prompts for a given user prompt.
11. The system of claim 10, wherein the number of additional prompts is fixed.
12. The system of claim 10, wherein the number of additional prompts is user configurable.
13. The system of claim 10, wherein at least one of the additional prompts is a variant of the given user prompt.
14. The system of claim 10, wherein the chatbot is configured to use prompt engineering techniques to automatically generate the one or more additional prompts.
15. The system of claim 10, wherein the chatbot executes the one or more additional prompts in addition to the given user prompt.
16. The system of claim 10, wherein the chatbot executes the one or more additional prompts in lieu of the given user prompt.
17. The system of claim 10, wherein the chatbot executes a plurality of prompts from among the given user prompt and the one or more additional prompts and identifies a best prompt from among the plurality of executed prompts.
18. The system of claim 17, wherein the chatbot identifies the best prompt by comparing an output of each executed prompt to a baseline.
19. The system of claim 18, wherein the chatbot identifies the best prompt by comparing outputs of the plurality of executed prompts.
20. The system of claim 1, wherein the plurality of scripts includes an interrogation script to interrogate a specified data item.
21. The system of claim 1, wherein the plurality of scripts includes a report generation script.
22. The system of claim 21, wherein the report generation script produces a retrieval- augmented generation (RAG) gen Al generated report.
23. The system of claim 1, wherein the plurality of scripts includes a visualization script.
24. The system of claim 23, wherein the visualization script produces a retrieval-augmented generation (RAG) gen Al generated visualization.
25. The system of claim 24, wherein the visualization is output as a python script that can be rendered in the chatbot user interface.
26. A conversation search method comprising: providing natural language prompts to a conversation search system comprising(a) a chatbot comprising a natural language user interface through which a user can enter natural language prompts for a conversation search of a reviewset containing data items for review;(b) a plurality of scripts including at least: a first pass script configured to characterize and label the data items of the reviewset to filter irrelevant data items out of the reviewset; and a second pass script configured to predict relevance of each data item of the reviewset remaining after the filtering of the first pass script; and(c) an orchestrator that selects and triggers execution of the scripts based on the natural language prompts, the natural language prompts providing search criteria and prompting execution of the first pass script followed by the second pass script.
27. The method of claim 26, wherein the data items include at least one of messages or documents.
28. The method of claim 26, wherein the natural language user interface uses a large language model (LLM) to process user prompts.
29. The method of claim 26, wherein the natural language user interface allows natural language prompts to be entered textually.
30. The method of claim 26, wherein the natural language user interface allows natural language prompts to be entered verbally.
31. The method of claim 26, wherein the first pass script uses natural language processing (NLP) to characterize and label the data items.
32. The method of claim 26, wherein the second pass script uses natural language processing to predict relevance of each data item.
33. The method of claim 32, wherein the natural language processing uses a Bert model.
34. The method of claim 33, wherein the Bert model includes one of a fraud detection model, a harassment model, or a compliance model.
35. The method of claim 26, wherein the chatbot includes an efficacy engine that automatically generates one or more additional prompts for a given user prompt.
36. The method of claim 35, wherein the number of additional prompts is fixed.
37. The method of claim 35, wherein the number of additional prompts is user configurable.
38. The method of claim 35, wherein at least one of the additional prompts is a variant of the given user prompt.
39. The method of claim 35, wherein the chatbot is configured to use prompt engineering techniques to automatically generate the one or more additional prompts.
40. The method of claim 35, wherein the chatbot executes the one or more additional prompts in addition to the given user prompt.
41. The method of claim 35, wherein the chatbot executes the one or more additional prompts in lieu of the given user prompt.
42. The method of claim 35, wherein the chatbot executes a plurality of prompts from among the given user prompt and the one or more additional prompts and identifies a best prompt from among the plurality of executed prompts.
43. The method of claim 42, wherein the chatbot identifies the best prompt by comparing an output of each executed prompt to a baseline.
44. The method of claim 43, wherein the chatbot identifies the best prompt by comparing outputs of the plurality of executed prompts.
45. The method of claim 26, wherein the plurality of scripts includes an interrogation script to interrogate a specified data item.
46. The method of claim 26, wherein the plurality of scripts includes a report generation script.
47. The method of claim 46, wherein the report generation script produces a retrieval- augmented generation (RAG) gen Al generated report.
48. The method of claim 26, wherein the plurality of scripts includes a visualization script.
49. The method of claim 48, wherein the visualization script produces a retrieval-augmented generation (RAG) gen Al generated visualization.
50. The method of claim 49, wherein the visualization is output as a python script that can be rendered in the chatbot user interface.
Citation Information
Patent Citations
Methods and apparatus for a knowledge-based deep learning refactoring model with tightly integrated functional nonparametric memory
US20220101096A1
Apparatus and method for generating a digital assistant
US20240126794A1
Methods and systems for improved document processing and information retrieval
WO2024015321A1